Lexicoblog

The occasional ramblings of a freelance lexicographer

Monday, January 07, 2013

The future of dictionaries (2): lexicographer versus computer



Some 20-odd years ago, as a young, Linguistics undergraduate, I became interested in the concept of computers ‘understanding’ human language. I did my undergraduate dissertation on Natural Language Processing (NLP), considering how far computers might go in really understanding language in all its subtle, complex, nuanced detail, and holding up the talking computer Hal, from Kubrick’s 1968 film 2001, as what I then suggested was an unobtainable goal. I went on to start a Master’s course in Computer Speech and Language Processing. I only lasted a term – mainly because I discovered I really hated all the computer programming involved, but also because I was disappointed to find that most of the course seemed to revolve around the speech processing side (i.e. voice recognition) and the language processing component came down to rather vague theoretical discussions that didn’t go much beyond my basic undergraduate research. Okay, that may be a bit of a distorted recollection of the actual course content, but I was only 21 and that’s how you see things when you’re barely out of your teens!

Obviously, in the intervening decades, technology has come on in leaps and bounds. Speech recognition has improved immeasurably - I'm actually dictating this blog post using speech recognition software and while it's not perfect, it's considerably more impressive than my early efforts are programming! I have to hold up my hands here and admit that I haven't kept up-to-date with developments in NLP, but I suspect progress has been much slower; we're still an incredibly long way from communicating with our technology in the same fluent way we can chat to our friends.

So, what’s any of this got to do with dictionaries? Well, let me try and explain my train of thought, triggered by the announcement by Macmillan back at the start of November that they are to stop printing paper dictionaries and focus on their online content:

  • If publishers aren't actually selling paper dictionaries but are mostly focusing on a free online service, how much are they going to be prepared to spend on the time-consuming and labour-intensive work of lexicography?
  • Of course, they'll be looking into other related income streams, selling dictionary data for other uses, and online advertising, but without a tangible, on-the-shelf product, will that justify quite the same budget?
  • Reduced budgets often suggest a drive towards more automation, something we've already seen with the emergence of developments such as ‘TickBox lexicography’.
  • Will more automation and “more efficient” ways of working inevitably lead to a drop in standards?

Clever developments in making the dictionary compilation process more automated do supposedly speed it up, for example, by automatically selecting ‘good’ dictionary examples from a corpus, to save a human lexicographer having to trawl through by hand. But any lexicographer who's worked with them will know that they only work to a degree and only speed things up to a certain extent … probably not quite compensating for the increased rate expected of said lexicographer without a drop in quality.

And then there's the whole established process for keeping dictionaries up-to-date. Currently, most dictionaries undergo a revision and a new edition every five years or so. This is a long, slow, and labour-intensive process that involves a team of lexicographers (mostly freelancers nowadays) going through the whole A-Z, looking at each entry and checking whether it needs updating. This doesn’t just involve adding trendy new buzzwords like ‘omnishambles’ or whatever – which are rarely of much use, or interest, to the average foreign learner anyway. There are all kinds of more subtle changes in the usage of existing words, sometimes due to linguistic trends and sometimes just as a result of changes in the real world. As one commenter on the Macmillan dictionaries blog pointed out, MED still contains an entry for Inland Revenue as the name of the UK tax authority, even though it changed its name to HMRC in 2005. And having done a quick search myself, I found it also has a couple of example sentences that rather unhelpfully in a digital age refer to cassettes (She slotted another tape into the cassette player. @ slot into, He quickly undid the screws that held the cassette together. @ undo).
 And I’m not just trying to pick holes in Macmillan here; all dictionaries naturally date as language and usage changes. Thus the need for new editions. And there are changes in style and presentation too as different aspects of language come to the fore within language research and teaching. More information about collocations has become de rigueur over recent years, for example. And whilst corpora are wonderful tools for researching collocational information, it still needs a team of lexicographers to trawl through each entry and decide where it’s worth adding a bolded collocate, or in some cases, whether a particularly strong collocation should actually be shown as a phrase or an idiom.

Which comes back to where I started … computer technology can do lots of wonderful things, but for me, when it comes to language, there still needs to be a human drudge working their way through that data to make intelligent decisions about what to present in a dictionary and how. In a world of online-only dictionaries, will dictionary departments have the clout to take on a team of lexicographers to do those regular sweeps through the database or will they just have a couple of people on the lookout for interesting, newsworthy nuggets that give the appearance of being “up-to-date”?

Labels: , , , ,

Monday, January 30, 2012

Corpus frequencies: what exactly counts?

Senses, idioms and phrasal verbs

Following on from my last post prompted by Michael Rundell’s webinar: Tweets, blogs and corpora: How computer technology helps us make better dictionaries. There was one more question from a webinar participant which I think opens up a whole area of corpus research and word frequencies as shown in dictionaries that tends to get glossed over.

“Do these numbers [corpus frequencies] consider all the meanings of a word or only the common ones?”

Response from another participant:
“I’d guess that the counting program [the corpus software] doesn’t understand the meaning so it is for all meanings of the word.”

An astute question and a correct answer! Corpus software is very clever at number crunching and identifying patterns, but computers still fall down when it comes down to actually understanding language. When you do a corpus search, you can choose the part of speech you’re interested in (separating out noun and verb senses of a word like walk, for example) and you can search for a ‘lemma’ rather than just a string of letters (so searching for the verb walk will include walk, walks, walked and walking). When it comes down to differentiating between different senses or uses of a word though, that can still only be done “by hand” by a human being sorting through a sample of corpus lines one-by-one. Sometimes, where one sense is overwhelmingly more frequent, the sense frequencies are obvious at a glance. In other cases, especially with very polysemous words, it’s a trickier business. Thus, sense ordering by frequency is, to a degree, impressionistic and doesn’t involve exact statistics.

Does this matter? Again, my answer is “not really”, provided we’re only taking frequency information in a dictionary as a general guide. For many words, the most frequent sense(s) of a word will probably account for the majority of its occurrences, so it’s fair to say that overall its core meaning(s) will fall within a general frequency band. It’s unlikely that in many cases there will be lots of obscure senses of a word that significantly distort the frequency statistics.

Where caution may be required though is where it’s the less frequent senses you’re actually interested in. To take an example I came across recently working on EAP vocabulary, if you do a corpus search for chemist, physicist and biologist, chemist comes up as much more frequent – as reflected in most of the learner’s dictionaries. Now that isn’t because there are more scientists studying chemistry than there are physics or biology. But of course, in British English at least, a chemist can be a pharmacy or a pharmacist as well as a scientist in a lab with their test tubes.

And it’s not just the effect that different senses of a word might have on frequency that needs considering. Another big area to take into account is words that form part of a phrase of some kind. Going back to my example of walk, it crops up in various phrases or idioms – walk the walk, run before you can walk, walk free, etc. – and a whole list of phrasal verbs – walk away with, walk in on, walk off, walk out … In most learner’s dictionaries, these come at the end of the entry for the headword and in most of the major dictionaries (I think with the exception of Cambridge), they don’t have frequency information in their own right. Instead, they get lumped into the overall frequency for the whole entry. This has two consequences; firstly, it means that learners can’t see which phrases and phrasal verbs are most frequent and also, it further undermines the frequency information for some words as we can’t be certain what it’s referring to. Take the verb deal as an example, highlighted as frequent in most dictionaries. In fact, something like 85% of occurrences of the verb deal are actually instances of the phrasal verb deal with. Yet, in most dictionaries, it appears that the basic verb senses (giving out cards or drugs) are common, while the phrasal verb deal with has no highlighting at all.

It is possible to construct corpus searches to find particular phrases or phrasal verbs, even where their form varies slightly - for example, where phrasal verbs have moveable particles. So it is possible to get frequency information for them, albeit not as simply or reliably as for single words. And if you look in specialist dictionaries of phrasal verbs or idioms, you’ll often find the most common ones highlighted. So why don’t most general learner’s dictionaries include this information? Well, firstly, it’s very time-consuming to research and secondly, it isn’t easy within a traditional dictionary format to devise a system that encompasses frequency information for both whole words (with senses lumped together) and individual usages in the form of phrases and phrasal verbs.

So having completely ripped apart the frequency information in dictionaries, am I saying that it’s useless and should be ignored? No, far from it! I think as a broad guide to which words are generally more frequent (and so worth focusing on), I still think it’s an incredibly useful tool. But as in any area of life, statistics should always be approached critically and before you rely too much on them, you need to understand what’s behind them, how they’re compiled and what caveats you might need to take into consideration.

Labels: , , , ,

Friday, January 27, 2012

Corpus: gospel or guide?

A response to a webinar:

Lately, I’ve been starting to explore Twitter and linking through to various blogs and websites to see what folks are talking about in ELT at the moment, all in the name of “professional development”. Yesterday, I also ventured into the world of webinars for the first time. I started off with a recording from Macmillan’s Interactive webinars series. My first aim was really just to explore the medium, so I chose a familiar topic with Michael Rundell’s Tweets, blogs and corpora: How computer technology helps us make better dictionaries. The actual content wasn’t particularly exciting - not because there was anything wrong with Michael’s presentation, but as an experienced lexicographer, I clearly wasn’t his target audience. That meant though that I could focus more on what actually goes on in a webinar.

Rather unexpectedly, I got particularly caught up with the reactions of the participants as they appeared in the little text box on the side of the screen. Unfortunately, Michael ran out of time, so didn’t get to address any of the comments or questions that popped up. I, however, was itching to respond to them! So I thought I’d tackle some of the points here which I’ve been mulling over since. And in fact, I’ve got so much to say, I’m going to split this into two posts.

The part of the webinar that interested me most in terms of participant feedback was when Michael was talking about how we use frequency information from corpora to highlight the most common and so "useful" words in a dictionary. Below are some of the comments and questions and my reactions:

“Are there standard lists with the top 250 words?”
“But where can we find the list of words (to know if they are frequent or not)”

I always find interest in wordlists from teachers and students a little bit worrying. It seems to suggest that language learners are rather like computers and if we can just input the right list of words, then they’ll output English at a given level! Whilst I think frequency lists can have a role to play in helping prioritise what to focus on, my feeling is that generally checking frequency should be something that comes after you encounter new vocabulary. You look up a new word you’ve come across in the dictionary and you might use the information about frequency to decide whether it’s worth putting in your vocabulary notebook or whether it’s a word that you can naturally drop into conversation or not. Language is a wonderfully messy, organic, personal sort of a thing and what vocabulary you choose to teach or learn should be governed by all sorts of different factors - interests, needs, context, personality - not some (inevitably very dull) standard list of frequent words.

“Are the top 3000 words in Oxford the same top 3000 words in Macmillan?”

I haven’t researched the answer to this one, but I think I can fairly confidently say “more or less” if we’re just talking about frequency (more on that below). Each of the major dictionary publishers uses a different corpus – or rather a different collection of corpora, some of which overlap (like the BNC). In the early days, with relatively small corpora, you would have expected some variation, with different corpora slightly skewed towards particular types of language. Nowadays though, with all the big publishers using really huge and diverse collections of corpora, I think you’d probably expect a straightforward frequency list (at least at the most frequent end) to come out more or less the same, with only minor variations.

Having said that, each dictionary publisher has it’s own criteria for how it shows frequency information – where it sets it’s limits and how it puts words into frequency bands. The Oxford 3000™, for example, isn’t just based on frequency, but was put together using three criteria; frequency, range and familiarity (if you're interested, you can read more about it here). Does this variation matter? Personally, I don’t think so. How many students, or even teachers, ever read the blurb in the front (or back) of a dictionary that explains the frequency information? My feeling is that most students either don’t even notice it, or if they do, it’s just some general sense that a word is highlighted or has stars next to it, therefore it must be useful to learn. Of course, there’ll be times when some teachers (esp. vocab nerds like myself!) will make a point in class about frequent and more marked synonyms by pointing to the frequency information (and often register labels) in the dictionary. But my feeling is that’s the exception, not the rule. And that’s fine. It’s still worth the information being there as one more tool in the language learning toolbox.

Coming back to the title of this post, corpus information has been incredibly useful over the past couple of decades in understanding how language is actually used and in making teaching materials more natural, but it's still only a guide. Despite much of my work being in the area of corpus research, I'm still very wary about taking corpus data as gospel, partly for some of the reasons I'll talk about in my next post ...

Labels: , , , ,