<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>Book Reviews: Lexicography and Natural Language Processing: A Festschrift in Honour of B. T. S. Atkins edited by Marie-Hélène Corréard</title><author surname="Haynes" givenname="Woody"><org  name="URA" country="France" city="Marseille"/></author><author surname="Evens" givenname="Martha"><org  name="Illinois Institute of Technology" country="USA" city="Chicago"/></author></firstpageheader><frontmatter><p><b>Book Reviews</b></p><p><b>Lexicography and Natural Language Processing: A Festschrift in Honour of B. </b><b>T.</b><b> S. Atkins</b></p><p><b>Marie-Heiene Correard (editor)</b></p><p>Grenoble, France: EURALEX, 2002, viii+247 pp; paperbound, ISBN 2-9518583-0-2, €30.00 (www.ims.uni-stuttgart.de/euralex)</p><p><i>Reviewed by</i></p><p><i>Woody Haynes and Martha Evens Illinois Institute of Technology</i></p></frontmatter><abstract></abstract></header><body><section title=""><p>This volume is a festschrift for Sue (B. T. S.) Atkins, who is perhaps best known for her work as general editor of the <i>Collins-Robert English/French Dictionaries </i>and a consultant-advisor for the <i>Oxford-Hachette English/French Dictionary, </i>which led to a revolution in the construction of bilingual dictionaries. She has also done a great deal to bridge the gap between professional lexicography, academic linguistics, and computational lin­guistics. Most recently she has been working with Charles Fillmore to build FrameNet and to adapt his ideas for use in a dictionary framework. The lead-off paper in this volume is Atkins's keynote address from the 1996 EURALEX meeting, the inspira­tion for this volume. The contributions of the other authors, all of them new papers, examine how lexicography is responding to Atkins's call for a "radical new type of dictionary" in her 1996 address. They consider the current state of dictionaries, dis­cuss their strengths and weaknesses, and describe new computational tools that are facilitating exploration in new directions and providing new insights.</p><p>The increasing availability of diverse electronic text has had a profound effect on the process of creating dictionaries, both in the compilation process and in the depth and structure of the result. It is true that most of the papers in this volume talk about how to use natural language processing in lexicography rather than how to make use of various lexical resources in natural language processing. But we believe that the volume includes much useful material for all those whose work in computational linguistics makes them consumers of lexical resources, since it discusses many of the current ideas about how those resources should be constructed. Even those papers that focus on bilingual resources provide many interesting ideas about the practice of lexicography today. Most of the recommendations for bilingual dictionaries can equally address issues involved in the use of machine-readable dictionaries of all kinds and apply to deficiencies seen in all available lexicons.</p><p>Atkins's keynote address, entitled "Bilingual Dictionaries: Past, Present and Fu­ture," looks at the various types of information available and needed in various types of bilingual and monolingual dictionaries, categorizes it, and argues for a truly elec­tronic dictionary that can adapt itself to the needs of the "multifarious users." Atkins identifies the strengths of current dictionaries in their wealth of information, scholarly work, and concern for the needs of the dictionary user. She sees weaknesses in the re­dundancy, coverage gaps, inflexible equivalence and collocational selection, distortion caused by disparate needs of source and target languages and by monolingual information omitted from bilingual dictionaries, inextensibility of bilingual dictionaries to multilingual dictionaries, lack of integrated thesaural functions, and the user learning curve for dictionary metalanguage.<page local="2" global="318"/></p><p>In "Use and Usability of Dictionaries: Common Sense and Context Sensibility?" Krista Varantola discusses the disparate needs of lay dictionary users and language professionals. She suggests adapting frame semantics, as proposed by Fillmore and Atkins (1998), to facilitate tailoring the electronic dictionary to give users what they need in terms they understand, relying less on context-free, impenetrable text definitions.</p><p>Alain Duval, in "La metalangue, un mal necessaire du dictionnaire actif," the only paper in the volume not in English, addresses the problems of communicating with the user of a bilingual dictionary. The new bilingual dictionary tries to function as an "active dictionary" that supports the user who is trying to generate text in a second language, while still doing the job of the "passive dictionary" that helps the user who is merely trying to understand that language. Duval illustrates the differences between the old and new with examples from several older bilingual dictionaries and points out the advantages of the new approach for users as well as the demands on the user who must understand the expanded metalanguage.</p><p>In "Word Groups in Bilingual Dictionaries: OHFD and After," Richard Wakely and Henri Bejoint describe their approach to usage notes in the <i>Oxford-Hachette French Dictionary. </i>They discuss their method of identifying lexical sets exhibiting sufficient size, frequency, and behavior commonality. These sets could then be described once with the entries for each set member pointing to the page containing the usage note. They note the pluses and minuses of this approach in terms of practicality, convenience, and usability.</p><p>In "Examples and Collocations in the French 'Dictionnaire de langue,'" A. P. Cowie, current editor of the <i>International Journal of Lexicography, </i>looks at the treatment of examples in a number of French monolingual dictionaries, including <i>Dictionnaire du francais contemporain, Le Petit Robert, Le Grand Robert, </i>and <i>Le Tresor. </i>He contrasts their methods of blending examples constructed by lexicographers with quotations, exact or adapted. He contends that "the richness, diversity and fitness for purpose of exam­ples in <i>Le Grand Robert </i>and <i>Le Tresor, </i>especially, are among the finest achievements in modern lexicography."</p><p>Juri Apresjan has been a leading figure in lexicography in the Soviet Union and Russia for over 30 years, since he worked with Igor Mel'cuk on the development of the <i>Explanatory-Combinatory Dictionary. </i>More recently he has been head of the major Russian machine translation project. In his paper, "Principles of Systematic Lexicogra­phy," he argues for the importance of building a systematic lexicon that can interact effectively with a system of grammar rules in the ECD tradition, and he sketches a linguistic basis for this effort.</p><p>Charles Fillmore, the creator of frame semantics and the father of the FrameNet lexical resource (Fillmore and Atkins 1998), discusses the problem of "Lexical Isolates," lexical items that "appear to be of unique semantic or syntactic type." He illustrates some of these behaviors with those problem children <i>let alone, mention, else, </i>and <i>ilk.</i></p><p>In "Sketching Words," Adam Kilgarriff and David Tugwell describe their method of identifying English word sketches from a corpus with part-of-speech tags and a shallow parse, producing an automatic summary of a word's behavior that can assist lexicographers in describing that behavior and can help NLP systems subsequently to perform word sense disambiguation reliably. Each word sketch consists of one of twenty-six word relations, with one, two, or three operands. The salience of a word sketch is defined as a function of mutual information and log frequencies.</p><page local="3" global="319"/><p>In "Good Old-Fashioned Lexicography: Human Judgment and the Limits of Au­tomation," Michael Rundell, editor-in-chief of the <i>Macmillan English Dictionary, </i>consid­ers whether the advances in automation of dictionary development will lead to the demise of the lexicographer. He argues against this view with several compelling ex­amples, suggesting that each advance identifies new layers of complexity that depend on the lexicographer for analysis.</p><p>Patrick Hanks, lead editor on a number of Collins and Oxford dictionaries, sug­gests, in "Mapping Meaning onto Use," that frame semantics provides a "richer schema for representing meaning than is used in any current dictionary." He advo­cates using a syntagmatic organizing principle in adjective and verb dictionary entries "rather than (or rather, in tandem with) perceived meaning."</p><p>Gregory Grefenstette is probably best known for his work on cross-language in­formation retrieval and its application to Internet text. Here he presents "The WWW as a Resource for Lexicography" in a wide range of languages and the tools needed to extract lexical information effectively. He argues that it is feasible to port a number of tools such as shallow parsers to other languages, especially those using some variant of the Roman alphabet, and that it is time to get to work on this project, as significant amounts of text begin to appear in a number of previously unrepresented languages.</p><p>Most of the papers in the volume view natural language processing as a tool for building lexicons. In "Lexical Knowledge and Natural Language Processing," Thierry Fontenelle talks about what is needed in a lexical database to support nat­ural language processing and discusses where that knowledge can be found in exist­ing lexical resources, especially collocational dictionaries, thesauri, and semantic net­works. This leads naturally to the problems of representing knowledge about verb alternations and other collocations, using the lexical functions of the <i>Explanatory-Combinatory Dictionary </i>(Apresjan, Mel'cuk, and Zholkovsky 1970) and Fillmore's frame semantics.</p><p>Annie Zaenen, principal scientist and area manager for Multilingual Theory and Technology at the Xerox Research Centre in Grenoble and co-author of several books about lexical-functional grammar and natural language understanding, makes a con­vincing case for a depressing conclusion in her "Musings about the Impossible Elec­tronic Dictionary." She looks at the complexity and pressures stifling progress in the creation of multifunctional lexicons and concludes that current trends will continue to produce disparate resources for disparate consumption rather than a unified lexical database.</p><p>This book is a EURALEX production in every way, and it is certainly a success. Anyone interested in lexicography should read this volume. It might have been even better, however, if the editors had given some of Sue Atkins's many admirers on other continents a chance to join in. The occasional typographical error should certainly be overlooked in view of the bargain price, which should allow many readers to buy copies of their own.</p></section><references><p>relatively thorough index of terms and rich bibliographic references at the end of each chapter.</p><p>Bean, Carol A. and Rebecca Green, editors. 2001. <i>Relationships in the Organization of Knowledge. </i>Dordrecht, Kluwer Academic Publishers.</p><p>Casagrande, Joseph B. and Kenneth L. Hale. 1967. Semantic relations in Papago folk-definitions. In Dell H. Hymes and</p><p>William E. Bittle, editors, <i>Studies in Southwestern Ethnolinguistics: Meaning and History in the Languages ofthe American Southwest. </i>Mouton, The Hague, The Netherlands, pages 165-193. Fellbaum, Christiane, editor. 1998. <i>WordNet: An Electronic Lexical Database. </i>MIT Press, Cambridge, MA.</p><p><i>Maria Lapata </i>is a research fellow at the University of Edinburgh, Division of Informatics. Her research interests include semantic knowledge acquisition, linguistically informed statistical methods for ambiguity resolution, and computational psycholinguistics. Lapata's address is Division of Informatics, University of Edinburgh, 2 Buccleuch Place, Edinburgh EH3 9LW, U.K.; e-mail: mlap@inf.ed.ac.uk.</p><page local="12" global="328"/><p><b>Recent Advances in Computational Terminology</b></p><p><b>Didier Bourigault, Christian Jacquemin, and Marie-Claude L'Homme (editors)</b> (University Toulouse-le-Mirail, CNRS Orsay, and University de Montreal)</p><p>Amsterdam: John Benjamins (Natural language processing series, edited by Ruslan Mitkov, volume 2), 2001, xviii+379 pp; hardbound, ISBN <i>Reviewed by Robert Gaizauskas University of Sheffield</i></p><doubt alpha="0.0" length="21" tooSmall="False" monospace="0.0">1-58811-016-8, $99.00</doubt><p>This collection of papers derives from the <i>Proceedings of the First Workshop on Compu­tational Terminology </i>(Computerm '98), held at COLING-ACL '98 in Montreal, but is a substantial revision thereof. The current volume comprises seventeen papers plus a brief introduction by the editors. The original workshop proceedings also had sev­enteen papers. However, seven of these original papers have disappeared, and seven new papers have taken their place. Furthermore, of the remaining papers, most have been significantly extended. Thus, this book should not be thought of as a simple reissue, in hardcover, of the workshop proceedings.</p><p>The words <i>Recent Advances </i>in the title might be taken to suggest brave strides forward in a clear-cut research program. Nothing could be further from the truth. This is an area of largely pretheoretical research, in which researchers are struggling bravely to use computational techniques to gain some foothold in dealing with the protean complexities of real lexical usage in a variety of technical domains and in a variety of applications. Consequently the book reads a bit like the conversation of the proverbial blind men feeling an elephant, each describing the part he is feeling. This is not meant to be a criticism, for probably nothing else is possible at this time, and besides, this tends to be a feature of edited collections. It does mean, however, that a reader should not come to this volume expecting to find a coherent account of the research issues and approaches in computational terminology. There should be something in here for everyone with any interest in terminology; the danger, however, is that there may not be a meal for anyone.</p><p>Classifying the work reported in this volume is not easy. I shall cluster the pa­pers along two dimensions, a major dimension—the task or problem addressed—and a minor dimension—the intended application. This crude structuring should help to convey some notion of the scope and content of the work. In order of ascending com­plexity the problems addressed by papers in the collection can be characterized as (1) term extraction—the problem of extracting a list of all and only the terms from texts in a given domain, (2) synonymy detection or semantic clustering—the problem of recog­nizing which terms are synonyms or belong to the same semantic class or cluster, and (3) term-oriented knowledge extraction from text—the problem of building knowl­edge structures in technical domains, identifying the underlying conceptual entities, attributes, and relations via terminology. The principal application areas addressed by the papers are information retrieval, terminology construction and maintenance, machine translation, automatic index extraction, and automatic abstract generation.</p><p>Consider first term extraction, the most basic of the three preceding tasks. Au­tomatic term extraction has potential application in automatic indexing, either for back-of-book indices or for document-collection navigation, and also for compiling controlled vocabulary terminologies such as are used in, for example, medical cod­ing applications.<page local="13" global="329"/> Several papers address this topic. Most generically, a review paper by M. Teresa Cabre Castellvî, Rosa Estopa Bagot, and Jordi Vivaldi Palatresi reviews twelve current term extraction systems, including well-known systems such as LEX-TER, FASTR, TERMIGHT, and TERMS, giving a brief description of each, as well as a contrastive analysis. A paper by Lee-Feng Chien and Chun-Liang Chen addresses the problem of incremental update of domain-specific Chinese term lexicons from on-line news sources. Terms are identified and allocated in real time to topic-specific lexicons corresponding to news categories, using highly efficient data structures called PAT trees. To be acceptable for a specific lexicon a term must be <i>complete </i>(have no left or right context dependency and have an internal association norm above threshold) and must be <i>significant </i>(have a relative frequency in a document collection corresponding to the target lexicon that compares favorably to its relative frequency in a general reference collection). Beatrice Daille supplies a linguistically interesting paper on rela­tional adjectives as signals of terms in French scientific text. Contrasting, e.g., <i>production importante </i>('significant production') with <i>production laitière </i>('dairy production'), she ar­gues that in the latter type of construction, such relational adjectives frequently signal terms. She goes on to describe an automatic technique for identifying such terms that is based on looking for paraphrases of the relational adjective + noun expressed as noun + prepositional phrase, where the complement of the preposition is the nominal form of the relational adjective (so, <i>production du lait). </i>A paper by Diana Maynard and Sophia Ananiadou extends their earlier work on term recognition by operationalizing two intuitions about the role of context in termhood: first, that a candidate term that has other candidate terms in its local context is more likely to be a term, and second, that a candidate term that is similar in meaning to domain-specific terms in its local context is more likely to be a term. Toru Hisamitsu and Yoshiki Niwa focus on the specific problem of extracting terms from parenthetical expressions in Japanese news wire text. In expressions of the form <i>A(B), B </i>might or might not be an abbreviated form of A. Segmentation problems in Japanese mean that if <i>B </i>is not correctly recog­nized as an abbreviation it will be oversegmented into single characters, causing real problems for IR systems. Hisamitsu and Niwa propose a neat solution to the prob­lem of identifying which parenthetical expressions are genuine abbreviations that is based on a combination of statistical and rule-based techniques. Finally, a paper by Hiroshi Nakagawa carefully compares two techniques for term extraction, one based on earlier work by Frantzi and Ananiadou and the other an interesting new proposal that assesses termhood according to how productive a noun in a candidate term is in occurring in many other distinct terms.</p><p>The second of the three broad problems or tasks introduced above is the prob­lem of synonymy detection or semantic clustering. Here the problem is not just to discover terms in text, but to relate them in basic ways. Clearly this capability is sig­nificant for information retrieval, in which documents similar in meaning to a query, but differing in expression, must be retrieved. Such a capability is also relevant, how­ever, for automatic index creation and for automatic abstracting. Again, several papers in the collection address this topic. Akiko Aizawa and Kyo Kageura propose a tech­nique that cleverly exploits parallel Japanese-English keyword pairs associated with academic papers to build multilingual semantically related keyword clusters for use in monolingual or cross-lingual IR applications. Peter Anick proposes to use <b>lexical dispersion, </b>a measure of the extent to which a given word is used in multiple NP constructs, to identify generic concepts in retrieval results and to structure these re­sults accordingly. Hongyan Jing and Evelyne Tzoukermann present a stimulating new approach to a classical problem in IR.<page local="14" global="330"/> Intuition suggests that both collapsing vari­ant morphological forms of words and distinguishing different word senses ought to improve retrieval. But previous attempts to do so, through stemming and sense dis­ambiguation, have not led to a reliable increase in performance. The authors present an approach based on full morphological analysis, rather than stemming, and on using context vectors to represent word sense distinctions and to determine whether iden­tical strings in the query and the document should be matched. The approach shows improvement over a more conventional model in which string identity is the only test of synonymy. Thierry Hamon and Adeline Nazarenko start with an existing lexical re­source containing synonym links and a term extractor and use these to bootstrap term synonym sets by (1) analyzing each compound term in a technical corpus into a head + expansion (modifiers) and then (2) forming candidate synonyms of the compound by combining (a) synonyms of the original head with the original expansion, (b) the orig­inal head with synonyms of the original expansion, and (c) synonyms of the original head with synonyms of the original expansion. The resulting synonym sets are to be used in a document-consulting system to help users navigate complex technical docu­ments. Adeline Nazarenko, Pierre Zweigenbaum, Benoit Habert, and Jacques Bouaud contribute a paper that describes an approach to classifying unknown words in a medical corpus into one of the eleven top-level semantic categories in the SNOMED hierarchical terminology. They parse the corpus for NPs, extract dependency relations between the words in the parsed NPs, e.g., <i>W</i>1 <i>RW</i>2, and construct a graph wherein words are the nodes and edges are labeled with shared contexts between the con­nected words: <i>W</i>1 and <i>W</i>3 share a context if for some <i>R</i><i> </i>and <i>W</i>2, both <i>W</i>1 <i>RW</i>2 and <i>W</i>3 <i>RW</i>2 are attested in the corpus. In this similarity graph, words whose semantic category is known from the SNOMED resource are labeled with their category, and categories are then propagated to uncategorized nodes via a voting mechanism be­tween the uncategorized nodes' nearest neighbors. Finally, Michael Oakes and Chris Paice describe a technique for validating terms that can occur in particular semantic roles, or slots, in an information extraction-like template structure designed to cap­ture details of scientific papers for use in generating abstracts. Starting with an initial, corpus-derived thesaurus containing domain-specific high-frequency words and mul­tiword units (MWUs), each manually tagged with its semantic role, the MWUs are analyzed to reveal any that contain as substrings words or shorter MWUs already in the thesaurus. For such MWUs a semantic grammar rule is generated whose pattern is the MWU with the substring replaced by its semantic role and whose action is to label matching strings with the semantic role of the MWU. Such rules, which implic­itly define a class of semantically equivalent terms, generalize the thesaurus beyond observed examples and are used to validate proposed slot fillers in the template.</p><p>The third problem area, and the most challenging, is that of building knowledge-rich terminologies—terminologies that contain not only terms, but attributes and re­lationships of the concepts denoted by the terms, frequently for use in applications requiring controlled terminologies. James Cimino contributes a paper describing the methodology employed in maintaining a large-scale knowledge-based controlled med­ical terminology used to encode patient data and to provide aggregation classes for a variety of applications, such as billing and decision support. In such a critical and knowledge-rich environment, terms cannot be automatically added to the terminol­ogy as a consequence of language processing. However, Cimino describes how simple language processing, together with knowledge-based reasoning, can be used to guide a terminologist in the process of, for example, adding the name of a new drug. Anne Condamines and Josette Rebeyrolle describe a corpus-driven approach to construct­ing a terminological knowledge base. First they use Bourigault's LEXTER to identify candidate terms, then initiate a search for conceptual relationships among them.<page local="15" global="331"/> Tax­onomies are constructed by using a fixed set of linguistic patterns to identify can­didate hypernymic and meronymic binary relations. Then pairwise comparisons are made between terms in different taxonomies, and recurrent contexts in the corpus are sought in which these term pairs co-occur. If such contexts are found, a concep­tual relationship is proposed, linguistic patterns are created to match the context, and these patterns applied to the corpus to identify new terms. The process is then re­peated until no more relationships or terms are found. Although highly suggestive, this paper is unclear in critical places, particularly regarding which steps are carried out manually and which automatically. Finally, Ingrid Meyer presents a framework for building knowledge-rich terminological dictionaries. Her approach is firmly semiauto­matic, tools being provided to assist, rather than replace, a human terminologist. The method depends on acquiring knowledge patterns, which may be lexical, grammatical, or paralinguistic (relying on, e.g., punctuation), to find <b>knowledge-rich contexts </b>from which hypernymic or other attribute or relational knowledge may be extracted. Such patterns are acquired through an iterative manual process of refinement in conjunction with a corpus.</p><p>Not fitting neatly into the above classification are two papers on bilingual term alignment for machine translation (MT). This is an important application area for terminology systems, as the translation of terminology-laden technical documents is commercial MT's bread and butter and an area in which human translators' lack of domain-specific knowledge is likely to be a bottleneck. Eric Gaussier's paper gives a general overview of issues faced in bilingual terminology extraction from a parallel corpus, particularly choices between (1) extracting terms in each language indepen­dently, then aligning terms, or (2) parsing terms in one language, then projecting, by alignment, terms onto the second language, or (3) parallel parsing. He explores an idea, referred to as <b>pattern affinities, </b>that candidate terms expressed via one syntac­tic pattern in one language are more likely to be rendered in the other language by some other specific syntactic pattern but shows via an implementation of this idea using the EM algorithm that results are not significantly improved. David Hull, in a very clear and convincing piece, describes a method of bilingual lexicon construction from translated sentence pairs that relies on term extraction in the source language and a probabilistic word translation model to propose term translations in the target language. Although not perfect, this model can, the author argues, lead to significant productivity gain in constructing bilingual term lexica when used in a semiautomated mode by a human terminologist.</p><p>This is a wide-ranging collection, and the editors are to be congratulated for pulling together so much interesting material. That said, there a few reproaches to be leveled at them too. First, the level of copyediting is very poor. Spelling mistakes and minor grammatical errors abound; figures and tables have incorrect captions; formulas have undefined terms. At best this is irritating; at worst it seriously impedes understanding and conveys an impression of sloppiness that undermines the reader's trust. Perhaps this falls in the crack between what the editors and the publishers feel is their respon­sibility. But one expects somewhat better. Second, the book would have been much more readable and more generally useful had the papers been structured into related subareas, with introductory overviews in each area setting the stage for, and compar­ing and contrasting, the relevant papers. As it is, the editors have opted to present the papers in alphabetical order by first author's surname, hardly the most cognitively compelling of structural principles. The editors' introduction is a step in the right di­rection, but a small one, and given the wide range of topics addressed, some further analysis and guidance would have been welcome. As a consequence, although this is a book I would regret not having in my university library, it is not one I would regret not owning myself.<page local="16" global="332"/></p><p><i>Robert Gaizauskas </i>is Professor of Computer Science at the University of Sheffield. His research interests are in applied natural language processing, specifically information extraction, most recently concentrating on biomedical texts. Gaizauskas's address is Department of Computer Science, University of Sheffield, Regent Court, 211 Portobello St., Sheffield, S1 4DP, U.K.; e-mail: R.Gaizauskas@dcs.shef.ac.uk.</p></references></body></article>