<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="21"/><title>English Tasks: All-Words and Verb Lexical Sample</title><author surname="Palmer" givenname="Martha"><org  name="University of Pennsylvania" country="USA" city="Philadelphia"/></author><author surname="Fellbaum" givenname="Christiane"><org  name="University of Pennsylvania" country="USA" city="Philadelphia"/></author><author surname="Cotton" givenname="Scott"><org  name="University of Pennsylvania" country="USA" city="Philadelphia"/></author><author surname="Delfs" givenname="Lauren"><org  name="University of Pennsylvania" country="USA" city="Philadelphia"/></author><author surname="Dang" givenname="Hoa Trang"><org  name="University of Pennsylvania" country="USA" city="Philadelphia"/></author></firstpageheader><frontmatter><p>English Tasks: All-Words and Verb Lexical Sample</p><p><b>Martha Palmer, Christiane Fellbaum, Scott Cotton, Lauren Delfs, </b>and <b>Hoa Trang Dang</b></p><p>University of Pennsylvania {mpalmer,fellbaum,cotton,lcdelfs,htd}@linc.cis.upenn.edu</p></frontmatter><abstract><b>We describe our experience in preparing the lexicon and sense-tagged corpora used in the English all-words and lexical sample tasks of </b><b>Senseval</b><b>-2.</b> </abstract></header><body><section number="1" title="Overview"><p><b>The English lexical sample task is the result of a coordinated effort between the University of Pennsylvania, which provided training/test data for the verbs, and Adam Kilgarriff at Brighton, who provided the training/test data for the nouns and adjectives (see Kilgarriff, this issue). In addition, we provided the test data for the English all-words task. The pre-release version of WordNet 1.7 from Princeton was used as the sense inventory. Most of the revisions of sense definitions relevant to the English tasks were done prior to the bulk of the tagging.</b></p><p><b>The manual annotation for both the English all-words and verb lexical sample tasks was done by researchers and students in linguistics and computational linguistics at the University of Pennsylvania. All of the verbs in both the lex­ical sample and all-words tasks were annotated using a graphical tagging interface that allowed the annotators to tag instances by verb type and view the sentences surrounding the instances. Well over 1000 person hours went into the tag­ging tasks.</b></p></section><section number="2" title="English All-Words Task"><p><b>The test data for the English all-words task con­sisted of 5,000 words of running text from three Wall Street Journal articles representing varied domains from the Penn Treebank II. Annota­tors preparing the data were allowed to indi-</b></p><p><b>Christiane Fellbaum is at Princeton University, fell-</b>baum@clarity.princeton.edu</p><p><b>Table 1: System performance on English all-words task (fine-grained scores); (*) indicates system results that were submitted after the </b><b>senseval</b><b>-2 workshop and official deadline.</b></p><p><b>cate at most one multi-word construction for each content word to be tagged, but could give multiple senses for the construction. In some cases, a multi-word construction was annotated with senses associated with just the head word of the phrase in addition to more specific senses based on the entire phrase. The annotations were done under a double-blind scheme by two linguistics students, and were then adjudicated and corrected by a different person.</b></p><p><b>Task participants were supplied with test data only, in the standard all-words format for </b><b>senseval</b><b>-2, as well as the original syntactic</b> and part-of-speech annotations from the Tree-bank.<page local="2" global="22"/> Table 1 shows the system performance on the task. Most of the systems tagged al­most all the content words. This included not only indicating the appropriate sense from the WordNet 1.7 pre-release (as it stood at the time of annotation), but also marking multi-word constructions appropriate to the corresponding sense tags. If given a perfect lemmatizer, a sim­ple baseline strategy which does not attempt to find the satellite words in multi-word construc­tions, but which simply tags each head word with the first WordNet sense for the correspond­ing Treebank part-of-speech tag, would result in precision and recall of about 0.57.</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>System</b></p></td><td class="cell"><p><b>Precision</b></p></td><td class="cell"><p><b>Recall</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>SMUaw-</b></p></td><td class="cell"><p><b>0.690</b></p></td><td class="cell"><p><b>0.690</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>AVe-Antwerp</b></p></td><td class="cell"><p><b>0.636</b></p></td><td class="cell"><p><b>0.636</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>LIA-Sinequa-AHWords</b></p></td><td class="cell"><p><b>0.618</b></p></td><td class="cell"><p><b>0.618</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>david-fa-UNED-AW-T</b></p></td><td class="cell"><p><b>0.575</b></p></td><td class="cell"><p><b>0.569</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>david-fa-UNED-AW-U</b></p></td><td class="cell"><p><b>0.556</b></p></td><td class="cell"><p><b>0.550</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>gchao2-</b></p></td><td class="cell"><p><b>0.475</b></p></td><td class="cell"><p><b>0.454</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>gchao3-</b></p></td><td class="cell"><p><b>0.474</b></p></td><td class="cell"><p><b>0.453</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Ken-Litkowski-clr-aw (*)</b></p></td><td class="cell"><p><b>0.451</b></p></td><td class="cell"><p><b>0.451</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Ken-Litkowski-clr-aw</b></p></td><td class="cell"><p><b>0.416</b></p></td><td class="cell"><p><b>0.451</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>gchao</b></p></td><td class="cell"><p><b>0.500</b></p></td><td class="cell"><p><b>0.449</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>cm.guo-usm-english-tagger2</b></p></td><td class="cell"><p><b>0.360</b></p></td><td class="cell"><p><b>0.360</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>magnini2-irst-eng-all</b></p></td><td class="cell"><p><b>0.748</b></p></td><td class="cell"><p><b>0.357</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>cmguo-usm-english-tagger</b></p></td><td class="cell"><p><b>0.345</b></p></td><td class="cell"><p><b>0.338</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>c.guo-usm-english-tagger3</b></p></td><td class="cell"><p><b>0.336</b></p></td><td class="cell"><p><b>0.336</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>agirre2-ehu-dlist-all</b></p></td><td class="cell"><p><b>0.572</b></p></td><td class="cell"><p><b>0.291</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>judita-</b></p></td><td class="cell"><p><b>0.440</b></p></td><td class="cell"><p><b>0.200</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>dianam-system3ospdana</b></p></td><td class="cell"><p><b>0.545</b></p></td><td class="cell"><p><b>0.169</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>dianam-system2ospd</b></p></td><td class="cell"><p><b>0.566</b></p></td><td class="cell"><p><b>0.169</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>dianam-system 1</b></p></td><td class="cell"><p><b>0.598</b></p></td><td class="cell"><p><b>0.140</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>woody-IIT2</b></p></td><td class="cell"><p><b>0.328</b></p></td><td class="cell"><p><b>0.038</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>woody-IIT3</b></p></td><td class="cell"><p><b>0.294</b></p></td><td class="cell"><p><b>0.034</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>woody-IITl</b></p></td><td class="cell"><p><b>0.287</b></p></td><td class="cell"><p><b>0.033</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></section><section number="3" title="English Lexical Sample Task"><p>The data for the verb lexical sample task came primarily from the Penn Treebank II Wall Street Journal corpus. However, where that did not supply enough samples to approximate 75+15*n instances per verb, where n is the num­ber of senses for the verb, we supplemented with British National Corpus instances. We did not find sentences for every sense of every word we tagged. We also sometimes found sentences for which none of the available senses were appro­priate, and these were discarded. The instances for each verb were partitioned into training/test data using a ratio of 2:1.</p><p>We also grouped the nouns, adjectives and verbs for the lexical sample task, attempting to be explicit about the criteria for each grouping. In particular, the criteria for grouping verbs included differences in semantic classes of ar­guments, differences in the number and type of arguments, whether an argument refers to a created entity or a resultant state, whether an event involves concrete or abstract entities or constitutes a mental act, whether there is a specialized subject domain, etc. All of the verbs were grouped by two or more people, with differences being reconciled. In some cases the groupings of the verbs are identical to the ex­isting WordNet groupings; in some cases they are quite different. The nouns and adjectives were grouped by the primary annotator in the project; WordNet does not have comparable groups for nouns and adjectives.</p><p>These groupings were used for coarse-grained scoring, under the framework of S <b>ense val-</b>1.</p><p>After the <b>senseval</b>-2 workshop, participant* were invited to retrain their systems on th( groups; only a handful of participants chose tc do this, and in the end the results were uni­formly only slightly better than training on the fine-grained senses with coarse-grained scoring.</p><p>Table 2 shows the system performance on just the verbs of the lexical sampk task. For comparison we ran several sim­ple baseline algorithms that had been used in <b>Senseval</b>-1, including RANDOM, COMMON­EST, LESK, LESK-DEFINITION, and LESK-CORPUS (Kilgarriff and Rosenzweig, 2000). In contrast to <b>Senseval</b>-1, in which none of the competing systems performed significantly bet­ter than the highest baseline (LESK-CORPUS), the best-performing systems this time per­formed well above the highest baseline.</p><p>Overall, the performance of the systems was much lower than in <b>senseval</b>-1. Several fac­tors may have contributed to this. In addi­tion to the use of fine-grained WordNet senses instead of the smaller Hector sense inventory from <b>senseval</b>-1, most of the verbs included in this task were chosen specifically because we expected them to be difficult to tag. There was also generally less training data made available to the systems (ignoring outliers, there were on average twice as many training samples for each verb in <b>Senseval</b>-1 as there were in <b>Senseval-</b>2). Table 3 shows the correspondence between test data size (half of training data size), en­tropy, and system performance for each verb.</p></section><section number="4" title="Annotating the Gold Standard"><p>The annotators made every effort to match the target word to a WordNet sense both syntacti­cally and semantically, but sometimes this could not be done. Given a conflict between syntax and semantics, the annotators opted to match semantics. For example, the word "train" has an intransitive sense ("undergo training or in­struction in preparation for a particular role, function, or profession") as well as a related (causative) transitive sense ( "create by training and teaching"). Instances of "train" that were interpreted as having a dropped object were tagged with the transitive sense even though the overt syntax did not match the sense definition.</p><p>Some sentences seemed to fit equally well with two different senses, often because of am-<page local="3" global="23"/></p><p>Table <b>2: </b>System precision (P) and recall (R) for English verb lexical sample task (fine-grained scores); (*) indicates system results that were submitted after the <b>Senseval-2 </b>workshop and official deadline.</p><p>biguous context; others did not fit well under any sense. One of the solutions employed in these cases was the assignment of multiple sense tags. The taggers would choose two senses (on rare occasions, even three) that they felt made an approximation of the correct sense when used in combination. Sometimes this strategy was also used in arbitration, when it was decided that neither tagger's tag was better than the other. The taggers tried to use this strategy sparingly and chose single tags whenever possi­ble.</p><p>Often, a particular verb yielded multiple in-</p><p>Table <b>3: </b>Test corpus size, entropy (base <b>2) </b>of tagged data, and average system recall for each verb, using fine-grained and coarse-grained scor­ing.</p><p>stances of what was clearly a salient sense, but one not found in WordNet. One of the results was that sentences that should have received a clear sense tag ended up with something rather ad hoc, and often inconsistent. One of the most notorious examples was "call," which had no sense that fit sentences like "The restau­rant is called Marrakesh." WordNet contains some senses related to this one. One sense refers to the bestowing of a name; another to informal designations; another to greetings and vocatives. But there is no sense in WordNet for simply stating something's name without additional connotations, and the gap possibly caused some inconsistencies in the annotation. All these senses belonged to the same group, and if the annotators had been allowed to tag with the more general group sense, there may have been less inconsistency.<page local="4" global="24"/></p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>System</b></p></td><td class="cell"><p><b>P</b></p></td><td class="cell"><p><b>R</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>agirre3-ehu-dlist-best</b></p></td><td class="cell"><p><b>0.846</b></p></td><td class="cell"><p><b>0.229</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>magnini-irst-eng-sample</b></p></td><td class="cell"><p><b>0.660</b></p></td><td class="cell"><p><b>0.138</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>kunlp-</b></p></td><td class="cell"><p><b>0.576</b></p></td><td class="cell"><p><b>0.576</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>jhu-english-JHU-final (*)</b></p></td><td class="cell"><p><b>0.566</b></p></td><td class="cell"><p><b>0.566</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>SMUls-</b></p></td><td class="cell"><p><b>0.563</b></p></td><td class="cell"><p><b>0.563</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>LIA-Sinequa-Lexsample</b></p></td><td class="cell"><p><b>0.535</b></p></td><td class="cell"><p><b>0.535</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>manning-cs224n</b></p></td><td class="cell"><p><b>0.523</b></p></td><td class="cell"><p><b>0.523</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>agirre3-ehu-dlist-all</b></p></td><td class="cell"><p><b>0.514</b></p></td><td class="cell"><p><b>0.493</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>talp-TALP</b></p></td><td class="cell"><p><b>0.513</b></p></td><td class="cell"><p><b>0.513</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>umcp-englishl-</b></p></td><td class="cell"><p><b>0.494</b></p></td><td class="cell"><p><b>0.493</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>jhu-english-JHU-ENGLISH</b></p></td><td class="cell"><p><b>0.489</b></p></td><td class="cell"><p><b>0.489</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>montoyo-Univ.-Alicante-System</b></p></td><td class="cell"><p><b>0.486</b></p></td><td class="cell"><p><b>0.480</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>jhu-english-JHU</b></p></td><td class="cell"><p><b>0.485</b></p></td><td class="cell"><p><b>0.485</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tdpl-duluth3</b></p></td><td class="cell"><p><b>0.465</b></p></td><td class="cell"><p><b>0.465</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tdpla-duluthC</b></p></td><td class="cell"><p><b>0.453</b></p></td><td class="cell"><p><b>0.453</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tdpl-duluth5</b></p></td><td class="cell"><p><b>0.450</b></p></td><td class="cell"><p><b>0.450</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tdpl-duluth4</b></p></td><td class="cell"><p><b>0.446</b></p></td><td class="cell"><p><b>0.446</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>baseline-lesk-corpus</b></p></td><td class="cell"><p><b>0.445</b></p></td><td class="cell"><p><b>0.445</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tdpl-duluth2</b></p></td><td class="cell"><p><b>0.440</b></p></td><td class="cell"><p><b>0.440</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tdpla-duluthA</b></p></td><td class="cell"><p><b>0.439</b></p></td><td class="cell"><p><b>0.439</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tdpl-duluthl</b></p></td><td class="cell"><p><b>0.437</b></p></td><td class="cell"><p><b>0.437</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tdpla-duluthB</b></p></td><td class="cell"><p><b>0.404</b></p></td><td class="cell"><p><b>0.404</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>baseline-commonest</b></p></td><td class="cell"><p><b>0.403</b></p></td><td class="cell"><p><b>0.403</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>david-fal-UNED-LS-T</b></p></td><td class="cell"><p><b>0.388</b></p></td><td class="cell"><p><b>0.387</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>david-fal-UNED-LS-U</b></p></td><td class="cell"><p><b>0.288</b></p></td><td class="cell"><p><b>0.287</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Haynes-IIT2</b></p></td><td class="cell"><p><b>0.233</b></p></td><td class="cell"><p><b>0.232</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Haynes-IITl</b></p></td><td class="cell"><p><b>0.220</b></p></td><td class="cell"><p><b>0.220</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Kenneth-Litkowski-clr-ls</b></p></td><td class="cell"><p><b>0.218</b></p></td><td class="cell"><p><b>0.218</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Haynes-IIT2 (*)</b></p></td><td class="cell"><p><b>0.199</b></p></td><td class="cell"><p><b>0.192</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Haynes-IITl (*)</b></p></td><td class="cell"><p><b>0.193</b></p></td><td class="cell"><p><b>0.186</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>baseline-lesk</b></p></td><td class="cell"><p><b>0.181</b></p></td><td class="cell"><p><b>0.181</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>michael-oakes.suss2</b></p></td><td class="cell"><p><b>0.094</b></p></td><td class="cell"><p><b>0.094</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>baseline-lesk-def</b></p></td><td class="cell"><p><b>0.088</b></p></td><td class="cell"><p><b>0.088</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>baseline-random</b></p></td><td class="cell"><p><b>0.085</b></p></td><td class="cell"><p><b>0.085</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Verb</b></p></td><td class="cell"><p><b>Size</b></p></td><td class="cell"><p><b>Entropy</b></p></td><td class="cell"><p><b>Fine</b></p></td><td class="cell"><p><b>Coarse</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>ferret</b></p></td><td class="cell"><p><b>1</b></p></td><td class="cell"><p><b>0.00</b></p></td><td class="cell"><p><b>0.913</b></p></td><td class="cell"><p><b>0.913</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>collaborate</b></p></td><td class="cell"><p><b>30</b></p></td><td class="cell"><p><b>0.44</b></p></td><td class="cell"><p><b>0.898</b></p></td><td class="cell"><p><b>0.898</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>wander</b></p></td><td class="cell"><p><b>50</b></p></td><td class="cell"><p><b>0.96</b></p></td><td class="cell"><p><b>0.619</b></p></td><td class="cell"><p><b>0.786</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>face</b></p></td><td class="cell"><p><b>93</b></p></td><td class="cell"><p><b>1.09</b></p></td><td class="cell"><p><b>0.690</b></p></td><td class="cell"><p><b>0.785</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>replace</b></p></td><td class="cell"><p><b>45</b></p></td><td class="cell"><p><b>1.62</b></p></td><td class="cell"><p><b>0.471</b></p></td><td class="cell"><p><b>0.860</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>use</b></p></td><td class="cell"><p><b>76</b></p></td><td class="cell"><p><b>1.68</b></p></td><td class="cell"><p><b>0.558</b></p></td><td class="cell"><p><b>0.682</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>begin</b></p></td><td class="cell"><p><b>280</b></p></td><td class="cell"><p><b>1.76</b></p></td><td class="cell"><p><b>0.625</b></p></td><td class="cell"><p><b>0.625</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>treat</b></p></td><td class="cell"><p><b>44</b></p></td><td class="cell"><p><b>2.10</b></p></td><td class="cell"><p><b>0.453</b></p></td><td class="cell"><p><b>0.543</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>live</b></p></td><td class="cell"><p><b>67</b></p></td><td class="cell"><p><b>2.35</b></p></td><td class="cell"><p><b>0.455</b></p></td><td class="cell"><p><b>0.476</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>match</b></p></td><td class="cell"><p><b>42</b></p></td><td class="cell"><p><b>2.35</b></p></td><td class="cell"><p><b>0.398</b></p></td><td class="cell"><p><b>0.620</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>train</b></p></td><td class="cell"><p><b>63</b></p></td><td class="cell"><p><b>2.60</b></p></td><td class="cell"><p><b>0.394</b></p></td><td class="cell"><p><b>0.492</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>drift</b></p></td><td class="cell"><p><b>32</b></p></td><td class="cell"><p><b>2.77</b></p></td><td class="cell"><p><b>0.327</b></p></td><td class="cell"><p><b>0.354</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>dress</b></p></td><td class="cell"><p><b>59</b></p></td><td class="cell"><p><b>2.89</b></p></td><td class="cell"><p><b>0.434</b></p></td><td class="cell"><p><b>0.679</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>serve</b></p></td><td class="cell"><p><b>51</b></p></td><td class="cell"><p><b>3.02</b></p></td><td class="cell"><p><b>0.404</b></p></td><td class="cell"><p><b>0.445</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>drive</b></p></td><td class="cell"><p><b>42</b></p></td><td class="cell"><p><b>3.03</b></p></td><td class="cell"><p><b>0.308</b></p></td><td class="cell"><p><b>0.528</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>leave</b></p></td><td class="cell"><p><b>66</b></p></td><td class="cell"><p><b>3.06</b></p></td><td class="cell"><p><b>0.317</b></p></td><td class="cell"><p><b>0.428</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>develop</b></p></td><td class="cell"><p><b>69</b></p></td><td class="cell"><p><b>3.17</b></p></td><td class="cell"><p><b>0.301</b></p></td><td class="cell"><p><b>0.456</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>see</b></p></td><td class="cell"><p><b>69</b></p></td><td class="cell"><p><b>3.28</b></p></td><td class="cell"><p><b>0.278</b></p></td><td class="cell"><p><b>0.317</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>wash</b></p></td><td class="cell"><p><b>12</b></p></td><td class="cell"><p><b>3.31</b></p></td><td class="cell"><p><b>0.343</b></p></td><td class="cell"><p><b>0.535</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>work</b></p></td><td class="cell"><p><b>60</b></p></td><td class="cell"><p><b>3.54</b></p></td><td class="cell"><p><b>0.303</b></p></td><td class="cell"><p><b>0.442</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>keep</b></p></td><td class="cell"><p><b>67</b></p></td><td class="cell"><p><b>3.62</b></p></td><td class="cell"><p><b>0.336</b></p></td><td class="cell"><p><b>0.353</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>call</b></p></td><td class="cell"><p><b>66</b></p></td><td class="cell"><p><b>3.68</b></p></td><td class="cell"><p><b>0.246</b></p></td><td class="cell"><p><b>0.457</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>play</b></p></td><td class="cell"><p><b>66</b></p></td><td class="cell"><p><b>3.80</b></p></td><td class="cell"><p><b>0.323</b></p></td><td class="cell"><p><b>0.345</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>find</b></p></td><td class="cell"><p><b>68</b></p></td><td class="cell"><p><b>3.81</b></p></td><td class="cell"><p><b>0.178</b></p></td><td class="cell"><p><b>0.285</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>carry</b></p></td><td class="cell"><p><b>66</b></p></td><td class="cell"><p><b>3.97</b></p></td><td class="cell"><p><b>0.279</b></p></td><td class="cell"><p><b>0.332</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>strike</b></p></td><td class="cell"><p><b>54</b></p></td><td class="cell"><p><b>4.06</b></p></td><td class="cell"><p><b>0.248</b></p></td><td class="cell"><p><b>0.331</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>pull</b></p></td><td class="cell"><p><b>60</b></p></td><td class="cell"><p><b>4.24</b></p></td><td class="cell"><p><b>0.255</b></p></td><td class="cell"><p><b>0.414</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>draw</b></p></td><td class="cell"><p><b>41</b></p></td><td class="cell"><p><b>4.60</b></p></td><td class="cell"><p><b>0.195</b></p></td><td class="cell"><p><b>0.264</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>turn</b></p></td><td class="cell"><p><b>67</b></p></td><td class="cell"><p><b>4.79</b></p></td><td class="cell"><p><b>0.216</b></p></td><td class="cell"><p><b>0.327</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>It has been well-established that sense-tagging is a very difficult task (Kilgarriff, 1997; Hanks, 2000), even for experienced human tag­gers. If the sense inventory has gaps or redun­dancies, or if some of the sense glosses have ambiguous wordings, choosing the correct sense can be all but impossible. Even if the annotator is working with a very good entry, unforeseen instances of the word always arise.</p><p>The degree of polysemy does not affect the relative difficulty of tagging, at least not in the way it is often thought. Very polysemous words, such as "drive," are not necessarily harder to tag than less polysemous words like "replace." The difficulty of tagging depends much more on other aspects of the entry and of the word itself. Often very polysemous words <i>are </i>quite difficult to tag, because they are more likely to be un-derspecified or occur in novel uses; however, "re­place," with four senses, proved a difficult verb to tag, while "play," with thirty-five senses, was relatively straightforward.</p><p>In many ways, the grouped senses are very helpful for the sense-tagger. Grouping similar senses allows the sense-tagger to study side-by-side the senses that are perhaps most likely to be confused, which is helpful when the differences between the senses are very subtle. However, it would be a poor idea to attempt to tag a corpus using <i>only </i>the groups, and not the finer sense distinctions, because often some of the senses included in a group will have some properties that the others do not; it is always better to make the finest distinction possible and not just assign the same tag to everything that seems close.</p><p>Inter-annotator agreement figures for the hu­man taggers are quite low. However, in some respects they are not quite as low as they seem. Some of the apparent discrepancies were sim­ply the result of a technical error: the annota­tor accidentally picked the wrong tag, perhaps choosing one of its neighbors. Other differences resulted from the sense inventories themselves. Sometimes the taggers interpreted the wording of a given sense definition in different ways, which caused them to choose different tags, but does not entail that they had interpreted the instances differently; in fact, discussion of such cases usually revealed that the taggers had interpreted the instances themselves in the sam way. Additional apparent discrepancies resulte&lt; from the various strategies for dealing with case in which there was no single proper sense h WordNet. This was the case when an instance in the corpus was underspecified so as to al low multiple appropriate interpretations. Thii resulted in (a) multiple tags by one or botl taggers, and (b) each tagger making a differ ent choice. Here, again, the taggers often hac the same interpretation of the instance itself but because the sense inventory was insufficienl for their needs, they were forced to find différent strategies. Sometimes, in fact, one tagger would double-tag a particular instance while the sec­ond tagger chose a single sense that matched one of the two selected by the first annotator. This is considered a discrepancy for statistical purposes, but clearly reflects similar interpreta­tions on the part of the annotators.</p><p>In the most recent evaluation, with two new annotators tagging against the Gold Standard, the best fine-grained agreement figures for verbs were in the 70's, similar to Semcor figures. How­ever, when we used the groupings to do a more coarse-grained evaluation, and counted a match between a single tag and a member of a double tag as correct, the human annotator agreement figures rose to 90%.</p></section><section number="5" title="Acknowledgments"><p>Support for this work was provided by the National Science Foundation (grants NSF-9800658 and NSF-9910603), DARPA (grant 535626), and the CIA (contract number 2000*SO53100*000). We would also like to thank Joseph Rosenzweig for building the anno­tation tools, and Susanne Wolff for contribution to the manual annotation.</p></section><references><p>Patrick Hanks. 2000. Do word meanings ex­ist? <i>Computers and the Humanities, </i>34(1-2), April. Special Issue on SENSEVAL.</p><p>A. Kilgarriff and J. Rosenzweig. 2000. Frame­work and results for English SENSEVAL. <i>Computers and the Humanities, </i>34(1-2), April. Special Issue on SENSEVAL.</p><p>Adam Kilgarriff. 1997. I don't believe in word senses. <i>Computers and the Humanities, </i>31(2).</p></references></body></article>