<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="17"/><title>English Lexical Sample Task Description</title><author surname="Kilgarriff" givenname="Adam"><org  name="University of Brighton" country="United Kingdom" city="Brighton"/></author></firstpageheader><frontmatter><p>English Lexical Sample Task Description</p><p><b>Adam Kilgarriff</b></p><p><b>ITRJ, University of Brighton Brighton, UK adamOitri.bton</b><b>.</b>ac.uk</p><p>The English lexical sample task (adjectives and nouns) for SENSEVAL 2 was set up accord­ing to the same principles as for SENSEVAL-1, as reported in (Kilgarriff and Rosenzweig, 2000). (Adjectives and nouns only, because the data preparation for the verbs lexical sample was undertaken alongside that for the English all-words task, and is reported in Palmer et al (this volume). All discussion below up to the Results section covers only adjectives and nouns.)</p><p><b>1 Lexical sample</b></p></frontmatter><abstract>The lexicon was sampled to give a range of low, medium and high frequency words (see Table 1). These were all different words to the ones used in SENSEVAL 1. </abstract></header><body><section number="2" title="Corpus choice"><p>For the most part, the British National Cor­pus (New edition) was used. (The new edition has the advantage that it is available world­wide, so all participants had the opportunity of obtaining it for system training.) Our goal was to match this source, containing British En­glish, with another, of American English. In the event, only limited quantities of corpus data for American English were available without copy­right complications, so the lion's share of the data was from the BNC with a limited quantity from the Wall Street Journal.</p><p>In accordance with standard SENSEVAL pro­cedure, the goal was to have 75 + 15n + 6m in­stances for each lexical-sample word, where <i>n </i>is the number of senses the word has and <i>m </i>is the number of multiword expressions that the word is part of (both, of course, relative to a specific lexicon). In practice numbers varied slightly, as instances were deleted because they had the wrong part of speech or were otherwise unusable. See Table 1 for actual numbers of senses, multiwords expressions and instances.</p></section><section number="3" title="Lexicon choice"><p>Here lay the biggest contrast with the SENSEVAL-1 task, which had used Oxford University Press's experimental HECTOR lexi­con. This time, in response to popular acclaim, WordNet was used.</p><p>Since SENSEVAL was first mooted, in 1997, WordNet-or-not-WordNet has been a recurring theme. In favour was the argument that it was already very widely used, almost a <i>de facto </i>stan­dard. The argument against concerned its sense distinctions. WordNet, like thesauruses but un­like standard dictionaries, is organised around groups of words of similar meanings <i>(synsets), </i>not around words (with their various meanings). This means that the priority for the lexicogra­pher is building coherent synsets rather than the coherent analysis of the various meanings of a particular word. The writer of a thesaurus does not need to pay as much attention to the distinc­tion between two senses of a word, as the writer of a dictionary. Word sense disambiguation is a task which needs clear and well-motivated sense distinctions. In English SENSEVAL-1, Word-Net was not used because of concerns that it did not provide clean enough sense distinctions.</p><p>While HECTOR provided good sense distinc­tions, it was unsatisfactory in that it did not cover the whole lexicon so there was no pos­sibility of scaling up. The case for WordNet - that it was already integrated into so much NLP and WSD work - still stood, so the de­cision was made to use WordNet. To guard against cases where WordNet made a distinc­tion between two meanings, but it was not clear what the distinction was, all the words in the lexical sample had their entries reviewed by a<page local="2" global="18"/></p><doubt alpha="0.0" length="1" tooSmall="False" monospace="0.0">8</doubt><p>Table 1: Lexical sample: rubric for column headers: Ss=number of fine-grained senses; Mwe = number of multi-word expressions which the word participates in (as <i>bear </i>participates in WordNet headword <i>polar bear); </i>inst = number of instances tagged; ITA = inter-tagger agree­ment (fine-grained).</p><p>lexicographer, with a view particularly to merg­ing insufficiently-distinct senses. It was initially unclear how these revisions would relate to the publicly available version of WordNet (at that time, WordNet 1.6). We are very grateful to the Princeton WordNet team (George Miller, Chris­tiane Fellbaum and Randee Tengi) for their help at this point; they agreed to incorporate our proposed revisions into a new version of Word-Net (1.7) which was then made available in time (despite some very tight deadlines) for the SEN­SEVAL competition.</p><p>WordNet 1.7 was not available as a complete object at the time of the gold standard pro­duction, in Spring 2001, but the entries for the lexical sample words were fixed at that point. For each lexical sample entry, we produced an HTML version for the lexicographers to work from. In addition to all the relevant infor­mation in WordNet, this had a mnemonic for each sense, so that taggers could use mnemon­ics when doing the tagging, rather than easily-forgotten, easily-confused sense numbers. The mnemonics were selected by a lexicographer.</p></section><section number="4" title="Gold standard production"><p>Once the corpus sources and lexical entries were fixed, work could proceed with the Gold-Standard tagging.<footnote anchor="1"/></p><p>First, a team of three professional lexicogra­phers and fourteen students and others was re­cruited. Recruitment proceeded as follows: an aptitude test was set up on the web. The test involved sense-tagging some corpus instances (taken from SENSEVAL-1, so the gold-standard answers were known). Email postings were made asking interested people to visit the web­site and take the test. All applicants scoring sufficiently well on the test were then offered work, on a piecework basis.</p><p>An HTML version of the corpus for a word was prepared. This comprised a series of ten-sentence stretches of text, with one word in the last of the sentences highlighted; that was the word to be sense-tagged. The files were HTML versions of the XML files used for test and train­ing data.</p><p>A tagger was emailed the lexical entry and corpus for a word.  They then tagged it, and</p><p>lrThe tagging was supported by a grant from EPSRC, the UK funding council, under GR/R02337/01 (MATS).</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Word</p></td><td class="cell"><p>Ss</p></td><td class="cell"><p>Mwe</p></td><td class="cell"><p>inst</p></td><td class="cell"><p>ITA</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ADJS: lexical sample size: 15</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>blind</p></td><td class="cell"><p>3</p></td><td class="cell"><p>21</p></td><td class="cell"><p>163</p></td><td class="cell"><p>89.6</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>colorless</p></td><td class="cell"><p>2</p></td><td class="cell"><p>0</p></td><td class="cell"><p>103</p></td><td class="cell"><p>94.2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>cool</p></td><td class="cell"><p>6</p></td><td class="cell"><p>1</p></td><td class="cell"><p>158</p></td><td class="cell"><p>92.1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>faithful</p></td><td class="cell"><p>3</p></td><td class="cell"><p>0</p></td><td class="cell"><p>70</p></td><td class="cell"><p>94.6</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>fine</p></td><td class="cell"><p>9</p></td><td class="cell"><p>6</p></td><td class="cell"><p>212</p></td><td class="cell"><p>84.0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>fit</p></td><td class="cell"><p>3</p></td><td class="cell"><p>0</p></td><td class="cell"><p>86</p></td><td class="cell"><p>85.0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>free</p></td><td class="cell"><p>8</p></td><td class="cell"><p>36</p></td><td class="cell"><p>247</p></td><td class="cell"><p>79.2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>graceful</p></td><td class="cell"><p>2</p></td><td class="cell"><p>0</p></td><td class="cell"><p>85</p></td><td class="cell"><p>72.6</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>green</p></td><td class="cell"><p>7</p></td><td class="cell"><p>80</p></td><td class="cell"><p>284</p></td><td class="cell"><p>86.6</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>local</p></td><td class="cell"><p>3</p></td><td class="cell"><p>12</p></td><td class="cell"><p>113</p></td><td class="cell"><p>89.1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>natural</p></td><td class="cell"><p>10</p></td><td class="cell"><p>37</p></td><td class="cell"><p>309</p></td><td class="cell"><p>72.4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>oblique</p></td><td class="cell"><p>2</p></td><td class="cell"><p>5</p></td><td class="cell"><p>86</p></td><td class="cell"><p>96.4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>simple</p></td><td class="cell"><p>7</p></td><td class="cell"><p>19</p></td><td class="cell"><p>196</p></td><td class="cell"><p>67.8</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>solemn</p></td><td class="cell"><p>2</p></td><td class="cell"><p>0</p></td><td class="cell"><p>77</p></td><td class="cell"><p>84.1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>vital</p></td><td class="cell"><p>4</p></td><td class="cell"><p>7</p></td><td class="cell"><p>112</p></td><td class="cell"><p>93.7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ALL ADJS</p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>2301</p></td><td class="cell"><p>83.4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>NOUNS: lexical sample size: 29</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>art</p></td><td class="cell"><p>5</p></td><td class="cell"><p>35</p></td><td class="cell"><p>294</p></td><td class="cell"><p>78.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>authority</p></td><td class="cell"><p>7</p></td><td class="cell"><p>6</p></td><td class="cell"><p>276</p></td><td class="cell"><p>84.3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>bar</p></td><td class="cell"><p>13</p></td><td class="cell"><p>57</p></td><td class="cell"><p>455</p></td><td class="cell"><p>87.3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>bum</p></td><td class="cell"><p>4</p></td><td class="cell"><p>0</p></td><td class="cell"><p>137</p></td><td class="cell"><p>91.7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>chair</p></td><td class="cell"><p>4</p></td><td class="cell"><p>35</p></td><td class="cell"><p>207</p></td><td class="cell"><p>92.8</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>channel</p></td><td class="cell"><p>7</p></td><td class="cell"><p>10</p></td><td class="cell"><p>218</p></td><td class="cell"><p>84.8</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>child</p></td><td class="cell"><p>4</p></td><td class="cell"><p>16</p></td><td class="cell"><p>193</p></td><td class="cell"><p>92.3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>church</p></td><td class="cell"><p>3</p></td><td class="cell"><p>21</p></td><td class="cell"><p>192</p></td><td class="cell"><p>88.0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>circuit</p></td><td class="cell"><p>6</p></td><td class="cell"><p>31</p></td><td class="cell"><p>255</p></td><td class="cell"><p>93.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>day</p></td><td class="cell"><p>9</p></td><td class="cell"><p>82</p></td><td class="cell"><p>434</p></td><td class="cell"><p>76.3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>detention</p></td><td class="cell"><p>2</p></td><td class="cell"><p>5</p></td><td class="cell"><p>95</p></td><td class="cell"><p>98.7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>dyke</p></td><td class="cell"><p>2</p></td><td class="cell"><p>0</p></td><td class="cell"><p>86</p></td><td class="cell"><p>96.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>facility</p></td><td class="cell"><p>5</p></td><td class="cell"><p>9</p></td><td class="cell"><p>172</p></td><td class="cell"><p>89.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>fatigue</p></td><td class="cell"><p>4</p></td><td class="cell"><p>6</p></td><td class="cell"><p>128</p></td><td class="cell"><p>97.7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>feeling</p></td><td class="cell"><p>6</p></td><td class="cell"><p>5</p></td><td class="cell"><p>153</p></td><td class="cell"><p>77.0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>grip</p></td><td class="cell"><p>7</p></td><td class="cell"><p>3</p></td><td class="cell"><p>153</p></td><td class="cell"><p>85.2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>hearth</p></td><td class="cell"><p>3</p></td><td class="cell"><p>1</p></td><td class="cell"><p>96</p></td><td class="cell"><p>85.0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>holiday</p></td><td class="cell"><p>2</p></td><td class="cell"><p>9</p></td><td class="cell"><p>93</p></td><td class="cell"><p>90.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>lady</p></td><td class="cell"><p>3</p></td><td class="cell"><p>27</p></td><td class="cell"><p>158</p></td><td class="cell"><p>74.1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>material</p></td><td class="cell"><p>5</p></td><td class="cell"><p>39</p></td><td class="cell"><p>209</p></td><td class="cell"><p>85.1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>mouth</p></td><td class="cell"><p>8</p></td><td class="cell"><p>10</p></td><td class="cell"><p>179</p></td><td class="cell"><p>88.7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>nation</p></td><td class="cell"><p>3</p></td><td class="cell"><p>10</p></td><td class="cell"><p>112</p></td><td class="cell"><p>90.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>nature</p></td><td class="cell"><p>5</p></td><td class="cell"><p>8</p></td><td class="cell"><p>138</p></td><td class="cell"><p>86.7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>post</p></td><td class="cell"><p>8</p></td><td class="cell"><p>33</p></td><td class="cell"><p>236</p></td><td class="cell"><p>87.7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>restraint</p></td><td class="cell"><p>6</p></td><td class="cell"><p>3</p></td><td class="cell"><p>136</p></td><td class="cell"><p>80.4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>sense</p></td><td class="cell"><p>5</p></td><td class="cell"><p>37</p></td><td class="cell"><p>160</p></td><td class="cell"><p>87.1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>spade</p></td><td class="cell"><p>3</p></td><td class="cell"><p>7</p></td><td class="cell"><p>98</p></td><td class="cell"><p>95.1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>stress</p></td><td class="cell"><p>5</p></td><td class="cell"><p>7</p></td><td class="cell"><p>118</p></td><td class="cell"><p>74.7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>yew</p></td><td class="cell"><p>2</p></td><td class="cell"><p>15</p></td><td class="cell"><p>85</p></td><td class="cell"><p>97.1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ALL NOUNS</p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>5266</p></td><td class="cell"><p>86.3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ALL</p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>7567</p></td><td class="cell"><p>85.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="3" global="19"/><p>returned, by email, a file of answers. These files were checked automatically, and if they contained 'answers' which were not possible an­swers for the word, the suspect items were au­tomatically emailed back to the tagger for cor­rection,</p><p>The tagger guidelines are available along with other resources for the English-lexical-sample task. They developed in the course of the exer­cise; when a tagger asked a pertinent questions, I circulated the question and my answer to all taggers and incorporated them into the guide­lines.</p><p>As in SENSEVAL-1, "Unassignable" and "Proper-name" tags were always available alongside regular tags, and taggers were told to put down more than one tag, where multiple tags were equally applicable. Taggers were also asked to mark items where the part of speech was wrong; these were then deleted from the dataset.</p></section><section number="5" title="Tagger agreement procedures and scores"><p>As in all exercises where a gold standard corpus is the goal, it was necessary to have all data tagged by more than one person. The question then arises, how many taggings does each item need? The algorithm adopted here was:</p></section><section number="1." title="send item out to two taggers"></section><section number="2." title="if they agree completely, stop; return agreed answer"></section><section number="3." title="else, send out to another tagger"></section><section number="4." title="is there one or more tag that two agree on?"></section><section number="5." title="if yes, stop; return all tags which two people agree on"></section><section number="6." title="if no, return to step 3"><p>Thus, in simple cases, a minimum of effort was used, but in difficult cases, more opinions were obtained. The number of taggings per items is shown below. Note that the algorithm stops at step 2 if both taggers agree on one tag, or if both taggers agree on two or more tags.</p><p>Table 2: Patterns of (dis)agreement for 3-tagger cases. GS = gold standard tagging arising from these human taggings. ";" used as separator where a tagger (or the gold standard) gave mul­tiple tags.</p><p>Of the 5032 two-tagger items, in 4688 cases, the taggers agreed on one tag; in 340 cases, on two tags; and in 4 cases, on three tags.</p><p>For the 2446 cases which were tagged three times, 136 were cases where all three taggers agreed perfectly (so, had the algorithm been followed to the letter, the item would not have been tagged a third time; such cases were caused by delays in taggers returning answers.) The common patterns amongst the remainder are shown in Table 2.</p><p>For the 86 cases with four taggers, half the cases were {A, A, B, C} taggings.</p><p>Fine-grained inter-tagger agreement (ITA) figures was calculated using the same scoring algorithm as for the systems.<footnote anchor="2"/> For each pair of taggers tagging an instance, two scores were calculated, one with the one answer as the key, the other with the other. For each instance, scores were normalised so that the maximum score for each corpus instance was one, however many times it had been tagged. The overall ITA was 85.5%. A breakdown by word and by word class is given in Table l.<footnote anchor="3"/><page local="4" global="20"/></p><footnote label="2">All ITA figures and other results reported in this pa­per refer to fine-grained sense distinctions. The grouping of senses into coarse-grained categories took place inde­pendently of the gold-standard preparation, which was based entirely on fine sense distinctions.</footnote><footnote label="3">Kappa was not calculated because there were vari­ous ways in which it might have been calculated, so it was unclear which was appropriate, and it would have introduced more complication than clarification. Also</footnote><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>3 taggers' answers</p></td><td class="cell"><p>GS</p></td><td class="cell"><p>cases</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A       A B</p></td><td class="cell"><p>A</p></td><td class="cell"><p>651</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A      A;B A</p></td><td class="cell"><p>A</p></td><td class="cell"><p>550</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A      A;B B</p></td><td class="cell"><p>A;B</p></td><td class="cell"><p>209,</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A      A;B A;B</p></td><td class="cell"><p>A;B</p></td><td class="cell"><p>189</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A      A;B C</p></td><td class="cell"><p>A</p></td><td class="cell"><p>162</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A       A A;B;C</p></td><td class="cell"><p>A</p></td><td class="cell"><p>67</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A      A;B A;C</p></td><td class="cell"><p>A</p></td><td class="cell"><p>51</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A       A B;C</p></td><td class="cell"><p>A</p></td><td class="cell"><p>44</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A;B     A;C C</p></td><td class="cell"><p>A;C</p></td><td class="cell"><p>41</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A;B   A:B;C C</p></td><td class="cell"><p>A;B;C</p></td><td class="cell"><p>38</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Taggings</p></td><td class="cell"><p>Number</p></td><td class="cell"><p>%</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>2</p></td><td class="cell"><p>5032</p></td><td class="cell"><p>66.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>3</p></td><td class="cell"><p>2446</p></td><td class="cell"><p>32.3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>4</p></td><td class="cell"><p>86</p></td><td class="cell"><p>1.1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>5</p></td><td class="cell"><p>4</p></td><td class="cell"><p>0.05</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>As argued in (Kilgarriff and Rosenzweig, 2000) (also (Kilgarriff, 1999)) the inter-tagger agreement figure for a gold standard is of less interest than the replicability figure: if a com­pletely different team of taggers used the same methodology to do the same task, what would the agreement level between the two teams' out­puts be? It is the replicability figure, rather than ITA, which defines an upper bound for the task. We have not yet had time to conduct such a study.</p></section><section number="6" title="Task organisation"><p>The organisation followed standard SENSEVAL procedure. The data was prepared in XML us­ing SENSEVAL DTDs, with the data for each word split in a ration of 2:1 between training and test data. Data distribution, results up­loads, baselines and scoring were handled at UPenn (see paper by Cotton and Edmonds).</p></section><section number="7" title="Results"><p>Results are presented in the table below. Owing to space constraints, where a team submitted multiple systems with similar results, only the best result is shown. Full results are available at the SENSEVAL website, as are decodings of system names. At the SENSEVAL workshop (5-6 July 2001) it was agreed that there should also be a later deadline (end July 2001) so that 'egregious bugs' could be fixed. In order to hon­our both standard practice in evaluation exer­cises (eg, no extension of deadlines) and also the agreement made at the workshop, both results sets are presented, with later-deadline results marked with (R) as a suffix to the name.</p><p>There has not yet been time for an analy­sis of the results. The one comment that does seem pertinent is the contrast with the English-lexical-sample task in SENSEVAL-1. The tasks were organised in similar ways, and some of the systems were improved versions of systems par­ticipating in 1998. Yet the performance of the best systems has, apparently, dropped around 14%. We may well ask, why?</p><p>We believe the drop is due to the choice of lexicon. As discussed above, using Word-Net for SENSEVAL has drawbacks. High-</p><p>the figures shown, unlike kappa figures, have the merit of being directly comparable with system performance scores.</p><p>Table 3: PR=system precision; ATT= percent­age of cases for which an answer was returned ("attempted").</p><p>accuracy word sense disambiguation is only pos­sible where the lexicon makes clear and well-motivated sense distinctions, and provides suf­ficient information about the distinctions for the disambiguation algorithm to build on. An im­plication for future WSD research is that it is time to turn our attention from algorithms, to sense distinctions.</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>PR</b></p></td><td class="cell"><p><b>ATT</b></p></td><td class="cell"><p><b>System</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Supervised systems</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.82</b></p></td><td class="cell"><p><b>28</b></p></td><td class="cell"><p><b>BCU ehu-dlist-best</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.67</b></p></td><td class="cell"><p><b>25</b></p></td><td class="cell"><p><b>IRST</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.64</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>JHU (R)</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.64</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>SMUls</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.63</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>KUNLP</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.62</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>Stanford-CS224</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.61</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>Sinequa-LIA SCT</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.59</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>TALP</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.57</b></p></td><td class="cell"><p><b>98</b></p></td><td class="cell"><p><b>BCU ehu-dlist-all</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.57</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>Duluth-3</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.57</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>UMD-SST</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.50</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>UNED LS-T</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.42</b></p></td><td class="cell"><p><b>98</b></p></td><td class="cell"><p><b>Alicante</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Supervised baselines</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.51</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>Base Lesk</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.48</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>Base Commonest</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Unsupervised systems</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.58</b></p></td><td class="cell"><p><b>55</b></p></td><td class="cell"><p><b>ITRI-WASPS</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>40</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>UNED-LS-U</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.29</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>CLresearch DIMAP</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.25</b></p></td><td class="cell"><p><b>99</b></p></td><td class="cell"><p><b>IIT-2 (R)</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Unsupervised baselines</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.16</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>Base Lesk-defs</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>.14</b></p></td><td class="cell"><p><b>100</b></p></td><td class="cell"><p><b>Base random</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></section><references><p>Adam Kilgarriff and Joseph Rosenzweig. 2000. Framework and results for English Senseval. <i>Computers and the Humanities, </i>34(1-2): 15-48. Special Issue on Senseval, edited by Adam Kil­garriff and Martha Palmer.</p><p>Adam Kilgarriff. 1999. 95% replicability for manual word sense tagging. In <i>Proc. EACL, </i>pages 277­278, Bergen, June.</p></references></body></article>