<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="63"/><title>SemEval-2010 Task 14: Word Sense Induction &amp;Disambiguation</title><pubinfo>Proceedings of the 5th International Workshop on Semantic Evaluation, ACL 2010,pages 63-68, Uppsala, Sweden, 15-16 July 2010. ©2010 Association for Computational Linguistics</pubinfo><author surname="Manandhar" givenname="Suresh"><org  name="University of York" country="United Kingdom" city="York"/></author><author surname="Klapaftis" givenname="Ioannis"><org  name="University of York" country="United Kingdom" city="York"/></author><author surname="Dligach" givenname="Dmitriy"><org  name="University of Oslo" country="Norway" city="Oslo"/></author><author surname="Pradhan" givenname="Sameer"><org  name="University of Oslo" country="Norway" city="Oslo"/></author></firstpageheader><frontmatter><p><b>SemEval-2010 Task 14: Word Sense Induction &amp; Disambiguation</b></p><p><b>Suresh Manandhar Ioannis </b><b>P.</b><b> Klapaftis</b></p><p>Department of Computer Science Department of Computer Science</p><p>University of York, UK University of York, UK</p><p><b>Dmitriy Dligach</b></p><p>Department of Computer Science University of Colorado, USA</p></frontmatter><abstract>This paper presents the description and evaluation framework of SemEval-2010 Word Sense Induction &amp; Disambiguation task, as well as the evaluation results of 26 participating systems. In this task, partici­pants were required to induce the senses of 100 target words using a training set, and then disambiguate unseen instances of the same words using the induced senses. Sys­tems' answers were evaluated in: (1) an unsupervised manner by using two clus­tering evaluation measures, and (2) a su­pervised manner in a WSD task. </abstract></header><body><section number="1" title="Introduction"><p>Word senses are more beneficial than simple word forms for a variety of tasks including Information Retrieval, Machine Translation and others (Pantel and Lin, 2002). However, word senses are usually represented as a fixed-list of definitions of a manu­ally constructed lexical database. Several deficien­cies are caused by this representation, e.g. lexical databases miss main domain-specific senses (Pan­tel and Lin, 2002), they often contain general defi­nitions and suffer from the lack of explicit seman­tic or contextual links between concepts (Agirre et al., 2001). More importantly, the definitions of hand-crafted lexical databases often do not reflect the exact meaning of a target word in a given con­text (Veronis, 2004).</p><p>Unsupervised Word Sense Induction (WSI) aims to overcome these limitations of hand-constructed lexicons by learning the senses of a target word directly from text without relying on any hand-crafted resources. The primary aim of SemEval-2010 WSI task is to allow comparison of unsupervised word sense induction and disam­biguation systems.</p><p><b>Sameer S. Pradhan</b></p><p>BBN Technologies Cambridge, USA</p><p>The target word dataset consists of 100 words, 50 nouns and 50 verbs. For each target word, par­ticipants were provided with a training set in or­der to learn the senses of that word. In the next step, participating systems were asked to disam­biguate unseen instances of the same words using their learned senses. The answers of the systems were then sent to organisers for evaluation.</p></section><section number="2" title="Task description"><p>Figure 1 provides an overview of the task. As can be observed, the task consisted of three separate phases. In the first phase, <i>train­ing phase, </i>participating systems were provided with a training dataset that consisted of a set of target word (noun/verb) instances (sen­tences/paragraphs). Participants were then asked to use this training dataset to induce the senses of the target word. No other resources were al­lowed with the exception of NLP components for morphology and syntax. In the second phase, <i>testing phase, </i>participating systems were pro­vided with a testing dataset that consisted of a set of target word (noun/verb) instances (sen­tences/paragraphs). Participants were then asked to tag (disambiguate) each testing instance with the senses induced during the <i>training phase. </i>In the third and final phase, the tagged test instances were received by the organisers in order to evalu­ate the answers of the systems in a supervised and an unsupervised framework. Table 1 shows the to­tal number of target word instances in the training and testing set, as well as the average number of senses in the gold standard.</p><p>The main difference of the SemEval-2010 as compared to the SemEval-2007 sense induction task is that the training and testing data are treated separately, i.e the testing data are only used for sense tagging, while the training data are only used<page local="2" global="64"/></p><p><b>Training Phase</b></p><p>Training Instances</p><p>Word Sense Induction &amp; Disambiguation System <b>Testing Phase</b></p><p>Testing Instances</p><doubt alpha="92.9" length="14" tooSmall="True" monospace="0.0">\lnducedSenses</doubt><doubt alpha="79.2" length="24" tooSmall="True" monospace="0.0">Tagged Test/ Instances /</doubt><p>Test Instance Tagging <b>Evaluation Phase</b></p><figure caption="Figure 1: Training, testing and evaluation phases of SemEval-2010 Task 14"></figure><p>Table 1 : Training &amp; testing set details for sense induction. Treating the testing data as new unseen instances ensures a realistic evalua­tion that allows to evaluate the clustering models of each participating system.</p><p>The evaluation framework of SemEval-2010 WSI task considered two types of evaluation. In the first one, <i>unsupervised evaluation, </i>sys­tems' answers were evaluated according to: (1) <i>V-Measure </i>(Rosenberg and Hirschberg, 2007), and (2) <i>paired F-Score </i>(Artiles et al., 2009). Nei­ther of these measures were used in the SemEval-2007 WSI task. Manandhar &amp; Klapaftis (2009) provide more details on the choice of this evalu­ation setting and its differences with the previous evaluation. The second type of evaluation, <i>super­vised evaluation, </i>follows the supervised evalua­tion of the SemEval-2007 WSI task (Agirre and Soroa, 2007). In this evaluation, induced senses are mapped to gold standard senses using a map­ping corpus, and systems are then evaluated in a standard WSD task.</p><subsection number="2.1" title="Training dataset"><p>The target word dataset consisted of 100 words, i.e. 50 nouns and 50 verbs. The training dataset for each target noun or verb was created by follow­ing a web-based semi-automatic method, similar to the method for the construction of <i>Topic Signa­tures </i>(Agirre et al., 2001). Specifically, for each WordNet (Fellbaum, 1998) sense of a target word, we created a query of the following form:</p><p><i>&lt;Target Word&gt; <b>AND </b>&lt;Relative Set&gt;</i></p><p>The <i>&lt;Target Word&gt; </i>consisted of the target word stem. The <i>&lt;Relative Set&gt; </i>consisted of a disjunctive set of word lemmas that were related to the target word sense for which the query was created. The relations considered were WordNet's hypernyms, hyponyms, synonyms, meronyms and holonyms. Each query was manually checked by one of the organisers to remove ambiguous words. The following example shows the query created for the first<footnote anchor="1"/> and second<footnote anchor="2"/> WordNet sense of the target <i>noun failure.</i></p><p>The created queries were issued to Yahoo! search API<footnote anchor="3"/> and for each query a maximum of 1000 pages were downloaded. For each page we extracted fragments of text that occurred in &lt;p&gt; &lt;/p&gt; html tags and contained the target word stem. In the final stage, each extracted fragment of text was POS-tagged using the Genia tagger (Tsu-ruoka and Tsujii, 2005) and was only retained, if the POS of the target word in the extracted text matched the POS of the target word in our dataset.</p></subsection><subsection number="2.2" title="Testing dataset"><p>The testing dataset consisted of instances of the same target words from the training dataset. This dataset is part of OntoNotes (Hovy et al., 2006). We used the sense-tagged dataset in which sen­tences containing target word instances are tagged with OntoNotes (Hovy et al., 2006) senses. The texts come from various news sources including CNN, ABC and others.</p><footnote label="1">An act that fails</footnote><footnote label="2">An event that does not accomplish its intended purpose 3 http://developer.yahoo.com/search/ [Access: 10/04/2010]</footnote><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p><b>Training set</b></p></td><td class="cell"><p><b>Testing set</b></p></td><td class="cell"><p><b>Senses(#)</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>All</b></p></td><td class="cell"><p>879807</p></td><td class="cell"><p>8915</p></td><td class="cell"><p>3.79</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Nouns</b></p></td><td class="cell"><p>716945</p></td><td class="cell"><p>5285</p></td><td class="cell"><p>4.46</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Verbs</b></p></td><td class="cell"><p>162862</p></td><td class="cell"><p>3630</p></td><td class="cell"><p>3.12</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 2: Training set creation: example queries for target wordfailure" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Word Sense</b></p></td><td class="cell"><p><b>Query</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Sense 1</p></td><td class="cell"><p>failure <b>AND </b>(loss <b>OR </b>nonconformity <b>OR </b>test <b>OR </b>surrender <b>OR </b>"force play" <b>OR ...)</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Sense 2</p></td><td class="cell"><p>failure <b>AND </b>(ruination <b>OR </b>flop <b>OR </b>bust <b>OR </b>stall <b>OR </b>ruin <b>OR </b>walloping <b>OR ...)</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="3" global="65"/><p>For the purposes of this section we provide an ex­ample (Table 3) in which a target word has 181 instances and 3 GS senses. A system has gener­ated a clustering solution with 4 clusters covering all instances. Table 3 shows the number of com­mon instances between clusters and GS senses.</p></subsection><subsection number="3.1" title="Unsupervised evaluation"><p>This section presents the measures of unsuper­vised evaluation, i.e <i>V-Measure </i>(Rosenberg and Hirschberg, 2007) and (2) <i>paired F-Score </i>(Artiles et al., 2009).</p><subsubsection number="3.1.1" title="V-Measure evaluation"><p>Let <i>w </i>be a target word with <i>N </i>instances (data points) in the testing dataset. Let <i>K</i><i> </i><i>=</i><i> </i><i>{Cj\j</i><i> </i><i>=</i><i> </i>1... <i>n}</i><i> </i>be a set of automatically generated clus­ters grouping these instances, and <i>S = {Gi\i = </i>1... m} the set of gold standard classes contain­ing the desirable groupings of <i>w </i>instances.</p><p>V-Measure (Rosenberg and Hirschberg, 2007) assesses the quality of a clustering solution by ex­plicitly measuring its <i>homogeneity </i>and its <i>com­pleteness. </i>Homogeneity refers to the degree that each cluster consists of data points primarily be­longing to a single GS class, while completeness refers to the degree that each GS class consists of data points primarily assigned to a single cluster (Rosenberg and Hirschberg, 2007). Let <i>h </i>be ho­mogeneity and c completeness. V-Measure is the harmonic mean of <i>h</i><i> </i>and c, i.e. <i>VM</i><i> </i><i>=</i><i> </i><b>Homogeneity. </b>The homogeneity, <i>h, </i>of a clus­tering solution is defined in Formula 1, where <i>H(S\K)</i><i> </i>is the conditional entropy of the class distribution given the proposed clustering and <i>H(S)</i><i> </i>is the class entropy.</p><doubt alpha="33.3" length="3" tooSmall="False" monospace="0.0">h =</doubt><doubt alpha="41.7" length="12" tooSmall="False" monospace="0.0">H(S)=H{S\K)-</doubt><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">1,</doubt><doubt alpha="29.4" length="17" tooSmall="False" monospace="0.0">-, _h(s\k)1h(s) '</doubt><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">N</doubt><doubt alpha="28.6" length="7" tooSmall="False" monospace="0.0">\k\ \s\</doubt><p>if <i>H(S)</i><i> </i><i>=</i><i> </i>0 otherwise</p><doubt alpha="25.0" length="4" tooSmall="False" monospace="0.0">,\k\</doubt><doubt alpha="100.0" length="2" tooSmall="False" monospace="0.0">EE</doubt><doubt alpha="33.3" length="6" tooSmall="False" monospace="0.0">3=1i=l</doubt><p><i>a.</i></p><doubt alpha="83.3" length="6" tooSmall="False" monospace="0.0">N8sHsl</doubt><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">(1)</doubt><doubt alpha="0.0" length="7" tooSmall="False" monospace="0.0">(2) (3)</doubt><p>When <i>H{S\K)</i><i> </i>is 0, the solution is perfectly homogeneous, because each cluster only contains data points that belong to a single class. How­ever in an imperfect situation, <i>H(S\K)</i><i> </i>depends on the size of the dataset and the distribution of class sizes. Hence, instead of taking the raw con­ditional entropy, V-Measure normalises it by the maximum reduction in entropy the clustering in­formation could provide, i.e. <i>H(S).</i><i> </i>When there is only a single class <i>(H(S)</i><i> </i><i>=</i><i> </i>0), any clustering would produce a perfectly homogeneous solution. <b>Completeness. </b>Symmetrically to homogeneity, the completeness, c, of a clustering solution is de­fined in Formula 4, where <i>H(K\S)</i><i> </i>is the condi­tional entropy of the cluster distribution given the class distribution and <i>H(K)</i><i> </i>is the clustering en­tropy. When <i>H{K\S)</i><i> </i>is 0, the solution is perfectly complete, because all data points of a class belong to the same cluster.</p><p>For the clustering example in Table 3, homo­geneity is equal to 0.404, completeness is equal to 0.37 and V-Measure is equal to 0.386.</p><doubt alpha="33.3" length="3" tooSmall="False" monospace="0.0">c =</doubt><doubt alpha="45.5" length="11" tooSmall="False" monospace="0.0">h{k\s)h{k):</doubt><p>if <i>H{K)</i><i> </i><i>=</i><i> </i>0 otherwise</p><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">(4)</doubt><doubt alpha="45.5" length="11" tooSmall="False" monospace="0.0">H(K)=H{K\S)</doubt><doubt alpha="33.3" length="3" tooSmall="False" monospace="0.0">\k\</doubt><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">E</doubt><doubt alpha="28.6" length="7" tooSmall="False" monospace="0.0">\s\ \k\</doubt><doubt alpha="100.0" length="2" tooSmall="False" monospace="0.0">ai</doubt><doubt alpha="75.0" length="4" tooSmall="False" monospace="0.0">log-</doubt><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">13</doubt><doubt alpha="66.7" length="6" tooSmall="False" monospace="0.0">Nst\k\</doubt><doubt alpha="62.5" length="16" tooSmall="False" monospace="0.0">i=l j=lZ^fc=la%k</doubt><doubt alpha="0.0" length="7" tooSmall="False" monospace="0.0">(5) (6)</doubt></subsubsection><subsubsection number="3.1.2" title="Paired F-Score evaluation"><p>In this evaluation, the clustering problem is trans­formed into a classification problem. For each cluster <i>Ci </i>we generate ('^') instance pairs, where I Ci I is the total number of instances that belong to cluster <i>Ci. </i>Similarly, for each GS class <i>Gi </i>we gen­erate ('&lt;^i') instance pairs, where <i>\Gi\ </i>is the total number of instances that belong to GS class <i>Gi.</i></p><p>Let <i>F(K)</i><i> </i>be the set of instance pairs that ex­ist in the automatically induced clusters and <i>F{S)</i><i> </i>be the set of instance pairs that exist in the gold standard. Precision can be defined as the number of common instance pairs between the two sets to the total number of pairs in the clustering solu­tion (Equation 7), while recall can be defined as the number of common instance pairs between the two sets to the total number of pairs in the gold standard (Equation 8).<page local="4" global="66"/> Finally, precision and re­call are combined to produce the harmonic mean</p><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">3=</doubt><doubt alpha="71.4" length="7" tooSmall="False" monospace="0.0">VlSla -</doubt><doubt alpha="100.0" length="5" tooSmall="False" monospace="0.0">VlSla</doubt><doubt alpha="50.0" length="2" tooSmall="False" monospace="0.0">i=</doubt><doubt alpha="20.0" length="5" tooSmall="False" monospace="0.0">J3 =_</doubt><doubt alpha="40.0" length="5" tooSmall="False" monospace="0.0">-V1a-</doubt><table caption="Table 3: Clusters &amp; GS senses matrix.3   Evaluation framework" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Gi</p></td><td class="cell"><p>G2</p></td><td class="cell"><p>G3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Ci</p></td><td class="cell"><p>10</p></td><td class="cell"><p>10</p></td><td class="cell"><p>15</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>c2</p></td><td class="cell"><p>20</p></td><td class="cell"><p>50</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>c3</p></td><td class="cell"><p>1</p></td><td class="cell"><p>10</p></td><td class="cell"><p>60</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>Ci</i></p></td><td class="cell"><p>5</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><doubt alpha="40.0" length="5" tooSmall="False" monospace="0.0">(FS =</doubt><doubt alpha="40.0" length="5" tooSmall="False" monospace="0.0">2-P-R</doubt><doubt alpha="66.7" length="3" tooSmall="False" monospace="0.0">P+R</doubt><doubt alpha="33.3" length="3" tooSmall="False" monospace="0.0">P =</doubt><doubt alpha="33.3" length="3" tooSmall="False" monospace="0.0">R =</doubt><doubt alpha="45.5" length="11" tooSmall="False" monospace="0.0">\F(K)nF(S)\</doubt><doubt alpha="41.2" length="17" tooSmall="False" monospace="0.0">\F(K)\\F(K)nF(S)\</doubt><doubt alpha="33.3" length="6" tooSmall="False" monospace="0.0">\F(S)\</doubt><doubt alpha="0.0" length="7" tooSmall="False" monospace="0.0">(7) (8)</doubt><p>For example in Table 3, we can generate ( 2 ) in­stance pairs for Ci , (<footnote anchor="7"/>2°) for G2, f^<footnote anchor="1"/>) for G3 and (2) for C4, resulting in a total of 5505 instance pairs. In the same vein, we can generate (<footnote anchor="3"/>2<footnote anchor="6"/>) in­stance pairs for Gi, (<footnote anchor="7"/>2°) for G 2 and (<footnote anchor="7"/>2<footnote anchor="5"/>) for G3. In total, the GS classes contain 5820 instance pairs. There are 3435 common instance pairs, hence pre­cision is equal to 62.39%, recall is equal to 59.09% and paired F-Score is equal to 60.69%.</p></subsubsection></subsection><subsection number="3.2" title="Supervised evaluation"><p>In this evaluation, the testing dataset is split into a mapping and an evaluation corpus. The first one is used to map the automatically induced clusters to GS senses, while the second is used to evaluate methods in a WSD setting. This evaluation fol­lows the supervised evaluation of SemEval-2007 WSI task (Agirre and Soroa, 2007), with the dif­ference that the reported results are an average of 5 random splits. This repeated random sam­pling was performed to avoid the problems of the SemEval-2007 WSI challenge, in which different splits were providing different system rankings.</p><p>Let us consider the example in Table 3 and as­sume that this matrix has been created by using the mapping corpus. Table 3 shows that <i>C\ </i>is more likely to be associated with G3, G2 is more likely to be associated with <i>G</i><i>2, </i>G3 is more likely to be associated with G3 and G4 is more likely to be as­sociated with Gi. This information can be utilised to map the clusters to GS senses.</p><p>Particularly, the matrix shown in Table 3 is nor­malised to produce a matrix <i>M,</i><i> </i>in which each entry depicts the estimated conditional probabil­ity <i>P(Gi\Cj).</i><i> </i>Given an instance / of <i>tw </i>from the evaluation corpus, a row cluster vector <i>IC </i>is created, in which each entry <i>k </i>corresponds to the score assigned to G&amp; to be the winning cluster for instance I. The product of <i>IC </i>and <i>M </i>provides a row sense vector, <i>IG,</i><i> </i>in which the highest scor­ing entry <i>a </i>denotes that <i>Ga </i>is the winning sense. For example, if we produce the row cluster vector [Gi = 0.8, G2 = 0.1, G3 = 0.1, G4 = 0.0], and multiply it with the normalised matrix of Table 3, then we would get a row sense vector in which G3 would be the winning sense with a score equal to 0.43.</p></subsection></section><section number="4" title="Evaluation results"><p>In this section, we present the results of the 26 systems along with two baselines. The first base­line, Most Frequent Sense <i>(MFS),</i><i> </i>groups all test­ing instances of a target word into one cluster. The second baseline, <i>Random, </i>randomly assigns an in­stance to one out of four clusters. The number of clusters of <i>Random </i>was chosen to be roughly equal to the average number of senses in the GS. This baseline is executed five times and the results are averaged.</p><subsection number="4.1" title="Unsupervised evaluation"><p>Table 4 shows the V-Measure (VM) performance of the 26 systems participating in the task. The last column shows the number of induced clusters of each system in the test set.The <i>MFS </i>baseline has a V-Measure equal to 0, since by definition its com­pleteness is 1 and homogeneity is 0. All systems outperform this baseline, apart from one, whose V-Measure is equal to 0. Regarding the <i>Random </i>baseline, we observe that 17 perform better, which indicates that they have learned useful information better than chance.</p><table caption="Table 4 also shows that V-Measure tends to favour systems producing a higher number of clus-"></table><table caption="Table 4: V-Measure unsupervised evaluation" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>System</b></p></td><td class="cell"><p><b>VM (%) (All)</b></p></td><td class="cell"><p><b>VM(%) (Nouns)</b></p></td><td class="cell"><p><b>VM(%) (Verbs)</b></p></td><td class="cell"><p><b>#C1</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Hermit</p></td><td class="cell"><p>16.2</p></td><td class="cell"><p>16.7</p></td><td class="cell"><p>15.6</p></td><td class="cell"><p>10.78</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>UoY</p></td><td class="cell"><p>15.7</p></td><td class="cell"><p>20.6</p></td><td class="cell"><p>8.5</p></td><td class="cell"><p>11.54</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KSU KDD</p></td><td class="cell"><p>15.7</p></td><td class="cell"><p>18</p></td><td class="cell"><p>12.4</p></td><td class="cell"><p>17.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI</p></td><td class="cell"><p>9</p></td><td class="cell"><p>11.4</p></td><td class="cell"><p>5.7</p></td><td class="cell"><p>4.15</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-SVD</p></td><td class="cell"><p>9</p></td><td class="cell"><p>11.4</p></td><td class="cell"><p>5.7</p></td><td class="cell"><p>4.15</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-110</p></td><td class="cell"><p>8.6</p></td><td class="cell"><p>8.6</p></td><td class="cell"><p>8.5</p></td><td class="cell"><p>9.71</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Co</p></td><td class="cell"><p>7.9</p></td><td class="cell"><p>9.2</p></td><td class="cell"><p>6</p></td><td class="cell"><p>2.49</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PCGD</p></td><td class="cell"><p>7.8</p></td><td class="cell"><p>7.3</p></td><td class="cell"><p>8.4</p></td><td class="cell"><p>2.9</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PC</p></td><td class="cell"><p>7.5</p></td><td class="cell"><p>7.7</p></td><td class="cell"><p>7.3</p></td><td class="cell"><p>2.92</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PC-2</p></td><td class="cell"><p>7.1</p></td><td class="cell"><p>7.7</p></td><td class="cell"><p>6.1</p></td><td class="cell"><p>2.93</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Narrow-Gap</p></td><td class="cell"><p>6.9</p></td><td class="cell"><p>8</p></td><td class="cell"><p>5.1</p></td><td class="cell"><p>2.42</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GD-2</p></td><td class="cell"><p>6.9</p></td><td class="cell"><p>6.1</p></td><td class="cell"><p>8</p></td><td class="cell"><p>2.82</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GD</p></td><td class="cell"><p>6.9</p></td><td class="cell"><p>5.9</p></td><td class="cell"><p>8.5</p></td><td class="cell"><p>2.78</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Narrow-PK2</p></td><td class="cell"><p>6.8</p></td><td class="cell"><p>7.8</p></td><td class="cell"><p>5.5</p></td><td class="cell"><p>2.68</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-MIX-PK2</p></td><td class="cell"><p>5.6</p></td><td class="cell"><p>5.8</p></td><td class="cell"><p>5.2</p></td><td class="cell"><p>2.66</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-15</p></td><td class="cell"><p>5.3</p></td><td class="cell"><p>5.4</p></td><td class="cell"><p>5.1</p></td><td class="cell"><p>4.97</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Co-Gap</p></td><td class="cell"><p>4.8</p></td><td class="cell"><p>5.6</p></td><td class="cell"><p>3.6</p></td><td class="cell"><p>1.6</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Random</p></td><td class="cell"><p>4.4</p></td><td class="cell"><p>4.2</p></td><td class="cell"><p>4.6</p></td><td class="cell"><p>4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-13</p></td><td class="cell"><p>3.6</p></td><td class="cell"><p>3.5</p></td><td class="cell"><p>3.7</p></td><td class="cell"><p>3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Gap</p></td><td class="cell"><p>3.1</p></td><td class="cell"><p>4.2</p></td><td class="cell"><p>1.5</p></td><td class="cell"><p>1.4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Gap</p></td><td class="cell"><p>3</p></td><td class="cell"><p>2.9</p></td><td class="cell"><p>3</p></td><td class="cell"><p>1.61</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Uni-PK2</p></td><td class="cell"><p>2.4</p></td><td class="cell"><p>0.8</p></td><td class="cell"><p>4.7</p></td><td class="cell"><p>2.04</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-12</p></td><td class="cell"><p>2.3</p></td><td class="cell"><p>2.2</p></td><td class="cell"><p>2.5</p></td><td class="cell"><p>2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PT</p></td><td class="cell"><p>1.9</p></td><td class="cell"><p>1</p></td><td class="cell"><p>3.1</p></td><td class="cell"><p>1.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Uni-Gap</p></td><td class="cell"><p>1.4</p></td><td class="cell"><p>0.2</p></td><td class="cell"><p>3</p></td><td class="cell"><p>1.39</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GDC</p></td><td class="cell"><p>7</p></td><td class="cell"><p>6.2</p></td><td class="cell"><p>7.8</p></td><td class="cell"><p>2.83</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>MFS</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-SVD-Gap</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0.1</p></td><td class="cell"><p>1.02</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="5" global="67"/><p>Table 5 : Paired F-Score unsupervised evaluation ters than the number of GS senses, although V-Measure does not increase monotonically with the number of clusters increasing. For that reason, we introduced the second unsupervised evaluation measure (paired F-Score) that penalises systems when they produce: (1) a higher number of clus­ters (low recall) or (2) a lower number of clusters (low precision), than the GS number of senses.</p><p>Table 5 shows the performance of systems us­ing the second unsupervised evaluation measure. In this evaluation, we observe that most of the sys­tems perform better than <i>Random. </i>Despite that, none of the systems outperform the <i>MFS </i>baseline. It seems that systems generating a smaller number of clusters than the GS number of senses are bi­ased towards the <i>MFS, </i>hence they are not able to perform better. On the other hand, systems gen­erating a higher number of clusters are penalised by this measure. Systems generating a number of clusters roughly the same as the GS tend to con­flate the GS senses lot more than the <i>MFS.</i></p></subsection><subsection number="4.2" title="Supervised evaluation results"><p>Table 6 shows the results of this evaluation for a 80-20 test set split, i.e. 80% for mapping and 20% for evaluation. The last columns shows the aver­age number of GS senses identified by each sys­tem in the five splits of the evaluation datasets. Overall, 14 systems outperform the MFS, while 17 of them perform better than <i>Random. </i>The ranking of systems in nouns and verbs is different. For instance, the highest ranked system in nouns is <i>UoY, </i>while in verbs <i>Duluth-Mix-Narrow-Gap.</i><i> </i>It seems that depending on the part-of-speech of the target word, different algorithms, features and parame­ters' tuning have different impact.</p><p>The supervised evaluation changes the distri­bution of clusters by mapping each cluster to a weighted vector of senses. Hence, it can poten­tially favour systems generating a high number of homogeneous clusters. For that reason, we applied a second testing set split, where 60% of the testing corpus was used for mapping and 40% for eval­uation. Reducing the size of the mapping corpus allows us to observe, whether the above statement is correct, since systems with a high number of clusters would suffer from unreliable mapping.</p><p>Table 7 shows the results of the second super­vised evaluation. The ranking of participants did not change significantly, i.e. we observe only dif­ferent rankings among systems belonging to the same participant. Despite that, Table 7 also shows that the reduction of the mapping corpus has a dif­ferent impact on systems generating a larger num­ber of clusters than the GS number of senses.</p><p>For instance, <i>UoY </i>that generates 11.54 clusters outperformed the <i>MFS </i>by 3.77% in the 80-20 split and by 3.71% in the 60-40 split. The reduction of the mapping corpus had a minimal impact on its performance. In contrast, <i>KSU KDD </i>that gener­ates 17.5 clusters was below the <i>MFS </i>by 6.49% in the 80-20 split and by 7.<page local="6" global="68"/>83% in the 60-40 split. The reduction of the mapping corpus had a larger impact in this case. This result indicates that the performance in this evaluation also depends on the distribution of instances within the clusters. Sys­tems generating a skewed distribution, in which a small number of homogeneous clusters tag the ma­jority of instances and a larger number of clusters tag only a few instances, are likely to have a bet­ter performance than systems that produce a more uniform distribution.</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>System</b></p></td><td class="cell"><p><b>FS(%) (AH)</b></p></td><td class="cell"><p><b>FS(%) (Nouns)</b></p></td><td class="cell"><p><b>FS(%) (Verbs)</b></p></td><td class="cell"><p><b>#C1</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>MFS</p></td><td class="cell"><p>63.5</p></td><td class="cell"><p>57.0</p></td><td class="cell"><p>72.7</p></td><td class="cell"><p>1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-SVD-Gap</p></td><td class="cell"><p>63.3</p></td><td class="cell"><p>57.0</p></td><td class="cell"><p>72.4</p></td><td class="cell"><p>1.02</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PT</p></td><td class="cell"><p>61.8</p></td><td class="cell"><p>56.4</p></td><td class="cell"><p>69.7</p></td><td class="cell"><p>1.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GD</p></td><td class="cell"><p>59.2</p></td><td class="cell"><p>51.6</p></td><td class="cell"><p>70.0</p></td><td class="cell"><p>2.78</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Gap</p></td><td class="cell"><p>59.1</p></td><td class="cell"><p>54.5</p></td><td class="cell"><p>65.8</p></td><td class="cell"><p>1.61</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Uni-Gap</p></td><td class="cell"><p>58.7</p></td><td class="cell"><p>57.0</p></td><td class="cell"><p>61.2</p></td><td class="cell"><p>1.39</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GD-2</p></td><td class="cell"><p>58.2</p></td><td class="cell"><p>50.4</p></td><td class="cell"><p>69.3</p></td><td class="cell"><p>2.82</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GDC</p></td><td class="cell"><p>57.3</p></td><td class="cell"><p>48.5</p></td><td class="cell"><p>70.0</p></td><td class="cell"><p>2.83</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Uni-PK2</p></td><td class="cell"><p>56.6</p></td><td class="cell"><p>57.1</p></td><td class="cell"><p>55.9</p></td><td class="cell"><p>2.04</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PC</p></td><td class="cell"><p>55.5</p></td><td class="cell"><p>50.4</p></td><td class="cell"><p>62.9</p></td><td class="cell"><p>2.92</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PC-2</p></td><td class="cell"><p>54.7</p></td><td class="cell"><p>49.7</p></td><td class="cell"><p>61.7</p></td><td class="cell"><p>2.93</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Gap</p></td><td class="cell"><p>53.7</p></td><td class="cell"><p>53.4</p></td><td class="cell"><p>53.9</p></td><td class="cell"><p>1.4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PCGD</p></td><td class="cell"><p>53.3</p></td><td class="cell"><p>44.8</p></td><td class="cell"><p>65.6</p></td><td class="cell"><p>2.9</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Co-Gap</p></td><td class="cell"><p>52.6</p></td><td class="cell"><p>53.3</p></td><td class="cell"><p>51.5</p></td><td class="cell"><p>1.6</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-MIX-PK2</p></td><td class="cell"><p>50.4</p></td><td class="cell"><p>51.7</p></td><td class="cell"><p>48.3</p></td><td class="cell"><p>2.66</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>UoY</p></td><td class="cell"><p>49.8</p></td><td class="cell"><p>38.2</p></td><td class="cell"><p>66.6</p></td><td class="cell"><p>11.54</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Narrow-Gap</p></td><td class="cell"><p>49.7</p></td><td class="cell"><p>47.4</p></td><td class="cell"><p>51.3</p></td><td class="cell"><p>2.42</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Co</p></td><td class="cell"><p>49.5</p></td><td class="cell"><p>50.2</p></td><td class="cell"><p>48.2</p></td><td class="cell"><p>2.49</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Narrow-PK2</p></td><td class="cell"><p>47.8</p></td><td class="cell"><p>37.1</p></td><td class="cell"><p>48.2</p></td><td class="cell"><p>2.68</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-12</p></td><td class="cell"><p>47.8</p></td><td class="cell"><p>44.3</p></td><td class="cell"><p>52.6</p></td><td class="cell"><p>2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-SVD</p></td><td class="cell"><p>41.1</p></td><td class="cell"><p>37.1</p></td><td class="cell"><p>46.7</p></td><td class="cell"><p>4.15</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI</p></td><td class="cell"><p>41.1</p></td><td class="cell"><p>37.1</p></td><td class="cell"><p>46.7</p></td><td class="cell"><p>4.15</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-13</p></td><td class="cell"><p>38.4</p></td><td class="cell"><p>36.2</p></td><td class="cell"><p>41.5</p></td><td class="cell"><p>3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KSUKDD</p></td><td class="cell"><p>36.9</p></td><td class="cell"><p>24.6</p></td><td class="cell"><p>54.7</p></td><td class="cell"><p>17.5</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Random</p></td><td class="cell"><p>31.9</p></td><td class="cell"><p>30.4</p></td><td class="cell"><p>34.1</p></td><td class="cell"><p>4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-15</p></td><td class="cell"><p>27.6</p></td><td class="cell"><p>26.7</p></td><td class="cell"><p>28.9</p></td><td class="cell"><p>4.97</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Hermit</p></td><td class="cell"><p>26.7</p></td><td class="cell"><p>24.4</p></td><td class="cell"><p>30.1</p></td><td class="cell"><p>10.78</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-110</p></td><td class="cell"><p>16.1</p></td><td class="cell"><p>15.8</p></td><td class="cell"><p>16.4</p></td><td class="cell"><p>9.71</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 6: Supervised recall (SR) (test set split:80% mapping, 20% evaluation)" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>System</b></p></td><td class="cell"><p><b>SR(%) (All)</b></p></td><td class="cell"><p><b>SR (%) (Nouns)</b></p></td><td class="cell"><p><b>SR(%) (Verbs)</b></p></td><td class="cell"><p><b>#S</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>UoY</p></td><td class="cell"><p>62.4</p></td><td class="cell"><p>59.4</p></td><td class="cell"><p>66.8</p></td><td class="cell"><p>1.51</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI</p></td><td class="cell"><p>60.5</p></td><td class="cell"><p>54.7</p></td><td class="cell"><p>68.9</p></td><td class="cell"><p>1.66</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-SVD</p></td><td class="cell"><p>60.5</p></td><td class="cell"><p>54.7</p></td><td class="cell"><p>68.9</p></td><td class="cell"><p>1.66</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Co-Gap</p></td><td class="cell"><p>60.3</p></td><td class="cell"><p>54.1</p></td><td class="cell"><p>68.6</p></td><td class="cell"><p>1.19</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Co</p></td><td class="cell"><p>60.8</p></td><td class="cell"><p>54.7</p></td><td class="cell"><p>67.6</p></td><td class="cell"><p>1.51</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Gap</p></td><td class="cell"><p>59.8</p></td><td class="cell"><p>54.4</p></td><td class="cell"><p>67.8</p></td><td class="cell"><p>1.11</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PC-2</p></td><td class="cell"><p>59.8</p></td><td class="cell"><p>54.1</p></td><td class="cell"><p>68.0</p></td><td class="cell"><p>1.21</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PC</p></td><td class="cell"><p>59.7</p></td><td class="cell"><p>54.6</p></td><td class="cell"><p>67.3</p></td><td class="cell"><p>1.39</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PCGD</p></td><td class="cell"><p>59.5</p></td><td class="cell"><p>53.3</p></td><td class="cell"><p>68.6</p></td><td class="cell"><p>1.47</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GDC</p></td><td class="cell"><p>59.1</p></td><td class="cell"><p>53.4</p></td><td class="cell"><p>67.4</p></td><td class="cell"><p>1.34</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GD</p></td><td class="cell"><p>59.0</p></td><td class="cell"><p>53.0</p></td><td class="cell"><p>67.9</p></td><td class="cell"><p>1.33</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PT</p></td><td class="cell"><p>58.9</p></td><td class="cell"><p>53.1</p></td><td class="cell"><p>67.4</p></td><td class="cell"><p>1.08</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GD-2</p></td><td class="cell"><p>58.7</p></td><td class="cell"><p>52.8</p></td><td class="cell"><p>67.4</p></td><td class="cell"><p>1.33</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-SVD-Gap</p></td><td class="cell"><p>58.7</p></td><td class="cell"><p>53.2</p></td><td class="cell"><p>66.7</p></td><td class="cell"><p>1.01</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>MFS</p></td><td class="cell"><p>58.7</p></td><td class="cell"><p>53.2</p></td><td class="cell"><p>66.6</p></td><td class="cell"><p>1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-12</p></td><td class="cell"><p>58.5</p></td><td class="cell"><p>53.1</p></td><td class="cell"><p>66.4</p></td><td class="cell"><p>1.25</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Hermit</p></td><td class="cell"><p>58.3</p></td><td class="cell"><p>53.6</p></td><td class="cell"><p>65.3</p></td><td class="cell"><p>2.06</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-13</p></td><td class="cell"><p>58.0</p></td><td class="cell"><p>52.3</p></td><td class="cell"><p>66.4</p></td><td class="cell"><p>1.46</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Random</p></td><td class="cell"><p>57.3</p></td><td class="cell"><p>51.5</p></td><td class="cell"><p>65.7</p></td><td class="cell"><p>1.53</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-15</p></td><td class="cell"><p>56.8</p></td><td class="cell"><p>50.9</p></td><td class="cell"><p>65.3</p></td><td class="cell"><p>1.61</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Narrow-Gap</p></td><td class="cell"><p>56.6</p></td><td class="cell"><p>48.1</p></td><td class="cell"><p>69.1</p></td><td class="cell"><p>1.43</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Narrow-PK2</p></td><td class="cell"><p>56.1</p></td><td class="cell"><p>47.5</p></td><td class="cell"><p>68.7</p></td><td class="cell"><p>1.41</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-110</p></td><td class="cell"><p>54.8</p></td><td class="cell"><p>48.3</p></td><td class="cell"><p>64.2</p></td><td class="cell"><p>1.94</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KSUKDD</p></td><td class="cell"><p>52.2</p></td><td class="cell"><p>46.6</p></td><td class="cell"><p>60.3</p></td><td class="cell"><p>1.69</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-MIX-PK2</p></td><td class="cell"><p>51.6</p></td><td class="cell"><p>41.1</p></td><td class="cell"><p>67.0</p></td><td class="cell"><p>1.23</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Gap</p></td><td class="cell"><p>50.6</p></td><td class="cell"><p>40.0</p></td><td class="cell"><p>66.0</p></td><td class="cell"><p>1.01</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Uni-PK2</p></td><td class="cell"><p>19.3</p></td><td class="cell"><p>1.8</p></td><td class="cell"><p>44.8</p></td><td class="cell"><p>0.62</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Uni-Gap</p></td><td class="cell"><p>18.7</p></td><td class="cell"><p>1.6</p></td><td class="cell"><p>43.8</p></td><td class="cell"><p>0.56</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></subsection></section><section number="5" title="Conclusion"><p>We presented the description, evaluation frame­work and assessment of systems participating in the SemEval-2010 sense induction task. The eval­uation has shown that the current state-of-the-art lacks unbiased measures that objectively evaluate clustering.</p><p>The results of systems have shown that their performance in the unsupervised and supervised evaluation settings depends on cluster granularity along with the distribution of instances within the clusters. Our future work will focus on the assess­ment of sense induction on a task-oriented basis as well as on clustering evaluation.</p></section><section title="Acknowledgements"><p>We gratefully acknowledge the support of the EU FP7 INDECT project, Grant No. 218086, the National Science Foundation Grant NSF-0715078, Consistent Criteria for Word Sense Disambigua­tion, and the GALE program of the Defense Ad­vanced Research Projects Agency, Contract No. HR0011-06-C-0022, a subcontract from the BBN-AGILE Team.</p><table caption="Table 7: Supervised recall (SR) (test set split:60% mapping, 40% evaluation)" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>System</b></p></td><td class="cell"><p><b>SR(%) (AH)</b></p></td><td class="cell"><p><b>SR(%) (Nouns)</b></p></td><td class="cell"><p><b>SR(%) (Verbs)</b></p></td><td class="cell"><p>#s</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>UoY</p></td><td class="cell"><p>62.0</p></td><td class="cell"><p>58.6</p></td><td class="cell"><p>66.8</p></td><td class="cell"><p>1.66</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Co</p></td><td class="cell"><p>60.1</p></td><td class="cell"><p>54.6</p></td><td class="cell"><p>68.1</p></td><td class="cell"><p>1.56</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Co-Gap</p></td><td class="cell"><p>59.5</p></td><td class="cell"><p>53.5</p></td><td class="cell"><p>68.3</p></td><td class="cell"><p>1.2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-SVD</p></td><td class="cell"><p>59.5</p></td><td class="cell"><p>53.5</p></td><td class="cell"><p>68.3</p></td><td class="cell"><p>1.73</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI</p></td><td class="cell"><p>59.5</p></td><td class="cell"><p>53.5</p></td><td class="cell"><p>68.3</p></td><td class="cell"><p>1.73</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-Gap</p></td><td class="cell"><p>59.3</p></td><td class="cell"><p>53.2</p></td><td class="cell"><p>68.2</p></td><td class="cell"><p>1.11</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PCGD</p></td><td class="cell"><p>59.1</p></td><td class="cell"><p>52.6</p></td><td class="cell"><p>68.6</p></td><td class="cell"><p>1.54</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PC-2</p></td><td class="cell"><p>58.9</p></td><td class="cell"><p>53.4</p></td><td class="cell"><p>67.0</p></td><td class="cell"><p>1.25</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PC</p></td><td class="cell"><p>58.9</p></td><td class="cell"><p>53.6</p></td><td class="cell"><p>66.6</p></td><td class="cell"><p>1.44</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GDC</p></td><td class="cell"><p>58.3</p></td><td class="cell"><p>52.1</p></td><td class="cell"><p>67.3</p></td><td class="cell"><p>1.41</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GD</p></td><td class="cell"><p>58.3</p></td><td class="cell"><p>51.9</p></td><td class="cell"><p>67.6</p></td><td class="cell"><p>1.42</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>MFS</p></td><td class="cell"><p>58.3</p></td><td class="cell"><p>52.5</p></td><td class="cell"><p>66.7</p></td><td class="cell"><p>1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-PT</p></td><td class="cell"><p>58.3</p></td><td class="cell"><p>52.2</p></td><td class="cell"><p>67.1</p></td><td class="cell"><p>1.11</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-WSI-SVD-Gap</p></td><td class="cell"><p>58.2</p></td><td class="cell"><p>52.5</p></td><td class="cell"><p>66.7</p></td><td class="cell"><p>1.01</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KCDC-GD-2</p></td><td class="cell"><p>57.9</p></td><td class="cell"><p>51.7</p></td><td class="cell"><p>67.0</p></td><td class="cell"><p>1.44</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-12</p></td><td class="cell"><p>57.7</p></td><td class="cell"><p>51.7</p></td><td class="cell"><p>66.4</p></td><td class="cell"><p>1.27</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-13</p></td><td class="cell"><p>57.6</p></td><td class="cell"><p>51.1</p></td><td class="cell"><p>67.0</p></td><td class="cell"><p>1.48</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Hermit</p></td><td class="cell"><p>57.3</p></td><td class="cell"><p>52.5</p></td><td class="cell"><p>64.2</p></td><td class="cell"><p>2.27</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-15</p></td><td class="cell"><p>56.5</p></td><td class="cell"><p>50.0</p></td><td class="cell"><p>66.1</p></td><td class="cell"><p>1.76</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Random</p></td><td class="cell"><p>56.5</p></td><td class="cell"><p>50.2</p></td><td class="cell"><p>65.7</p></td><td class="cell"><p>1.65</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Narrow-Gap</p></td><td class="cell"><p>56.2</p></td><td class="cell"><p>47.7</p></td><td class="cell"><p>68.6</p></td><td class="cell"><p>1.51</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Narrow-PK2</p></td><td class="cell"><p>55.7</p></td><td class="cell"><p>46.9</p></td><td class="cell"><p>68.5</p></td><td class="cell"><p>1.51</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-R-110</p></td><td class="cell"><p>53.6</p></td><td class="cell"><p>46.7</p></td><td class="cell"><p>63.6</p></td><td class="cell"><p>2.18</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-MIX-PK2</p></td><td class="cell"><p>50.5</p></td><td class="cell"><p>39.7</p></td><td class="cell"><p>66.1</p></td><td class="cell"><p>1.31</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>KSUKDD</p></td><td class="cell"><p>50.4</p></td><td class="cell"><p>44.3</p></td><td class="cell"><p>59.4</p></td><td class="cell"><p>1.92</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Gap</p></td><td class="cell"><p>49.8</p></td><td class="cell"><p>38.9</p></td><td class="cell"><p>65.6</p></td><td class="cell"><p>1.04</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Uni-PK2</p></td><td class="cell"><p>19.1</p></td><td class="cell"><p>1.8</p></td><td class="cell"><p>44.4</p></td><td class="cell"><p>0.63</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duluth-Mix-Uni-Gap</p></td><td class="cell"><p>18.9</p></td><td class="cell"><p>1.5</p></td><td class="cell"><p>44.2</p></td><td class="cell"><p>0.56</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></section><references><p>Eneko Agirre and Aitor Soroa. 2007. SemEval-2007 Task 02: Evaluating Word Sense Induction and Dis­crimination Systems. In <i>Proceedings of SemEval-2007, </i>pages 7-12, Prague, Czech Republic. ACL.</p><p>Eneko Agirre, Olatz Ansa, David Martinez, and Eduard Hovy. 2001. Enriching Wordnet Concepts With Topic Signatures. <i>ArXiv Computer Science e-prints.</i></p><p>Javier Artiles, Enrique Amigo, and Julio Gonzalo. 2009. The role of named entities in web people search. In <i>Proceedings ofEMNLP, </i>pages 534-542. ACL.</p><p>Christiane Fellbaum. 1998. <i>Wordnet: An Electronic Lexical Database. </i>MIT Press, Cambridge, Mas­sachusetts, USA.</p><p>Eduard Hovy, Mitchell Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel. 2006. Ontonotes: the 90% solution. In <i>Proceedings ofNAACL, Com­panion Volume: Short Papers on </i><i>XX,</i><i> </i>pages 57-60. ACL.</p><p>Suresh Manandhar and Ioannis P. Klapaftis. 2009. Semeval-2010 Task 14: Evaluation Setting for Word Sense Induction &amp; Disambiguation Systems. In <i>DEW '09: Proceedings of the Workshop on Se­mantic Evaluations: Recent Achievements and Fu­ture Directions, </i>pages 117-122, Boulder, Colorado, USA. ACL.</p><p>Patrick Pantel and Dekang Lin. 2002. Discovering Word Senses from Text. In <i>KDD '02: Proceedings of the 8th ACM SIGKDD Conference, </i>pages 613-619, New York, NY, USA. ACM.</p><p>Andrew Rosenberg and Julia Hirschberg. 2007. V-measure: A Conditional Entropy-based External Cluster Evaluation Measure. In <i>Proceedings of the 2007 EMNLP-CoNLL Joint Conference, </i>pages 410-420, Prague, Czech Republic.</p><p>Yoshimasa Tsuruoka and Jumchi Tsujii. 2005. Bidi­rectional Inference With the Easiest-first Strategy for Tagging Sequence Data. In <i>Proceedings of the HLT-EMNLP Joint Conference, </i>pages 467^174, Morristown, NJ, USA.</p><p>Jean Véronis. 2004. Hyperlex: Lexical Cartography for Information Retrieval. <i>Computer Speech &amp; Lan­guage, </i>18(3):223-252.</p></references></body></article>