<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="36"/><title>SemEval-2007 Task 08: Metonymy Resolution at SemEval-2007</title><pubinfo>Proceedings of the 4th International Workshop on Semantic Evaluations (SemEval-2007),pages 36-41, Prague, June 2007. ©2007 Association for Computational Linguistics</pubinfo><author surname="Markert" givenname="Katja"><org  name="University of Leeds" country="United Kingdom" city="Leeds"/></author><author surname="Nissim" givenname="Malvina"><org  name="University of Bologna" country="Italy" city="Bologna"/></author></firstpageheader><frontmatter><p><b>SemEval-2007 Task 08: Metonymy Resolution at SemEval-2007</b></p><p><b>Katja Markert</b></p><p>School of Computing University of Leeds, UK markert@comp.leeds.ac.uk</p><p><b>Malvina Nissim</b></p><p>Dept. of Linguistics and Oriental Studies University of Bologna, Italy malvina.nissim@unibo.it</p></frontmatter><abstract>We provide an overview of the metonymy resolution shared task organised within SemEval-2007. We describe the problem, the data provided to participants, and the evaluation measures we used to assess per­formance. We also give an overview of the systems that have taken part in the task, and discuss possible directions for future work. </abstract></header><body><section number="1" title="Introduction"><p>Both word sense disambiguation and named entity recognition have benefited enormously from shared task evaluations, for example in the Senseval, MUC and CoNLL frameworks. Similar campaigns have not been developed for the resolution of figurative language, such as metaphor, metonymy, idioms and irony. However, resolution of figurative language is an important complement to and extension of word sense disambiguation as it often deals with word senses that are not listed in the lexicon. For exam­ple, the meaning of <i>stopover </i>in the sentence <i>He saw teaching as a stopover on his way to bigger things </i>is a metaphorical sense of the sense "stopping place in a physical journey", with the literal sense listed in WordNet 2.0 but the metaphorical one not being listed.<footnote anchor="1"/> The same holds for the metonymic reading of <i>rattlesnake </i>(for the animal's meat) in <i>Roast rat­tlesnake tastes like chicken<footnote anchor="2"/> </i>Again, the meat reading of <i>rattlesnake </i>is not listed in WordNet whereas the meat reading for <i>chicken </i>is.</p><footnote label="1">This example was taken from the Berkely Master Metaphor list (Lakoff and Johnson, 1980) .</footnote><footnote label="2">From now on, all examples in this paper are taken from the British National Corpus (BNC) (Burnard, 1995), but Ex. 23.</footnote><p>As there is no common framework or corpus for figurative language resolution, previous computa­tional works (Fass, 1997; Hobbs et al., 1993; Barn-den et al., 2003, among others) carry out only small-scale evaluations. In recent years, there has been growing interest in metaphor and metonymy resolu­tion that is either corpus-based or evaluated on larger datasets (Martin, 1994; Nissim and Markert, 2003; Mason, 2004; Peirsman, 2006; Birke and Sarkaar, 2006; Krishnakamuran and Zhu, 2007). Still, apart from (Nissim and Markert, 2003; Peirsman, 2006) who evaluate their work on the same dataset, results are hardly comparable as they all operate within dif­ferent frameworks.</p><p>This situation motivated us to organise the first shared task for figurative language, concentrating on metonymy. In metonymy one expression is used to refer to the referent of a related one, like the use of an animal name for its meat. Similarly, in Ex. 1, <i>Vietnam, </i>the name of a location, refers to an event (a war) that happened there.</p><p>(1) Sex, drugs, and <b>Vietnam </b>have haunted Bill Clinton's campaign.</p><p>In Ex. 2 and 3, <i>BMW, </i>the name ofa company, stands for its index on the stock market, or a vehicle manu­factured by BMW, respectively.</p><doubt alpha="60.9" length="23" tooSmall="False" monospace="0.0">(2)BMWslipped 4p to 31p</doubt><p>(3) His <b>BMW </b>went on to race at Le Mans</p><p>The importance of resolving metonymies has been shown for a variety of NLP tasks, such as machine translation (Kamei and Wakao, 1992), ques­tion answering (Stallard, 1993), anaphora resolution geographical information retrieval (Leveling and<page local="2" global="37"/></p><doubt alpha="57.8" length="45" tooSmall="False" monospace="0.0">(Harabagiu, 1998; Markert and Hahn, 2002) and</doubt><p>Hartrumpf, 2006).</p><p>Although metonymic readings are, like all figu­rative readings, potentially open ended and can be innovative, the regularity of usage for word groups helps in establishing a common evaluation frame­work. Many other location names, for instance, can be used in the same fashion as <i>Vietnam </i>in Ex. 1. Thus, given a semantic class (e.g. location), one can specify several regular metonymic patterns (e.g. place-for-event) that instances of the class are likely to undergo. In addition to literal readings, regu­lar metonymic patterns and innovative metonymic readings, there can also be so-called mixed read­ings, similar to zeugma, where both a literal and a metonymic reading are evoked (Nunberg, 1995).</p><p>The metonymy task is a lexical sample task for English, consisting of two subtasks, one concentrat­ing on the semantic class <i>location, </i>exemplified by country names, and another one concentrating on <i>or­ganisation, </i>exemplified by company names. Partici­pants had to automatically classify preselected coun­try/company names as having a literal or non-literal meaning, given a four-sentence context. Addition­ally, participants could attempt finer-grained inter­pretations, further specifying readings into prespec-ified metonymic patterns (such as place-for-event) and recognising innovative readings.</p></section><section number="2" title="Annotation Categories"><p>We distinguish between literal, metonymic, and mixed readings for locations and organisations. In the case of a metonymic reading, we also specify the actual patterns. The annotation categories were motivated by prior linguistic research by ourselves (Markert and Nissim, 2006), and others (Fass, 1997; Lakoff and Johnson, 1980).</p><subsection number="2.1" title="Locations"><p><b>Literal </b>readings for locations comprise <i>locative </i>(Ex. 4) and <i>political </i>entity interpretations (Ex. 5).</p><p>(4) coral coast of <b>Papua New Guinea.</b></p><p>(5) <b>Britain</b>'s current account deficit. <b>Metonymic </b>readings encompass four types:</p><p><b>- place-for-people </b>a place stands for any per­sons/organisations associated with it. These can be governments (Ex. 6), affiliated organisations, incl. sports teams (Ex. 7), or the whole population (Ex. 8). Often, the referent is underspecified (Ex. 9).</p><p>(6) <b>America </b>did once try to ban alcohol.</p><p>(7) <b>England </b>lost in the semi-final.</p><p>(8) [... ] the incarnation was to fulfil the promise to <b>Israel </b>and to reconcile the world with God.</p><p>(9) The G-24 group expressed readiness to pro­vide <b>Albania </b>with food aid.</p><p><b>- place-for-event </b>a location name stands for an event that happened in the location (see Ex. 1).</p><p><b>- place-for-product </b>a place stands for a product manufactured in the place, as <i>Bordeaux </i>in Ex. 10.</p><p>(10) a smooth <b>Bordeaux </b>that was gutsy enough to cope with our food <b>- othermet </b>a metonymy that does not fall into any of the prespecified patterns, as in Ex. 11, where <i>New Jersey </i>refers to typical local tunes.</p><p>(11) The thing about the record is the influ­ences of the music. The bottom end is very New York<b>/New Jersey </b>and the top is very melodic.</p><p>When two predicates are involved, triggering a dif­ferent reading each (Nunberg, 1995), the annotation category is <b>mixed. </b>In Ex. 12, both a literal and a place-for-people reading are involved.</p><p>(12) they arrived in <b>Nigeria, </b>hitherto a leading critic of [. . . ]</p></subsection><subsection number="2.2" title="Organisations"><p>The <b>literal </b>reading for organisation names describes references to the organisation in general, where an organisation is seen as a legal entity, which consists of organisation members that speak with a collec­tive voice, and which has a charter, statute or defined aims. Examples of literal readings include (among others) descriptions of the structure of an organisa­tion (see Ex. 13), associations between organisations (see Ex. 14) or relations between organisations and products/services they offer (see Ex. 15).</p><page local="3" global="38"/><p>(13) <b>NATO </b>countries</p><p>(14) <b>Sun </b>acquired that part of Eastman-Kodak Cos Unix subsidary (15) <b>Intel</b>'s Indeo video compression hardware <b>Metonymic readings </b>include six types:</p><p><b>- org-for-members </b>an organisation stands for its members, such as a spokesperson or official (Ex. 16), or all its employees, as in Ex. 17.</p><doubt alpha="61.5" length="39" tooSmall="False" monospace="0.0">(16) Last FebruaryIBMannounced [. . . ]</doubt><p>(17) It's customary to go to work in black or white suits. [. . . ] <b>Woolworths </b>wear them <b>- org-for-event </b>an organisation name is used to re­fer to an event associated with the organisation (e.g. a scandal or bankruptcy), as in Ex. 18.</p><p>(18) the resignation of Leon Brittan from Trade and Industry in the aftermath of <b>Westland.</b></p><p><b>- org-for-product </b>the name of a commercial or­ganisation can refer to its products, as in Ex. 3.</p><p><b>- org-for-facility </b>organisations can also stand for the facility that houses the organisation or one of its branches, as in the following example.</p><p>(19) The opening of a <b>McDonald's </b>is a major event <b>- org-for-index </b>an organisation name can be used for an index that indicates its value (see Ex. 2).</p><p><b>- othermet </b>a metonymy that does not fall into any of the prespecified patterns, as in Ex. 20, where <i>Bar­clays Bank </i>stands for an account at the bank.</p><doubt alpha="59.6" length="47" tooSmall="False" monospace="0.0">(20) funds [. . . ]  had been paid intoBarclays</doubt><p><b>Bank.</b></p><p><b>Mixed </b>readings exist for organisations as well. In Ex. 21, both an org-for-index and an org-for-members pattern are invoked.</p><p>(21) <b>Barclays </b>slipped 4p to 351p after confirm­ing 3,000 more job losses.</p></subsection><subsection number="2.3" title="Class-independent categories"><p>Apart from class-specific metonymic readings, some patterns seem to apply across classes to all names. In the SemEval dataset, we annotated two of them.</p><p><b>object-for-name </b>all names can be used as mere signifiers, instead of referring to an object or set of objects. In Ex. 22, both <i>Chevrolet </i>and <i>Ford </i>are used as strings, rather than referring to the companies.</p><p>(22) <b>Chevrolet </b>is feminine because of its sound (it's a longer word than <b>Ford, </b>has an open vowel at the end, connotes Frenchness).</p><p><b>object-for-representation </b>a name can refer to a representation (such as a photo or painting) of the referent of its literal reading. In Ex. 23, <i>Malta </i>refers to a drawing of the island when pointing to a map.</p><doubt alpha="64.7" length="17" tooSmall="False" monospace="0.0">(23) This isMalta</doubt></subsection></section><section number="3" title="Data Collection and Annotation"><p>We used the CIA Factbook<footnote anchor="3"/> and the Fortune 500 list as sampling frames for country and company names respectively. All occurrences (including plu­ral forms) of all names in the sampling frames were extracted in context from all texts of the BNC, Ver­sion 1.0. All samples extracted are coded in XML and contain up to four sentences: the sentence in which the country/company name occurs, two be­fore, and one after. If the name occurs at the begin­ning or end of a text the samples may contain less than four sentences.</p><p>For both the location and the organisation subtask, two random subsets of the extracted samples were selected as training and test set, respectively. Before metonymy annotation, samples that were not under­stood by the annotators because of insufficient con­text were removed from the datsets. In addition, a sample was also removed if the name extracted was a homonym not in the desired semantic class (for ex­ample <i>Mr. Greenland </i>when annotating locations).<footnote anchor="4"/>For those names that do have the semantic class location or organisation, metonymy anno­tation was performed, using the categories described in Section 2. All training set annotation was carried out independently by both organisers. Annotation was highly reliable with a <i>kappa </i>(Carletta, 1996) of<page local="4" global="39"/></p><footnote label="3">https://www.cia.gov/cia/publications/ factbook/index.html</footnote><footnote label="4">Given that the task is not about standard Named Entity Recognition, we assume that the general semantic class of the name is already known.</footnote><p>Table 1 : Reading distribution for locations .88/.89 for locations/organisations.<footnote anchor="5"/> As agreement was established, annotation of the test set was car­ried out by the first organiser. All cases which were not entirely straightforward were then independently checked by the second organiser. Samples whose readings could not be agreed on (after a reconcil­iation phase) were excluded from both training and test set. The reading distributions of training and test sets for both subtasks are shown in Tables 1 and 2.</p><p>In addition to a simple text format including only the metonymy annotation, we provided participants with several linguistic annotations of both training and testset. This included the original BNC tokeni-sation and part-of-speech tags as well as manually annotated dependency relations for each annotated name (e.g. <i>BMW subj-of-slip </i>for Ex. 2).</p></section><section number="4" title="Submission and Evaluation"><p>Teams were allowed to participate in the location or organisation task or both. We encouraged super­vised, semi-supervised or unsupervised approaches.</p><p>Systems could be tailored to recognise metonymies at three different levels of granularity: <i>coarse, medium, </i>or <i>fine, </i>with an increasing number and specification of target classification categories, and thus difficulty. At the <i>coarse </i>level, only a distinction between literal and non-literal was asked for; <i>medium </i>asked for a distinction between literal, metonymic and mixed readings; <i>fine </i>needed a classification into literal readings, mixed readings, any of the class-dependent and class-independent metonymic patterns (Section 2) or an innovative metonymic reading (category othermet).</p><footnote label="5">The training sets are part of the already available Mascara corpus for metonymy (Markert and Nissim, 2006). The test sets were newly created for SemEval.</footnote><p>Systems were evaluated via accuracy (acc) and coverage (cov), allowing for partial submissions.</p><doubt alpha="60.0" length="5" tooSmall="False" monospace="0.0">acc =</doubt><doubt alpha="60.0" length="5" tooSmall="False" monospace="0.0">cov =</doubt><p>For each target category c we also measured:</p><p><i>precisioncrecallc = fscorec</i></p><p><i><u># correct assignments </u></i><i><u>of</u></i><i><u> </u></i><i><u>c</u></i><i><u> </u># assignments </i><i>of</i><i> </i><i>c</i><i> <u># correct assignments </u></i><i><u>of</u></i><i><u> </u></i><i><u>c</u></i><i><u> </u># dataset instances </i><i>of</i><i> </i><i>c</i><i> <u>2precisioncrecallc</u>precisionc-\-recallc</i></p><p>A baseline, consisting of the assignment of the most frequent category (always literal), was used for each task and granularity level.</p></section><section number="5" title="Systems and Results"><p>We received five submissions (FUH, GYDER, up13,  UTD-HLT-CG,  XRCE-M).   All tackled the location task; three (GYDER, UTD-HLT-CG, XRCE-M) also participated in the organisation task. All systems were full submissions (coverage of 1) and participated at all granularity levels.</p><subsection number="5.1" title="Methods and Features"><p>Out of five teams, four (FUH, GYDER, up13, UTD-HLT-CG) used supervised machine learning, including single (FUH,GYDER, up13 ) as well as multiple classifiers (UTD-HLT-CG). A range of learning paradigms was represented (including instance-based learning, maximum entropy, deci­sion trees, etc.). One participant (XRCE-M) built a hybrid system, combining a symbolic, supervised approach based on deep parsing with an unsuper-vised distributional approach exploiting lexical in­formation obtained from large corpora.</p><p>Systems up13 and FUH used mostly shallow fea­tures extracted directly from the training data (in­cluding parts-of-speech, co-occurrences and collo-</p><p><i><u># correct predictions</u></i> <i><u># predictions</u></i> <i># predictions</i> <i># samples</i><page local="5" global="40"/></p><table class="main" frame="box" rules="all" border="0" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>reading</p></td><td class="cell"><p>train</p></td><td class="cell"><p>test</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>literal</p></td><td class="cell"><p>737</p></td><td class="cell"><p>721</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>mixed</p></td><td class="cell"><p>15</p></td><td class="cell"><p>20</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>othermet</p></td><td class="cell"><p>9</p></td><td class="cell"><p>11</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>obj-for-name</p></td><td class="cell"><p>0</p></td><td class="cell"><p>4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>obj-for-representation</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>place-for-people</p></td><td class="cell"><p>161</p></td><td class="cell"><p>141</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>place-for-event</p></td><td class="cell"><p>3</p></td><td class="cell"><p>10</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>place-for-product</p></td><td class="cell"><p>0</p></td><td class="cell"><p>1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>total</p></td><td class="cell"><p>925</p></td><td class="cell"><p>908</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>2: Reading distribution for organis</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>reading</p></td><td class="cell"><p>train</p></td><td class="cell"><p>test</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>literal</p></td><td class="cell"><p>690</p></td><td class="cell"><p>520</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>mixed</p></td><td class="cell"><p>59</p></td><td class="cell"><p>60</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>othermet</p></td><td class="cell"><p>14</p></td><td class="cell"><p>8</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>obj-for-name</p></td><td class="cell"><p>8</p></td><td class="cell"><p>6</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>obj-for-representation</p></td><td class="cell"><p>1</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-members</p></td><td class="cell"><p>220</p></td><td class="cell"><p>161</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-event</p></td><td class="cell"><p>2</p></td><td class="cell"><p>1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-product</p></td><td class="cell"><p>74</p></td><td class="cell"><p>67</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-facility</p></td><td class="cell"><p>15</p></td><td class="cell"><p>16</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-index</p></td><td class="cell"><p>7</p></td><td class="cell"><p>3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>total</p></td><td class="cell"><p>1090</p></td><td class="cell"><p>842</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>cations). The other systems made also use of syn­tactic/grammatical features (syntactic roles, deter­mination, morphology etc.). Two of them (GYDER and UTD-HLT-CG) exploited the manually anno­tated grammatical roles provided by the organisers.</p><p>All systems apart from up13 made use of exter­nal knowledge resources such as lexical databases for feature generalisation (WordNet, FrameNet, VerbNet, Levin verb classes) as well as other cor­pora (the Mascara corpus for additional training ma­terial, the BNC, and the Web).</p></subsection><subsection number="5.2" title="Performance"><p>Tables 3 and 4 report accuracy for all systems.<footnote anchor="6"/> Ta­ble 5 provides a summary of the results with lowest, highest, and average accuracy and f-scores for each subtask and granularity level.<footnote anchor="7"/></p><p>The task seemed extremely difficult, with 2 of the 5 systems (up13 , FUH) participating in the location task not beating the baseline. These two systems re­lied mainly on shallow features with limited or no use of external resources, thus suggesting that these features might only be of limited use for identify­ing metonymic shifts. The organisers themselves have come to similar conclusions in their own ex­periments (Markert and Nissim, 2002). The sys­tems using syntactic/grammatical features (GYDER, UTD - HLT - CG, XRCE - M) could improve over the baseline whether using manual annotation or pars­ing. These systems also made heavy use of feature generalisation. Classification granularity had only a small effect on system performance.</p><p>Only few of the fine-grained categories could be distinguished with reasonable success (see the f-scores in Table 5). These include literal readings, and place-for-people, org-for-members, and org-for-product metonymies, which are the most frequent categories (see Tables 1 and 2). Rarer metonymic targets were either not assigned by the systems at all ("undef ' in Table 5) or assigned wrongly <u>Table 5 :</u><u> Overview of scores</u> (low f-scores). An exception is the object-for-name pattern, which XRCE-M and UTD-HLT-CG could distinguish with good success. Mixed read­ings also proved problematic since more than one pattern is involved, thus limiting the possibilities of learning from a single training instance. Only GYDER succeeded in correctly identifiying a variety of mixed readings in the organisation subtask. No systems could identify unconventional metonymies correctly. Such poor performance is due to the non-regularity of the reading by definition, so that ap­proaches based on learning from similar examples alone cannot work too well.</p><footnote label="6">Due to space limitations we do not report precision, recall, and f-score per class and refer the reader to each system de­scription provided within this volume.</footnote><footnote label="7">The value "undef" is used for cases where the system did not attempt any assignment for a given class, whereas the value "0" signals that assignments were done, but were not correct.</footnote><footnote label="8">Please note that results for the FUH system are slightly dif­ferent than those presented in the FUH system description pa­per. This is due to a preprocessing problem in the FUH system that was fixed only after the run submission deadline.</footnote><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>base</p></td><td class="cell"><p>min</p></td><td class="cell"><p>max</p></td><td class="cell"><p>ave</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>LOCATION-coarse</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>accuracy</p></td><td class="cell"><p>0.794</p></td><td class="cell"><p>0.754</p></td><td class="cell"><p>0.852</p></td><td class="cell"><p>0.815</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>literal-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.849</p></td><td class="cell"><p>0.912</p></td><td class="cell"><p>0.888</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>non-literal-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.344</p></td><td class="cell"><p>0.576</p></td><td class="cell"><p>0.472</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>LOCATION-medium</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>accuracy</p></td><td class="cell"><p>0.794</p></td><td class="cell"><p>0.750</p></td><td class="cell"><p>0.848</p></td><td class="cell"><p>0.812</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>literal-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.849</p></td><td class="cell"><p>0.912</p></td><td class="cell"><p>0.889</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>metonymic-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.331</p></td><td class="cell"><p>0.580</p></td><td class="cell"><p>0.476</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>mixed-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>0.083</p></td><td class="cell"><p>0.017</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>LOCATION-fine</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>accuracy</p></td><td class="cell"><p>0.794</p></td><td class="cell"><p>0.741</p></td><td class="cell"><p>0.844</p></td><td class="cell"><p>0.801</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>literal-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.849</p></td><td class="cell"><p>0.912</p></td><td class="cell"><p>0.887</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>place-for-people-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.308</p></td><td class="cell"><p>0.589</p></td><td class="cell"><p>0.456</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>place-for-event-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>0.167</p></td><td class="cell"><p>0.033</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>place-for-product-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>undef</p></td><td class="cell"><p>0.000</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>obj-for-name-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>0.667</p></td><td class="cell"><p>0.133</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>obj-for-rep-f</p></td><td class="cell"><p></p></td><td class="cell"><p>undef</p></td><td class="cell"><p>undef</p></td><td class="cell"><p>undef</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>othermet-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>undef</p></td><td class="cell"><p>0.000</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>mixed-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>0.083</p></td><td class="cell"><p>0.017</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ORGANISATION-coarse</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>accuracy</p></td><td class="cell"><p>0.618</p></td><td class="cell"><p>0.732</p></td><td class="cell"><p>0.767</p></td><td class="cell"><p>0.746</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>literal-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.800</p></td><td class="cell"><p>0.825</p></td><td class="cell"><p>0.810</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>non-literal-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.572</p></td><td class="cell"><p>0.652</p></td><td class="cell"><p>0.615</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ORGANISATION-medium</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>accuracy</p></td><td class="cell"><p>0.618</p></td><td class="cell"><p>0.711</p></td><td class="cell"><p>0.733</p></td><td class="cell"><p>0.718</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>literal-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.804</p></td><td class="cell"><p>0.825</p></td><td class="cell"><p>0.814</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>metonymic-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.553</p></td><td class="cell"><p>0.604</p></td><td class="cell"><p>0.577</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>mixed-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>0.308</p></td><td class="cell"><p>0.163</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ORGANISATION-fine</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>accuracy</p></td><td class="cell"><p>0.618</p></td><td class="cell"><p>0.700</p></td><td class="cell"><p>0.728</p></td><td class="cell"><p>0.713</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>literal-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.808</p></td><td class="cell"><p>0.826</p></td><td class="cell"><p>0.817</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-members-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.568</p></td><td class="cell"><p>0.630</p></td><td class="cell"><p>0.608</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-event-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>undef</p></td><td class="cell"><p>0.000</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-product-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.400</p></td><td class="cell"><p>0.500</p></td><td class="cell"><p>0.458</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-facility-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>0.222</p></td><td class="cell"><p>0.141</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>org-for-index-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>undef</p></td><td class="cell"><p>0.000</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>obj-for-name-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.250</p></td><td class="cell"><p>0.800</p></td><td class="cell"><p>0.592</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>obj-for-rep-f</p></td><td class="cell"><p></p></td><td class="cell"><p>undef</p></td><td class="cell"><p>undef</p></td><td class="cell"><p>undef</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>othermet-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>undef</p></td><td class="cell"><p>0.000</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>mixed-f</p></td><td class="cell"><p></p></td><td class="cell"><p>0.000</p></td><td class="cell"><p>0.343</p></td><td class="cell"><p>0.135</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="6" global="41"/><table caption="Table 3: Accuracy scores for all systems for all the location tasks.8"></table></subsection></section><section number="6" title="Concluding Remarks"><p>There is a wide range of opportunities for future fig­urative language resolution tasks. In the SemEval corpus the reading distribution mirrored the actual distribution in the original corpus (BNC). Although realistic, this led to little training data for several phenomena. A future option, geared entirely to­wards system improvement, would be to use a strat­ified corpus, built with different acquisition strate­gies like active learning or specialised search proce­dures. There are also several options for expand­ing the scope of the task, for example to a wider range of semantic classes, from proper names to common nouns, and from lexical samples to an all-words task. In addition, our task currently covers only metonymies and could be extended to other kinds of figurative language.</p></section><section title="Acknowledgements"><p>We are very grateful to the BNC Consortium for let­ting us use and distribute samples from the British National Corpus, version 1.0.</p><table caption="Table 4: Accuracy scores for all systems for all the organisation tasks" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>task <i>I </i><i>1 </i>system —&gt;</p></td><td class="cell"><p>baseline</p></td><td class="cell"><p>FUH</p></td><td class="cell"><p>UTD-HLT-CG</p></td><td class="cell"><p>XRCE-M</p></td><td class="cell"><p>GYDER</p></td><td class="cell"><p>upl3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>LOCATION-coarse</p></td><td class="cell"><p>0.794</p></td><td class="cell"><p>0.778</p></td><td class="cell"><p>0.841</p></td><td class="cell"><p>0.851</p></td><td class="cell"><p>0.852</p></td><td class="cell"><p>0.754</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>LOCATION-medium</p></td><td class="cell"><p>0.794</p></td><td class="cell"><p>0.772</p></td><td class="cell"><p>0.840</p></td><td class="cell"><p>0.848</p></td><td class="cell"><p>0.848</p></td><td class="cell"><p>0.750</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>LOCATION-fine</p></td><td class="cell"><p>0.794</p></td><td class="cell"><p>0.759</p></td><td class="cell"><p>0.822</p></td><td class="cell"><p>0.841</p></td><td class="cell"><p>0.844</p></td><td class="cell"><p>0.741</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>task <i>I </i><i>1 </i>system —&gt;</p></td><td class="cell"><p>baseline</p></td><td class="cell"><p>UTD-HLT-CG</p></td><td class="cell"><p>XRCE-M</p></td><td class="cell"><p>GYDER</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ORGANISATION-coarse</p></td><td class="cell"><p>0.618</p></td><td class="cell"><p>0.739</p></td><td class="cell"><p>0.732</p></td><td class="cell"><p>0.767</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ORGANISATION-medium</p></td><td class="cell"><p>0.618</p></td><td class="cell"><p>0.711</p></td><td class="cell"><p>0.711</p></td><td class="cell"><p>0.733</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ORGANISATION-fine</p></td><td class="cell"><p>0.618</p></td><td class="cell"><p>0.711</p></td><td class="cell"><p>0.700</p></td><td class="cell"><p>0.728</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></section><references><p>J.A. Barnden, S.R. Glasbey, M.G. Lee, and A.M. Walling­ton. 2003. Domain-transcending mappings in a system for metaphorical reasoning. In <i>Proc. ofEACL-2003, </i>57-61.</p><p>J. Birke and A Sarkaar. 2006. A clustering approach for the nearly unsupervised recognition of nonliteral language. In <i>Proc. ofEACL-2006.</i></p><p>L. Burnard, 1995. <i>Users' Reference Guide, British National Corpus. </i>BNC Consortium, Oxford, England.</p><p>J. Carletta. 1996. Assessing agreement on classification tasks: The kappa statistic. <i>Computational Linguistics, </i>22:249-254.</p><p>D. Fass. 1997. <i>Processing Metaphor and Metonymy. </i>Ablex,</p><p>Stanford, CA.</p><p>S. Harabagiu. 1998. Deriving metonymic coercions from WordNet. In <i>Workshop on the Usage ofWordNet in Natural Language Processing Systems, COLING-ACL '98, </i>142-148, Montreal, Canada.</p><p>J.R. Hobbs, M.E. Stickel, D.E. Appelt, and P. Martin. 1993. Interpretation as abduction. <i>Artificial Intelligence, </i>63:69­142.</p><p>S. Kamei and T. Wakao. 1992. Metonymy: Reassessment, sur­vey of acceptability and its treatment in machine translation systems. In <i>Proc. ofACL-92, </i>309-311.</p><p>S. Krishnakamuran and X. Zhu. 2007. Hunting elusive metaphors using lexical resources. In <i>NAACL 2007 Work­shop on Computational Approaches to Figurative Language.</i></p><p>G. Lakoff and M. Johnson. 1980. <i>Metaphors We Live By. </i>Chicago University Press, Chicago, Ill.</p><p>J. Leveling and S. Hartrumpf. 2006. On metonymy recogni­tion for gir. In <i>Proceedings of GIR-2006: 3rd Workshop on Geographical Information Retrieval.</i></p><p>K. Markert and U. Hahn. 2002. Understanding metonymies in discourse. <i>Artificial Intelligence, </i>135(1/2): 145—198.</p><p>K. Markert and M. Nissim. 2002. Metonymy resolution as a classification task. In <i>Proc. of EMNLP-2002,</i>204-213.</p><p>K. Markert and M. Nissim. 2006. Metonymic proper names: A corpus-based account. In A. Stefanowitsch, editor, <i>Corpora in Cognitive Linguistics. Vol. 1: Metaphor and Metonymy. </i>Mouton de Gruyter, 2006.</p><p>J. Martin. 1994. Metabank: a knowledge base of metaphoric language conventions. <i>Computational Intelli­gence, </i>10(2):134-149.</p><p>Z. Mason. 2004. Cormet: A computational corpus-based con­ventional metaphor extraction system. <i>Computational Lin­guistics, </i>30(1):23-44.</p><p>M. Nissim and K. Markert. 2003. Syntactic features and word similarity for supervised metonymy resolution. In <i>Proc. of</i> <i>ACL-2003, </i>56-63.</p><p>G. Nunberg. 1995. Transfers of meaning. <i>Journal of Seman­tics, </i>12:109-132.</p><p>Y Peirsman. 2006. Example-based metonymy recognition for proper nouns. In <i>Student Session ofEACL 2006.</i></p><p>D. Stallard. 1993. Two kinds of metonymy. In <i>Proc.ofACL-</i> <i>93, </i>87-94.</p></references></body></article>