<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="37"/><title>SENSEVAL-2 Japanese Translation Task</title><author surname="Kurohashi" givenname="Sadao"><org  name="University of Tokyo" country="Japan" city="Tokyo"/></author></firstpageheader><frontmatter><p>SENSEVAL-2 Japanese Translation Task</p><p><b>Sadao Kurohashi</b></p><p>University of Tokyo kuroOkc.t.u-tokyo.ac.jp</p></frontmatter><abstract>This paper reports an overview of <b>Senseval</b>-2 Japanese translation task. In this task, word senses are defined according to translation dis­tinction. A translation Memory (TM) was constructed, which contains, for each Japanese head word, a list of typical Japanese expressions and their English translations. For each target word instance, a TM record best approximating that usage had to be submitted. Alternatively, submission could take the form of actual target word translations. 9 systems from 7 organiza­tions participated in the task. </abstract></header><body><section number="1" title="Introduction"><p>In written texts, words which have multiple senses can be classified into two categories; homonyms and polysemous words. Generally speaking, while homonymy sense distinction is quite clear, polysemy sense distinction is very subtle and hard. English texts contain many homonyms. On the other hand, Japanese texts in which most content words are written by ideograms rarely contain homonyms. That is, the main target in Japanese WSD is polysemy, which makes Japanese WSD task setup very hard. What sense distinction of polysemous words is reasonable and effective heavily de­pends on how to use it, that is, an application of WSD.</p><p>Considering such a situation, in addition to the ordinary dictionary task we organized an­other task for Japanese, a translation task, in which word sense is defined according to trans­lation distinction. Here, we set up the task as­suming the example-based machine translation paradigm (Nagao, 1981). That is, first, a trans­lation memory (TM) is constructed which con­tains, for each Japanese head word, a list of typ­ical Japanese expressions (phrases/sentences) involving the head word and an English trans­lation for each (Figure 1). We call a pair of Japanese and English expressions in the TM as a TM record. Given an evaluation document containing a target word, participants have to submit the TM record best approximating that usage.</p><p>Alternatively, submissions can take the form of actual target word translations, or transla­tions of phrases or sentences including each tar­get word. This allows existing rule-based ma­chine translation (MT) systems to participate in the task, and we can compare TM based sys­tems with existing MT systems.</p><p>For evaluation, we distributed newspaper ar­ticles. The number of target words was 40, and 30 instances of each target word were provided, making for a total of 1,200 instances.</p></section><section number="2" title="Construction of Translation Memory"><p>The translation memory (TM) was constructed in two steps:</p><p>1. By referring to the KWIC (Key Word In Context) of a target word, its typical Japanese expressions are picked up by lex­icographers.</p><p>2. The Japanese expressions are translated by a translation company.</p><p>KWIC was made from the nine years vol­ume of Mainichi Newspaper corpus. They are morphologically analyzed and segmented into phrase sequences, and then the 100 most fre­quent phrase uni-grams, bi-grams (two types; the target word is in the first phrase or the sec­ond phrase) and tri-grams (the target word is in the middle phrase) are provided to lexicogra­phers (Figure 2).</p><page local="2" global="38"/><doubt alpha="85.7" length="7" tooSmall="False" monospace="0.0">MM muri</doubt><p>The lexicographers pick up a typical expres­sion of the target word from the KWIC. If its sense is context-independently clear, the expres­sion is adopted as it is. If its sense is not clear, some pre/post expressions are supplemented by referring original sentences in the newspaper corpus.</p><p>Then, we asked a translation company to translate the Japanese expressions. As a re­sult, a TM containing 320 head words and 6920 records was constructed (one head word has 21.6 records on average). The average number of words of a Japanese expression is 4.5.</p></section><section number="3" title="Gold Standard Data and the Evaluation of Translations"><p>As a gold standard data of the task, 40 target words were chosen out of 320 TM words. Con­sidering the possible comparison of the trans­lation task and the dictionary task, 40 target words were fully overlapped with 100 target words of the dictionary task.</p><p>In the Japanese dictionary task, target words are classified into three categories according to the difficulty (difficult, intermediate, easy), based on the entropy of word sense distri­bution in the training data of the dictionary task(Shirai, 2001). 40 target words of the tranj lation task consists of 20 nouns and 20 verbs: difficult nouns and verbs, 10 intermediate nom and verbs, and 5 easy nouns and verbs.</p><p>For each target word, 30 instances were ch&lt; sen from Mainichi Newspaper corpus (in tot. 1,200 instances) and they are also overlappe with the dictionary task. Since the dictionai task uses 100 instances for each target wor&lt; the translation task used 1st, 4th, 7th, ... 90t instances of the dictionary task.</p><p>As a gold standard data, zero or more a{ propriate TM records were assigned to each ii stance by the same translation company Aj propriate TM records were classified into th following three classes:</p><p>© : A TM record which can be used t translate the instance. POS, tense, plura singular, and subtle nuance do not <b><i>necei </i></b>sarily match.</p><p>O : If the instance is considered alone, th English translation is correct, but usin the TM record in the given context is nc so good, for example, making very round about translation.</p><table caption="Figure 1: An example of Translation Memory." class="main" frame="box" rules="all" border="0" regular="True"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>It is impossible to participate.</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>^ z mmmmm </i>\mmtc</p></td><td class="cell"><p>It is impossible to make use of the library in this hour.</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>This bill is hard to pass.</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>It is no wonder he got angry.</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>the most natural way</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>to work too much</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>unreasonable demand</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>passing by force</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>to commit a forced double suicide</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Figure 2: An example of KWIC (numbers indicate phrase frequency)." class="main" frame="box" rules="all" border="0" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Phrase</p></td><td class="cell"><p>uni-gram</p></td><td class="cell"><p>Phrase bi-gram</p></td><td class="cell"><p>Phrase tri-gram</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>597 *</p></td><td class="cell"><p></p></td><td class="cell"><p>151 <b><i>mï</i></b></p></td><td class="cell"><p></p></td><td class="cell"><p>i9 sitfctt <b>*a#</b></p></td><td class="cell"><p>7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>551 *</p></td><td class="cell"><p></p></td><td class="cell"><p>138 <b><i>Mi</i></b></p></td><td class="cell"><p></p></td><td class="cell"><p>14 <b><i>tXh</i></b><b><i> MM*</i></b></p></td><td class="cell"><p>6 jfcJÖäCDÖ <b>&amp;3Ät</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>416 *</p></td><td class="cell"><p><b>sa^o</b></p></td><td class="cell"><p>106 <b><i>Mi</i></b></p></td><td class="cell"><p></p></td><td class="cell"><p>i3 üttt <b>Mat</b></p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>413 *</p></td><td class="cell"><p><b><i>m</i></b><i>\z</i></p></td><td class="cell"><p>101 <b><i>Mi</i></b></p></td><td class="cell"><p>i &amp;&lt;</p></td><td class="cell"><p>10 <i>$iiö%&lt;D\Z </i><b><i>MMifi</i></b></p></td><td class="cell"><p><b>5 </b>i^&lt;cDfe <b>naa </b>&amp;H.</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>403 *</p></td><td class="cell"><p><b>S3*</b></p></td><td class="cell"><p>67 <b><i>Mi</i></b></p></td><td class="cell"><p>ICD &amp;V&gt;</p></td><td class="cell"><p>10 tTfe M3J t</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>351 <b><i>K</i></b></p></td><td class="cell"><p></p></td><td class="cell"><p>56 <b><i>Mi</i></b></p></td><td class="cell"><p></p></td><td class="cell"><p>9 V^50DÖ</p></td><td class="cell"><p>4lTfc <b>II« </b>&amp;V&gt;.</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="3" global="39"/><p><b>A </b>: If the instance is considered alone, the English translation is correct, but using the TM record in the given context is inappro­priate.</p><p>Out of 1,200 instances, 34 instances (2.8%) were assigned no TM records (there was no ap­propriate TM record). To one instance, on aver­age, 6.6 records were assigned as ©, 1.4 records as O, and 0.1 records as <b>A, </b>in total 8.1 records. If a system chooses a TM record randomly as an answer, the accuracy becomes 36.8% in case that all of ©, O and <b>A </b>records are regarded as correct, and 29.0% in case that only © is re­garded as correct (they are the baseline scores used in the next section).</p><p>In the gold standard data construction, 90 instances (9 words x 10 instances) were dealt with by two annotators doubly, and then their agreement were checked. For each instance one record is chosen randomly from annotator B's answers, and it was checked whether it is con­tained in annotator A's answers (annotator A made the whole gold standard data). The agree­ment was 86.6% in case that all of ©, O and <b>A </b>records are regarded as correct, and 80.9% in case that only © is regarded as correct.</p><p>In the case that the submission is in the form of translation data, translation experts (the same company as constructed the TM and the gold standard data) were asked to rank the supplied translation ©, O or X. This evalua­tion does not pay attention to the total transla­tion, but just the appropriateness of the target instance translation.</p></section><section number="4" title="Result"><p>In the Japanese translation task, 9 systems from 7 organizations submitted the answers. The characteristics of the systems are summarized as follows:</p><p>• AnonymX, Anonym Y Commercial, rule-based MT systems.</p><p>• CRL-NYU (Communications Research Laboratory &amp; New York Univ.)</p><p>TM records are classified according to the English head word, and each cluster is supplemented by several corpora. The system returns a TM record when the similarity between a TM record and an input sentence is very high. Otherwise, it returns the English head word of the most similar cluster by using several machine learning techniques.</p><p>• Ibaraki (Ibaraki Univ.)</p><p>A training data was constructed manually from newspaper articles, 170 instances for each target word. Features were collected in 7-word window around the target word, and decision list method was used for learn­ing.</p><p>• Stanford-Titechl (Stanford Univ. &amp; Tokyo Institute of Technology)</p><p>The system selects the appropriate TM record based on the character-bigram-based Dice's coefficient. It also utilized the context of the other target word instances in the evaluation text.</p><p>• AnonymZ</p><p>A sentence (TM records for learning, and an input for testing) is morphologically an­alyzed and converted into a semantic tag sequence, and maximum entropy method was used for learning.</p><doubt alpha="60.0" length="5" tooSmall="False" monospace="0.0">• ATR</doubt><p>The system selects the most similar TM record based on the cosine similarity be­tween context vectors, which were con­structed from semantic features and syn­tactic relations of neighboring words of the target word.</p><doubt alpha="66.7" length="21" tooSmall="False" monospace="0.0">• Kyoto (Kyoto Univ.)</doubt><p>The system selects the most similar TM record by bottom-up, shared-memory based matching algorithm.</p><p>• Stanford-Titech2 (Stanford Univ. &amp; Tokyo Institute of Technology)</p><p>The system selects the appropriate TM record based on the case-frame-based sim­ilarity, using NTT Goi-Taikei thesaurus.</p><p>The results of all systems are shown in Fig­ure 3. The left bar charts indicate the accuracy based on the lenient evaluation (©, O and <b>A </b>in TM selection and © and O in MT are re­garded as correct); the right bar charts indicate the accuracy based on the strict evaluation (© is only regarded as correct both in TM selection and MT). Note that since the TM does not have a hierarchical structure, there is no evaluation options such as fine, coarse, and mixed.<page local="4" global="40"/></p><doubt alpha="87.2" length="39" tooSmall="True" monospace="0.0">[□Lenient evaluationMStrict évaluation]</doubt><doubt alpha="66.7" length="3" tooSmall="False" monospace="0.0">«US</doubt><doubt alpha="0.0" length="11" tooSmall="False" monospace="0.0">'///'//''//</doubt><doubt alpha="26.7" length="15" tooSmall="False" monospace="0.0">J^ ^ .&lt;SP*&lt;ç» /</doubt><doubt alpha="7.7" length="26" tooSmall="False" monospace="0.0">* jr   £   &gt;   ^-   .&lt;-° *</doubt><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">/✓</doubt><figure caption="Figure 3: Result of the Japanese translation task."></figure><doubt alpha="13.8" length="29" tooSmall="False" monospace="0.0">^ ^ ^ ^ ^ ^ J" ^ ///*r+*&lt;f+ y</doubt><doubt alpha="66.7" length="3" tooSmall="False" monospace="0.0">yy+</doubt><figure caption="Figure 4: Scores for nouns and verbs."></figure><p>Figure 4 shows scores for nouns and verbs separately, and Figure 5 shows scores for dif­ficult/intermediate/easy words. Both of them were evaluated by the lenient criteria.</p><p>In these figures, "Agreement" and "Baseline" were as described in the previous section. When the system judges that there is no appropri­ate TM record for an instance, it can return "UNASSIGNABLE". In that case, if there is no appropriate TM record assigned in the gold standard data, the answer is regarded as cor­rect.</p><p>Among TM selection systems, systems using some extra learning data outperformed other systems just using the TM. The comparison be­tween TM selection systems and MT systems is not easy, but the result indicates the effec­tiveness of the accumulated know-how of MT systems. However, the performance of the best TM selection system is not so different from MT systems, which indicates the promising future ol TM based techniques.</p><figure caption="Figure 5: Scores for difficulty classes."></figure></section><section number="5" title="Conclusion"><p>This paper described an overview of <b>senseval-</b>2 Japanese translation task. The data used ir this task are available at <b>Senseval</b>-2 web site, We hope this valuable data helps improve WSE and MT systems.</p><p><b>Acknowledgment</b></p><p>I wish to express my gratitude to Mainichi Newspapers for providing articles. I would alsc like to thank Prof. Takenobu Tokunaga (Tokyc Institute of Technology) and Prof. Kiyoaki Shi-rai (JAIST) and Dr. Kiyotaka Uchimoto (CRL) for their valuable advise about task organiza­tion, Yuiko Igura (Kyoto Univ.) and Intel Group Corp. for data construction, and all par­ticipants to the task.</p></section><references><p>Makoto Nagao. 1981. A framework of mechan­ical translation between Japanease and En­glish by analogy priciple. In <i>Proc. of the In­ternational NATO Symposium on Artificia* and Human Intelligence.</i></p><p>Kiyoaki Shirai. 2001. SENSEVAL-2 Japanese dictionary task. In <i>Proceedings of the SENSEVAL-2 Workshop.</i></p></references></body></article>