<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>Comparing and Extracting Paraphrasing Words with 2-Way Bilingual Dictionaries.</title><author surname="Takao" givenname="Kazutaka"><org  name="BT Laboratories" country="United Kingdom" city="Ipswich"/></author><author surname="Imamura" givenname="Kenji"><org  name="ATR Spoken Language Translation Research Laboratories" country="Japan" city="Kyoto"/></author><author surname="Kashioka" givenname="Hideki"><org  name="ATR Spoken Language Translation Research Laboratories" country="Japan" city="Kyoto"/></author></firstpageheader><frontmatter><p><b>Comparing and Extracting Paraphrasing Words with 2-Way Bilingual</b></p><p><b>Dictionaries</b></p><p><b>Kazutaka Takao, Kenji Imamura, Hideki Kashioka</b></p><p>Advanced Telecommunications Research Institute Spoken Language Translation Research Laboratories 2-2-2 Hikaridai, Seika-cho, Soraku-gun, Kyoto 619-0288, Japan {kazutaka.takao, kenji.imamura, hideki.kashioka}@atr.co.jp</p></frontmatter><abstract>We analyze a variety of lexical expressions with 2-way bilingual dictionaries and propose a method for extracting paraphrasing words. First, we compare the coverage between an English-Japanese dictionary and a Japanese-English dictionary from the viewpoint of the returnability of the words by translating English to Japanese, and then back to English again. The variety is shown using examples. Next, we propose a method of automatically extracting English paraphrasing word groups; we gathered the English index words which have the same Japanese translation words in the E-J dictionary. The English words which are difficult to distinguish for native speakers of Japanese were then extracted into a paraphrasing group. We also extract the Japanese paraphrasing word groups for comparison. This method will be useful for sentence matching, especially in order to accept the variety of expressions. </abstract></header><body><section number="1." title="Introduction"><p>In machine translation there is a problem with the variety of translation words. We conventionally made each entry of our translation dictionary a one-to-one entry because we were in a hurry to build machine translation systems in the early stages. Therefore, the translation words of the target language generated from a word of the source language tend to be uniform in our machine translation systems. In case no suitable English words are found in the Japanese-English dictionary when somebody translates a Japanese sentence into English by hand, he/she recalls some English words he/she knows, looks up the English-Japanese dictionary in reverse, and tries to find the best word for the sentence.</p><p>We are importing the electronic data from commercial bilingual dictionaries into our machine translation systems. However, simply importing the lexical entries in the same direction as the direction of the translation, i.e., to import the entries in a J-E dictionary for our J-E machine translation, does not sufficiently handle the variety of expressions.</p><p>Only a few studies can be found about extracting information from a commercial English-Japanese dictionary. For example, Shirai et al. (2001) constructed J­E entries by reversing E-J from the viewpoint of a valency pattern.</p><p>In this paper, we used the electronic data of commercial 2-way bilingual dictionaries, i.e., a Japanese-English dictionary and an English-Japanese dictionary published in Japan. However, they have subtle differences among their entries and translations. Therefore, from the viewpoint of increasing the variety of the words, we report on a method for comparing the coverage between these two bilingual dictionaries, and extracting paraphrasing word groups.</p></section><section number="2." title="Returnability of bi-directional translation"><subsection number="2.1." title="Size of dictionaries"><p>We extracted the lexical entries from the dictionaries. The numbers of the words are shown in Table 1.</p></subsection><subsection number="2.2." title="Method of matching"><p>We compared the coverage by checking the returnability of E-J-E matches, i.e., translating English to Japanese, and then back to English again (see Figure 1). We enumerated each Japanese translation word (J1) for the English index word (E1) in the English-Japanese dictionary, and looked up the corresponding entry in the Japanese-English dictionary. Then, we compared the English words of both end results (E1, E2), and analyzed whether E1 returns to E2 when we search for the same word as E1 among plural translations.</p><p>We also carried out the J-E-J match and made a comparison with E-J-E. However, it is redundant to describe the analysis about both E-J-E and J-E-J. So that, we mainly describe about E-J-E as long as we do not annotate specially.</p><doubt alpha="58.5" length="41" tooSmall="False" monospace="0.0">Dictionary 1 (E-J)     Dictionary 2 (J-E)</doubt><doubt alpha="25.0" length="8" tooSmall="False" monospace="0.0">E1 =&gt; J1</doubt><doubt alpha="50.0" length="4" tooSmall="False" monospace="0.0">A_r&gt;</doubt><doubt alpha="25.0" length="8" tooSmall="False" monospace="0.0">J1 =&gt; E2</doubt><doubt alpha="50.0" length="4" tooSmall="False" monospace="0.0">S_A_</doubt><p>Source word   Matching word   Terminal word<page local="2"/></p><figure caption="Figure 1: E-J-E matching"></figure><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1016</doubt><table caption="Table 1: Size of dictionaries" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dic.</p></td><td class="cell"><p>Num. of index words(A)</p></td><td class="cell"><p>Total num. of trans.(B)</p></td><td class="cell"><p>B/A</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>E-J</p></td><td class="cell"><p>46469</p></td><td class="cell"><p>141726</p></td><td class="cell"><p>3.0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>J-E</p></td><td class="cell"><p>28395</p></td><td class="cell"><p>45934</p></td><td class="cell"><p>1.6</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></subsection><subsection number="2.3." title="Result of returnability"><p>We classified the result of the correspondence into four types, as given below, and the number of each type is shown in Tables 2 and 3. Note that if several translation words are described per index word in dictionary 1, they are counted separately.</p><p>(a) returns exactly (entry J1 exists in J-E and E1 exists as a translation)</p><p>(b) returns to a part of the morphemes (entry J1 exists in J-E and E1+x exists as a translation)</p><p>(c) entry of the matching word exists in dictionary 2, but the source word and the terminal word differ (entry J1 exists in J-E but E1 and E2 differ)</p><p>(d) entry of the matching word does not exist in dictionary 2 (entry J1 does not exist in J-E)</p><p>Table 3 : Classification of J-E-J matching result The following sections show some examples of each type.</p><subsubsection number="2.3.1." title="Type (a): returns exactly"><table caption="Table 4 shows examples of type (a). Japanese words are written in italics."></table></subsubsection><subsubsection number="2.3.2." title="Type (b): returns to a part of the morphemes"><p>Table 5 shows examples of type (b). From this, we can see that some of the examples specialize the part-of-speech or the usage by adding a functional word, such as "of" or "na", and some of the examples specialize the meaning by adding another noun, such as "office" or <i>"gakarf. </i>For example, <i>"uketsukë" </i>means both the place and the person in charge. However, by adding "gakari", the meaning is specialized to the person in charge.</p></subsubsection><subsubsection number="2.3.3." title="Type (c): entry exists but differs"><p>We chose 100 entries at random from type (c), i.e., the case in which an entry of the matching word was found in dictionary 2 but the source word and the terminal word differ, compared these two words in each entry, and classified them into the causes of the difference as shown in Tables 6 and 7. Note that there were many subtle cases and it was difficult to classify them clearly.</p><p>From this, despite some differences such as synonyms or narrowed or widened meanings, in many cases the terminal word is another expression of nearly the same meaning as the source word. This suggests that the number of the translation words in dictionary 2 is too small to obtain the same expression as the source word in dictionary 1. Furthermore, if we use the source words, we can increase the variety of the expressions of the terminal words. Moreover, we can obtain the paraphrasing word groups by grouping them efficiently.</p><p>On the other hand, some entries need remarks that the two words have a slight difference in nuance. For example, although both "blunder" and "fail" correspond to the same word <i>"yarisokonau", </i>there is a subtle difference in that "blunder" means making a stupid mistake but "fail" means being unsuccessful in achieving something. To take another example, although both "decay" and "devastation" correspond to the same word <i>"kouhai", </i>there is a subtle difference in that "decay" means destruction by natural causes but "devastation" means great destruction. Therefore, even if some English words correspond to a Japanese word, we should be careful to choose them properly. Moreover, those words are often difficult to distinguish from each other for Japanese speakers, and it is fruitful to extract these word groups as words that Japanese speakers often confuse.</p><p>Yet on the other hand, in about a quarter of the entries the difference comes from a difference in their part-of-speech. For example, sneeringly =&gt; <i>reishou-shite </i>: <i>reishou-suru </i>=&gt; sneer at "Sneeringly" and "sneer at" were different because we inflected the Japanese translation word in dictionary 1 (E-J) into the form that appears as the index word in dictionary 2 (J-E).</p></subsubsection><subsubsection number="2.3.4." title="Type (d): entry does not exist"><p>We chose 100 pairs at random from type (d), and classified them according to the reason for the miss by hand-checking (see Table 8). In E-J-E matching, we found that about half of them were a pattern of a Japanese compound word corresponding to one English word. By inverting the English and Japanese words in the E-J dictionary into the J-E translation pair, the lack of coverage in the J-E dictionary could be effectively supplemented. This means that we can achieve more natural English generation if we can use English specific words.<page local="3"/></p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1017</doubt><table caption="Table 5: Examples of type (b)" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>E-J-E</p></td><td class="cell"><p>branch =&gt; <i>shiten </i>=&gt; branch office amount =&gt; <i>sougaku </i>=&gt; total amount afraid =&gt; <i>osoreru </i>=&gt; be afraid of</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>J-E-J</p></td><td class="cell"><p><i>anzen </i>=&gt; safe =&gt; <i>anzen-na</i></p><p><i>uketsuke </i>=&gt; receptionist =&gt; <i>uketsuke-gakari</i></p><p><i>kigen </i>=&gt; deadline =&gt; <i>saishuu-kigen</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Type</p></td><td class="cell"><p>#(Entry)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>(a) E1=&gt;J1 : J1=&gt;E1</p></td><td class="cell"><p>14585 ( 10.3%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>(b) E1=&gt;J1 : J1=&gt;E1+x</p></td><td class="cell"><p>1664 ( 1.2%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>(c) E1=&gt;J1 : J1=&gt;E2</p></td><td class="cell"><p>52462 ( 37.0%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>(d) E1=&gt;J1 : x</p></td><td class="cell"><p>73015 ( 51.5%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Total</p></td><td class="cell"><p>141726 (100.0%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Table 2: Classification of E-J-E matching result</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Type</p></td><td class="cell"><p>#(Entry)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>(a) J1=&gt;E1 : E1=&gt;J1</p></td><td class="cell"><p>12717 ( 27.7%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>(b) J1=&gt;E1 : E1=&gt;J1+x</p></td><td class="cell"><p>2589 ( 5.6%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>(c) J1=&gt;E1 : E1=&gt;J2</p></td><td class="cell"><p>14088 ( 30.7%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>(d) J1=&gt;E1 : x</p></td><td class="cell"><p>16540 ( 36.0%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Total</p></td><td class="cell"><p>45934 (100.0%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 4: Examples of type (a)" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>E-J-E</p></td><td class="cell"><p>abandon =&gt; <i>houki-suru </i>=&gt; abandon lemon =&gt; <i>remon </i>=&gt; lemon reconstruction =&gt; <i>saiken </i>=&gt; reconstruction</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>J-E-J</p></td><td class="cell"><p><i>aisukunmu </i>=&gt; ice cream =&gt; <i>aisukunmu omoi </i>=&gt; heavy =&gt; <i>omoi atarashii </i>=&gt; fresh =&gt; <i>atarashii</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>We can find the same information in J-E-J matching (see Table 9). We can achieve more natural Japanese generation if we can use Japanese characteristic words, such as mimetic words (e.g., <i>garagara).</i><i></i></p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1018</doubt><table caption="Table 6: Analysis and examples of type (c) in E-J-E" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Class</p></td><td class="cell"><p>#</p></td><td class="cell"><p>Example</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Synonym</p></td><td class="cell"><p>37</p></td><td class="cell"><p>bluff =&gt; <i>zeppeki </i>=&gt; cliff help =&gt; <i>kyuujo </i>=&gt; rescue</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Part-of-speech differs</p></td><td class="cell"><p>26</p></td><td class="cell"><p>restful =&gt; <i>ochitsuita </i>=&gt; feel at ease sneeringly =&gt; <i>reishoushite </i>=&gt; sneer at</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Fine shade of meaning</p></td><td class="cell"><p>9</p></td><td class="cell"><p>blunder =&gt; <i>yarisokonau </i>=&gt; fail decay =&gt; <i>kouhai </i>=&gt; devastation</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Narrowed</p></td><td class="cell"><p>7</p></td><td class="cell"><p>past =&gt; <i>keireki </i>=&gt; record mate =&gt; <i>haiguusha </i>=&gt; spouse</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Lexical coverage</p></td><td class="cell"><p>5</p></td><td class="cell"><p>inner man =&gt; <i>ibukuro </i>=&gt; stomach nick =&gt; <i>keimusho </i>=&gt; prison</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Notation differs</p></td><td class="cell"><p>4</p></td><td class="cell"><p>make-up =&gt; <i>mëkyappu </i>=&gt; makeup</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dic. 2 gives example(s) only</p></td><td class="cell"><p>4</p></td><td class="cell"><p>fix =&gt; <i>kyuuchi </i>=&gt; drive into a tight corner</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Widened</p></td><td class="cell"><p>3</p></td><td class="cell"><p>write out =&gt; <i>nozoku </i>=&gt; remove</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Other</p></td><td class="cell"><p>5</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 7: Analysis and examples of type (c) in J-E-J" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Class</p></td><td class="cell"><p>#</p></td><td class="cell"><p>Example</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Synonym</p></td><td class="cell"><p>52</p></td><td class="cell"><p><i>atsukurushii </i>=&gt; stuffy =&gt; <i>kazetooshi-no-warui kazukazu </i>=&gt; lot of =&gt; <i>takusan ryouyou </i>=&gt; recuperation =&gt; <i>seiyou</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Widened</p></td><td class="cell"><p>15</p></td><td class="cell"><p><i>ikioi </i>=&gt; power =&gt; <i>chikara ikkou </i>=&gt; group =&gt; <i>shuudan kaishuu-suru </i>=&gt; repair =&gt; <i>shuuri-suru</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Notation differs</p></td><td class="cell"><p>11</p></td><td class="cell"><p><i>shiboru </i>=&gt; wring =&gt; <i>shiboru </i>(ideogram/phonogram) <i>tatakau </i>=&gt; fight =&gt; <i>tatakau </i>(different ideogram)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Part-of-speech differs</p></td><td class="cell"><p>7</p></td><td class="cell"><p><i>ne </i>=&gt; naturally =&gt; <i>umaretsuki rinban </i>=&gt; alternately =&gt; <i>kougo-ni</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Narrowed</p></td><td class="cell"><p>6</p></td><td class="cell"><p><i>ude </i>=&gt; ability =&gt; <i>nouryoku ebi </i>=&gt; spiny lobster =&gt; <i>iseebi</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Fine shade of meaning</p></td><td class="cell"><p>6</p></td><td class="cell"><p><i>iken </i>=&gt; advice =&gt; <i>jogen</i></p><p><i>ki-no-nagai </i>=&gt; long-range =&gt; <i>choukyori-no</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Intransitive/transitive</p></td><td class="cell"><p>1</p></td><td class="cell"><p><i>eiten-suru </i>=&gt; promoted =&gt; <i>shoushin-saseru</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Lexical coverage</p></td><td class="cell"><p>1</p></td><td class="cell"><p><i>wase </i>=&gt; early =&gt; <i>hayai</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Other</p></td><td class="cell"><p>1</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 8: Investigation and examples of type (d) in E-J-E" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Class</p></td><td class="cell"><p></p></td><td class="cell"><p>Example</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Single word to</p></td><td class="cell"><p>48</p></td><td class="cell"><p>accommodation =&gt;</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>compound word</p></td><td class="cell"><p></p></td><td class="cell"><p><i>shuuyou-nouryoku </i>obvious =&gt; <i>sugu-wakaru</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Notation differs</p></td><td class="cell"><p>10</p></td><td class="cell"><p>disengage =&gt; <i>hazusu</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Suffix exists</p></td><td class="cell"><p>10</p></td><td class="cell"><p>contagiously =&gt;</p><p><i>densentekini</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Coverage (compound)</p></td><td class="cell"><p>9</p></td><td class="cell"><p>fuzzy logic =&gt; <i>fajï-riron</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Coverage (single)</p></td><td class="cell"><p>8</p></td><td class="cell"><p>formation =&gt; <i>taikei</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Transliteration</p></td><td class="cell"><p>6</p></td><td class="cell"><p>housesitter =&gt; <i>hausushittä</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Explanatory</p></td><td class="cell"><p>6</p></td><td class="cell"><p>bobble =&gt; <i>chiisana-kyuujouno-fusakazari</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Part-of-speech differs</p></td><td class="cell"><p>3</p></td><td class="cell"><p>delicacy =&gt; <i>surudosa</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 9: Investigation and examples of type (d) in J-E-J" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Class</p></td><td class="cell"><p>#</p></td><td class="cell"><p>Example</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Single word to compound word</p></td><td class="cell"><p>49</p></td><td class="cell"><p><i>garagara </i>=&gt; almost empty <i>narau </i>=&gt; take lessons <i>youshoku </i>=&gt; important post</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Coverage (compound)</p></td><td class="cell"><p>21</p></td><td class="cell"><p><i>tokubetsu-ryoukin =&gt; </i>extra charge</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Part of idiom</p></td><td class="cell"><p>9</p></td><td class="cell"><p><i>biryoku </i>=&gt; though I'm afraid won't much</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Explanatory</p></td><td class="cell"><p>7</p></td><td class="cell"><p><i>ekiben </i>=&gt; box lunch sold at a railroad station</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Coverage (single)</p></td><td class="cell"><p>4</p></td><td class="cell"><p><i>kahogo </i>=&gt; overprotection</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Transliteration</p></td><td class="cell"><p>4</p></td><td class="cell"><p><i>aki </i>=&gt; risshuu</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Part-of-speech differs</p></td><td class="cell"><p>2</p></td><td class="cell"><p><i>kariire </i>=&gt; harvesting</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Other</p></td><td class="cell"><p>4</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="4"/><doubt alpha="66.7" length="3" tooSmall="False" monospace="0.0">E-J</doubt><doubt alpha="66.7" length="3" tooSmall="False" monospace="0.0">J-E</doubt></subsubsection></subsection></section><section number="3." title="Extracting paraphrasing words"><subsection number="3.1." title="Variety of translation words"><p>Generally speaking, only a limited number of Japanese translation words are shown in the commercial English-Japanese dictionary which are sufficient to understand the English word. For this reason, we can only get a few words by simply gathering the Japanese words for the translation word for each index word.</p><p>The Japanese-English dictionary shows a more marked tendency of the limitation of translation words than the English-Japanese dictionary. As seen in Table 1, the J-E dictionary not only has considerably fewer index words, but also considerably fewer translation words per index word (B/A) than the E-J dictionary. Therefore, the number of English translation words in the J-E dictionary seems to be too small. This is derived from an editorial policy in which the translation words were reduced intentionally and the example sentences were increased in order to prompt the users to avoid literal English translations, because the dictionary was published for human-reading. However, it is not suitable for obtaining a large vocabulary of English translation words. For example, Shirai &amp; Yamamoto (2001) carried out matching between a Korean-English dictionary and a Japanese-English dictionary and reported a problem that different words are used to describe the same meaning, such as "occupation" and "business", and this caused unmatchings between some entries in J-E and K-E.</p><p>Thereupon we gained a large vocabulary in English by gathering the English index words out of the E-J dictionary, which can be grouped into paraphrasing word groups by matching the Japanese translation words. If we use the paraphrasing word groups, the problem in Shirai &amp; Yamamoto (2001) can be resolved and it is fruitful to apply them into other matchings, such as sentence-matching.</p></subsection><subsection number="3.2." title="Method for extracting paraphrasing words"><p>We created the English paraphrasing word groups by gathering and grouping the English index words in the E-J dictionary according to the existence of the same Japanese word as a translation word. Consequently, we could obtain a greater variety of English words than we could by simply gathering the English translation words in the J-E dictionary. For example, in E-J-E matching, as shown in Figure 2, in the case where the Japanese matching word is "<i>''shokugyou'"</i>" if we simply gather E in J-E, we get only four words: "occupation", "profession", "trade" and "job". However, by adding E in E-J, we can get more words: "business", "calling", "career", etc.</p><p>On the other hand, when the Japanese matching word has a wide meaning, incompatible English words are classified in one group. Therefore, we excluded such groups from extraction into paraphrasing words. We defined the judgment for excluding or not as follows: if the English translation description in the J-E dictionary is classified into plural subsections, we regard the Japanese index word as having a wide meaning. For example, the translation of <i>"saisef </i>in J-E dictionary is classified into three subsections:</p><p>1: (produce newly) grow again 2: (use again) recycle 3: (playback music etc.) playback</p><p>Therefore, we exclude the matching of the word <i>"saisei". </i>As a result, we can omit the group that includes incompatible words such as "rebirth"and "replay".</p><p>In the same way, we also extracted Japanese paraphrasing word groups by J-E-J matching.</p><p><i>shokugyou </i>=&gt; occupation, profession, trade, job</p><p><b>Paraphrasing words</b></p><p>playback =&gt; <i>saisei</i> rebirth =&gt; <i>saisei</i> regeneration =&gt; <i>saisei</i> replay =&gt; <i>saisei</i> reproduction =&gt; <i>saisei</i> <i>saisei </i>=&gt; 1:grow again <i>saisei </i>=&gt; 2:recycle <i>saisei </i>=&gt; 3:playback</p><doubt alpha="100.0" length="7" tooSmall="False" monospace="0.0">Exclude</doubt><figure caption="Figure 2: Method for extracting paraphrasing words4. Extracted result"></figure></subsection><subsection number="4.1." title="Extracted result"><p>Table 10 shows the amount of the extracted English paraphrasing words by E-J-E matching and Table 11 shows the amount of the extracted Japanese paraphrasing words by J-E-J matching. The number 6676 of the extracted groups in Table 10 seems quite large as against 1257 in Table 11. This is because the E-J dictionary has more entries than the J-E dictionary and because the entries in the E-J dictionary tend to be classified into more subsections than the J-E dictionary.</p><p><u>Num. of extracted groups</u>_</p><p>Num. of words per group (E in J-E) <u>Num.</u><u> of extracted groups</u></p><doubt alpha="62.2" length="45" tooSmall="False" monospace="0.0">Num. of words per group (E in J-E + E in E-J)</doubt><doubt alpha="0.0" length="7" tooSmall="False" monospace="0.0">66761.8</doubt><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">5.0</doubt><table caption="Table 10: Amount of extracted English paraphrasing words by E-J-E"></table><p>Num. of words per group (J in E-J)</p><doubt alpha="62.2" length="45" tooSmall="False" monospace="0.0">Num. of words per group (J in E-J + J in J-E)</doubt><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1257</doubt><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">2.2</doubt><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">3.9</doubt><p>Table 11 : Amount of extracted Japanese paraphrasing words by J-E-J<page local="5"/></p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1019</doubt><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>bisuness =</p></td><td class="cell"><p>&gt; <i>shokugyou</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>calling =</p></td><td class="cell"><p>&gt; <i>shokugyou</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>career =</p></td><td class="cell"><p>&gt; <i>shokugyou</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>craft =</p></td><td class="cell"><p>&gt; <i>shokugyou</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>occupation =</p></td><td class="cell"><p>&gt; <i>shokugyou</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>profession =</p></td><td class="cell"><p>&gt; <i>shokugyou</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></subsection><subsection number="4.2." title="Evaluation experiment"><p>Next, in order to see the quality of the extracted result, we chose 100 groups at random out of the English paraphrasing word groups and conducted an evaluation with a native speaker of English. In each group, we chose one English word (X) at random that came from J-E, chose one English word (Y) at random that entered into the group newly by adding words in E-J, and compared X and Y for meaning and part-of-speech. We evaluated the meaning of the X and Y pairs by classifying them into five types as follows:</p><p>A: The meanings of X and Y are exactly the same.</p><p>B: Y has a wider meaning than X.</p><p>C: Y has a narrower meaning than X.</p><p>D: The meanings of X and Y overlap partly.</p><p>E: Nonsense.</p><p>The evaluator was not familiar with the part-of-speech system. Therefore, we asked him/her to recall some sentences using X, replace X into Y, and judge whether the sentences become ungrammatical or not as an evaluation of part-of-speech. We said that he/she need not consider whether the meaning changes or not in evaluating part-of-speech.</p><p>O: Still grammatical.</p><p>X: Becomes ungrammatical.</p><p>?: Both cases are possible.</p><p>Note that we hid the Japanese matching words and let him/her evaluate only with English words.</p><p>In the same way, we also conducted an evaluation of Japanese paraphrasing word groups with a native speaker of Japanese.</p></subsection><subsection number="4.3." title="Result of evaluation"><p>Table 12 shows the evaluation of English paraphrasing words. Almost half of the pairs were classified into type A, and the sum of type A to D amounts to 80%. Ten Xes in type A seems to be a lot. This is because some of them had a different part-of-speech in English but were the same word in Japanese with a change in the inflection by adding functional words:</p><p>frivolous / levity : <i>keihaku(+na)</i></p><p>Yet some of them seem to be misjudgments because the evaluator could not recall all usages written in the dictionary when he/she only saw the English words. For example, the next example was evaluated as "AX" but the correct answer is "A?". This is because the evaluator could recall only "surrender" as a verb in spite of the existence of a noun in the dictionary:</p><p>surrender / capitulation : <i>koufuku</i></p><p>We also evaluated the excluded words (see the example of <i>"saisef </i>in Figure 2) in order to see the performance of the exclusion (see Table 13). There are 20 pairs of type A. This is because the pairs are not always incompatible even if they are excluded. In fact, the words in the excluded groups should be subdivided into further groups to get proper paraphrasing words. Therefore, if we can subdivide them in some way, we will obtain more paraphrasing words. However, we must devise something for subdivision because we cannot know which subsection in J-E the J in E-J corresponds to.</p><p>We also evaluated Japanese. Table 14 shows the evaluation of Japanese paraphrasing words. Table 15 shows the evaluation of the excluded words in order to see the performance of the exclusion.</p><doubt alpha="64.3" length="70" tooSmall="False" monospace="0.0">X / Y : matching word; where X is E in J-E, Y is a new E by adding E-J</doubt><table caption="Table 12: Evaluation of English paraphrasing words"></table><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1020</doubt><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Type</p></td><td class="cell"><p>#</p></td><td class="cell"><p>POS detail</p></td><td class="cell"><p>Example *</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A</p></td><td class="cell"><p>44</p></td><td class="cell"><p>(O=26 ?= 8 X=10)</p></td><td class="cell"><p>name / appellation : <i>namae</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>B</p></td><td class="cell"><p>5</p></td><td class="cell"><p>(O= 3 ?= 1 X= 1)</p></td><td class="cell"><p>stovepipe / stack : <i>entotsu</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>C</p></td><td class="cell"><p>5</p></td><td class="cell"><p>(O= 4 ?= 0 X= 1)</p></td><td class="cell"><p>take off / undo : <i>nugu</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>D</p></td><td class="cell"><p>26</p></td><td class="cell"><p>(O=17 ?= 3 X= 6)</p></td><td class="cell"><p>principle / religion : <i>shinjou</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>E</p></td><td class="cell"><p>20</p></td><td class="cell"><p>(O= 7 ?= 2 X=11)</p></td><td class="cell"><p>toilet / cloakroom : <i>toire</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Total</p></td><td class="cell"><p>100</p></td><td class="cell"><p>(O=57 ?=14 X=29)</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 13: Evaluation of English excluded words" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Type</p></td><td class="cell"><p>#</p></td><td class="cell"><p>POS detail</p></td><td class="cell"><p>Example</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A</p></td><td class="cell"><p>20</p></td><td class="cell"><p>(O=13 ?= 2 X= 5)</p></td><td class="cell"><p>party / reception : <i>pâtï</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>B</p></td><td class="cell"><p>9</p></td><td class="cell"><p>(O= 6 ?= 0 X= 3)</p></td><td class="cell"><p>object / mark : <i>mokuhyou</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>C</p></td><td class="cell"><p>11</p></td><td class="cell"><p>(O= 8 ?= 1 X= 2)</p></td><td class="cell"><p>image / picture : <i>imëji</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>D</p></td><td class="cell"><p>22</p></td><td class="cell"><p>(O= 7 ?= 3 X=12)</p></td><td class="cell"><p>stamp / emblem : <i>shirushi</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>E</p></td><td class="cell"><p>38</p></td><td class="cell"><p>(O=19 ?= 4 X=15)</p></td><td class="cell"><p>tap / chef : <i>kokku</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Total</p></td><td class="cell"><p>100</p></td><td class="cell"><p>(O=53 ?=10 X=37)</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="6"/><p>toilet / cloakroom : <i>toire</i></p></subsection></section><section number="5." title="Discussion"><subsection number="5.1." title="Investigation of type E"><p>There are still 20 pairs of type E in Table 12. We further investigated the dictionary entries of the word Y in each pair to see the cause, i.e., E-J of E-J-E matching.</p><subsubsection number="5.1.1." title="Uncommon usage"><p>Firstly, some pairs of type E were not errors but uncommon usages. The dictionary entries for four pairs were marked as informal slang. The dictionary entries for eight pairs were placed in a rearward subsection, which means that the usages are rare and uncommon. For example:</p><p>If we look up "cloakroom" in a dictionary, the meaning of toilet is surely described. However, if we were to ask for the cloakroom during a trip, we would be led to the checkroom rather than the toilet. Therefore, the evaluator evaluated it as type E. It can be seen that type E is not always an error. Moreover, the entry of one pair was a technical term, and the entry of another pair was a British term.</p></subsubsection><subsubsection number="5.1.2." title="Doubtful description in E-J dictionary"><p>Secondly, two pairs of type E seem to be misdescriptions in the E-J dictionary. For example:</p><p>syringe / squirt : <i>chuushaki</i></p><p>Although the meaning of "syringe" was found in the index of "squirt" in the E-J dictionary, it was not found in the E­E dictionary. Yet another example:</p><p>shore / margin : <i>kishi</i></p><p>The evaluator commented that he/she had never heard "margin"used in the meaning of "shore".</p></subsubsection><subsubsection number="5.1.3." title="Difference of culture"><p>Thirdly, some pairs resulted from a difference of culture. For example:</p><p>study / library : <i>shosai</i></p><p>Although there is one word in Japanese, the place for working ("study") and the place for keeping books ("library") are clearly distinguished in English.</p></subsubsection></subsection><subsection number="5.2." title="Improving    accuracy   using subsection numbers"><p>If we want to obtain paraphrasing words accurately rather than widely, we can use subsection numbers to improve the accuracy. Generally, the usage with high frequency is placed first in a dictionary, and the usage with low frequency is placed rearward. Therefore, we can omit the uncommon usages by cutting off the rearward subsections. To see the effect of this cutting off, we classified the result in Table 12 into subsection numbers (see Table 16). The smaller the subsection number, the better the evaluation. Furthermore, if we want to get type A only, it is covered by subsections 1 to 3 in this example, i.e., 34+8+2=44.</p></subsection><subsection number="5.3." title="Comparison with WordNet"><p>We compared our paraphrasing word with WordNet, which is an English thesaurus, using the matching word <i>"shokugyou" </i>as an example (see Table 17). We picked some entries that have the meaning of <i>"shokugyou" </i>from WordNet. The words are divided into small hierarchies in WordNet, while the words are widely contained in the same group in our proposed method.<page local="7"/> In particular, the group of the proposed method contains words for which it is difficult for non-native speakers of English to distinguish whether there is a difference or not, such as "profession" and "work". Therefore, the proposed method is especially useful for sentence matching for non-native speakers of English.</p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1021</doubt><table caption="Table 15: Evaluation of Japanese excluded words" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Type</p></td><td class="cell"><p>#</p></td><td class="cell"><p>POS detail</p></td><td class="cell"><p>Example *</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A</p></td><td class="cell"><p>38</p></td><td class="cell"><p>(O=28 ?= 5 X= 5)</p></td><td class="cell"><p><i>shimoyake </i>/ <i>toushou </i>: chilblain</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>B</p></td><td class="cell"><p>21</p></td><td class="cell"><p>(O=17 ?= 0 X= 4)</p></td><td class="cell"><p><i>ensou-kaijou </i>/ <i>kaijou </i>: concert hall</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>C</p></td><td class="cell"><p>11</p></td><td class="cell"><p>(O= 9 ?= 0 X= 2)</p></td><td class="cell"><p><i>omamori </i>/ <i>mamorifuda </i>: amulet</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>D</p></td><td class="cell"><p>23</p></td><td class="cell"><p>(O=10 ?= 4 X= 9)</p></td><td class="cell"><p><i>tabitabi </i>/ <i>shikirini </i>: frequently</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>E</p></td><td class="cell"><p>7</p></td><td class="cell"><p>(O= 2 ?= 1 X= 4)</p></td><td class="cell"><p><i>yuubi </i>/ <i>hin-i </i>: elegance</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Total</p></td><td class="cell"><p>100</p></td><td class="cell"><p>(O=66 ?=10 X=24)</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>* X / Y : matching word; where X is J in E-J, Y is a new J by adding J-E</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Table 14: Evaluation of Japanese paraphrasing words</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Type</p></td><td class="cell"><p>#</p></td><td class="cell"><p>POS detail</p></td><td class="cell"><p>Example</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A</p></td><td class="cell"><p>6</p></td><td class="cell"><p>(O= 2 ?= 2 X= 2)</p></td><td class="cell"><p><i>kenkin </i>/ <i>kanpa </i>: contribution</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>B</p></td><td class="cell"><p>23</p></td><td class="cell"><p>(O=14 ?= 1 X= 8)</p></td><td class="cell"><p><i>kenkyuu-kadai </i>/ <i>kadai </i>: assignment</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>C</p></td><td class="cell"><p>14</p></td><td class="cell"><p>(O= 8 ?= 3 X= 3)</p></td><td class="cell"><p><i>nikki </i>/ <i>nisshi </i>: diary</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>D</p></td><td class="cell"><p>24</p></td><td class="cell"><p>(O=12 ?= 3 X= 9)</p></td><td class="cell"><p><i>yoku-naru </i>/ <i>agaru </i>: improve</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>E</p></td><td class="cell"><p>33</p></td><td class="cell"><p>(O=12 ?= 3 X=18)</p></td><td class="cell"><p><i>koshou-suru </i>/ <i>kowasu </i>: break</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Total</p></td><td class="cell"><p>100</p></td><td class="cell"><p>(O=48 ?=12 X=40)</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 16: Evaluation result of each category number in E-J dictionary" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsection</p></td><td class="cell"><p>A to E total</p></td><td class="cell"><p>A</p></td><td class="cell"><p>A to D</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>1</p></td><td class="cell"><p>57</p></td><td class="cell"><p>34 (60%)</p></td><td class="cell"><p>50 ( 88%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>2</p></td><td class="cell"><p>21</p></td><td class="cell"><p>8 (38%)</p></td><td class="cell"><p>13 ( 62%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>3</p></td><td class="cell"><p>12</p></td><td class="cell"><p>2 (17%)</p></td><td class="cell"><p>9 ( 75%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>over 3</p></td><td class="cell"><p>10</p></td><td class="cell"><p>...</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Total</p></td><td class="cell"><p>100</p></td><td class="cell"><p>44</p></td><td class="cell"><p>80</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>On the other hand, words of large ambiguity (e.g., "walk", "way") and words of small ambiguity (e.g., "business") are contained in the same group. This means that the proposed method can be used to replace a word of large ambiguity with a word of small ambiguity.</p></subsection></section><section number="6." title="Conclusion"><p>In this paper we proposed a method of extracting paraphrasing words with 2-way bilingual dictionaries by matching the dictionary entries. We used it with Japanese and English for a case study. We also conducted an evaluation with a native speaker. We think that the proposed method is useful for sentence-matching. We also are planning to apply this method to improve the results in the paper by Shirai &amp; Yamamoto (2001).</p><p>Furthermore, we showed that words that are difficult for native speakers of Japanese to distinguish from each other are contained in the same group. Natural language processing should be equally friendly to non-native speakers of English as it is to native speakers of English. We hope that the proposed method will help non-native speakers of English especially in handling the variety of the potential expressions that exist.</p></section><section number="7." title="Acknowledgements"><p>The research reported here was supported in part by a contract with the Telecommunications Advancement Organization of Japan entitled, "A study of speech dialogue translation technology based on a large corpus".</p><table caption="Table 17: Comparison example with WordNet" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>WordNet</p></td><td class="cell"><p>occupation / business / line of work / line</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>=&gt; profession</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>=&gt; trade / craft</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>=&gt; job / employment / work</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>=&gt; career / calling / vocation</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Proposed method:</p><p><i>shokugyou</i></p></td><td class="cell"><p>J-E: occupation / profession / trade / job E-J: business / calling / career / craft / game / line / occupation / profession / pursuit / trade / vocation / walk / way / work</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></section><references><p>Princeton Univ., 1997. <i>WordNet 1.6, </i>http ://www. cogsci.princeton. edu/~wn/.</p><p>Shirai, S., K. Yamamoto, 2001. Linking English Words in Two Bilingual Dictionaries to Generate Another Language Pair Dictionary. In <i>Proceedings of the 19th International Conference on Computer Processing of Oriental Languages, </i>174-179.</p><p>Shirai, S., K. Yamamoto, and K. Takao, 2001. Construction of a Dictionary for Translating Japanese Phrases into One English Word. In <i>Proceedings of the 19th International Conference on Computer Processing of Oriental Languages, </i>3-8.</p><p>Yamagishi, K., 1991. <i>The New Anchor Japanese-English</i></p><p><i>Dictionary, </i>Gakken. Yamagishi, K., 1996. <i>The Super Anchor English-Japanese</i> <i>Dictionary, </i>Gakken.</p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1022</doubt></references></body></article>