<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>TRANSLATION AMBIGUITY RESOLUTION BASED ON TEXT CORPORA OF SOURCE AND TARGET LANGUAGES</title><author surname="DOI" givenname="Shinichi"><org  name="URA" country="France" city="Marseille"/></author><author surname="MURAKI" givenname="Kazunori"><org  name="BT Laboratories" country="United Kingdom" city="Ipswich"/></author></firstpageheader><frontmatter><p><b>TRANSLATION AMBIGUITY RESOLUTION BASED ON TEXT CORPORA OF SOURCE AND TARGET LANGUAGES</b></p><p><b>Shinichi DOI and Kazunori MURAKI</b></p><p>NEC Corp. C&amp;C Information Technology Research Laboratories 4-1-1, Miyazaki, Miyamae-ku, Kawasaki 216, JAPAN e-mail: (loi%mtl.cl.nec.co.jp@sj.nec.com</p></frontmatter><abstract>We propose a new method to resolve am­biguity in translation and meaning in­terpretation using linguistic statistics ex­tracted from dual corpora of source; and target languages in addition to the logical restrictions described on dictionary and grammar rules for ambiguity resolution. It provides reasonable criteria for deter­mining a suitable equivalent translation or meaning by making the dependency re­lation in the source language be reflected in the translated text. The method can be tractable because the required statis­tics can be computed semi-automatically in advance from a source language corpus and a target language corpus, while an ordinal corpus-based translation method needs a large volume of bilingual corpus of strict pairs of a sentence and its transla­tion. Moreover, it also provides the means to compute the linguistic statistics on the pairs of meaning expressions. </abstract></header><body><section number="1" title="Introduction"><p>Recently many kinds of natural language pro­cessing systems like machine translation systems have been developed and put into practical use, but ambiguity resolution in translation and meaning in­terpretation is still the primary issue in such sys­tems. These systems have conventionally adopted a rule-based disambiguation method, using linguis­tic restrictions described logically in dictionary and grammar to select the suitable equivalent transla­tion and meaning. Generally speaking, it is impos­sible to provide all the restrictions systematically in advance. Furthermore, such machine transla­tion systems have suffered from inability to select the most suitable equivalent translation if the in­put expression meets two or more restrictions, and have difficulty in accepting any input expression that meets no restrictions.</p><p>In order to overcome these difficulties, following methods arc proposed these years:</p><p>1. Example-Based Translation : the method based on translation examples (pairs of source text and its translation) [Nagao 84, Sato 90, Sumitii 90]</p><p>2. Statistics-Based Translation : the method us­ing statistical or probabilistic information ex­tracted from a bilingual corpus [Brown 90, Nomiyama 91]</p><p>Still, each of thorn has inherent problems and is insufficient for ambiguity resolution. For example, either an ex ample-based translation method or a statistics-based translation method needs a large-scale database of translation examples, and it is difficult to collect an adequate amount of a bilin­gual corpus.</p><p>In this paper, we propose a new method to select the suitable equivalent translation using the sta­tistical data extracted independently from source and target language texts [Muraki 9l], The sta­tistical data used here is linguistic statistics repre­senting the dependency degree on the pairs of ex­pressions in each text, especially statistics for co­occurrence, i.e., how frequently the expressions co-occur in the same sentence, the same paragraph or the same chapter of each text. The dependency relation in the source language is reflected in the translated text through bilingual dictionary by se­lecting the equivalent translation which maximizes both statistics for co-occurrence in the source and target language text. Mormver, the method also provides the means to compute the linguistic statis­tics on the pairs of meaning expressions. We call this method for equivalent translation and meaning selection DMAX Criteria (Double Maximize Crite­ria based on Dual Corpora).</p><p>First, we make comments on the characteristics and the limits of the conventional methods of am­biguity resolution in translation and meaning inter­pretation in the second section. Next, we describe the details of DMAX Criteria for equivalent trans­lation selection in the third section. And last, we explain the means to compute the linguistic statis­tics on the pairs of meaning expressions.</p><doubt alpha="48.9" length="92" tooSmall="False" monospace="0.0">Actes de COLING-92, Nantes, 23-28 août 1992 S 2 S Troc,ovCOL1NG-92, Nantes, Aug. 23-28, 1992</doubt><page local="2"/></section><section number="2" title="Conventional Methods of Ambiguity Resolution"><subsection number="2.1" title="Rule-Based Translation"><p>In conventional methods, linguistic restrictions described in the dictionary and grammar are used to select the suitable equivalent translation or meaning. In general, these restrictions are de­scribed logically on characteristics of another ex­pression which modifies or is modified by the ex­pression to be processed. For example, to translate predicates (verbs and predicative adjectives), se­mantic restrictions are described on essential case arguments in forms of semantic markers to indicate features of words or terms in the thesaurus to show a hierarchy composed of word concepts.</p><p>Though these conventional methods have been very useful to realize natural language processing systems, they have the following problems:</p><p>1. It is impossible to decide the most suitable equivalent translation if the input expression meets two or more restrictions.</p><p>2. Analysis fails when the input expression can meet no restrictions.</p><p>3. Actually the practical systems depends on such heuristics as pre-decided application or­der of restrictions or some default equivalent translations or meanings.</p><p>4. The description of the restrictions is based on direct structural dependencies, therefore it is quite difficult to describe the restrictions based on sister-dependency or between expressions belong to different sentences or paragraphs.</p><p>5. Restrictions on any dependencies cannot be thoroughly described in advance.</p><p>For example, a Japanese word "boom" has two meanings, one is 'a ball(a <i>round object used in a game or sport)<footnote anchor="7"/> </i>and the other is 'a bowl(a <i>deep round container open at the toj) especially used in cooking)'. </i>When this word occurs in the following sentence, it must mean 'a bowl<footnote anchor="1"/>.</p><doubt alpha="75.0" length="4" tooSmall="False" monospace="0.0">Jap:</doubt><p>Booru-ni bowl dative or marker ball <i>mizu-o ireru </i>water obj, pour, marker   put in or fill <b>ENG:</b><b>   To </b>pour water into a bowl</p><p>In this case, to select the meaning by the logical re­strictions on dependencies, it is necessary to have described even the appearance or usage of the in­direct object of the verb "ireru". To describe such detail restrictions on «all expressions may be possi­ble, but it is quite difficult because the trouble of description and the cost of calculation.</p></subsection><subsection number="2.2" title="Example-Based Translation"><p>Besides the conventional translation method above, a machine translation system based on translation examples (pairs of source texts and their translations) is also proposed [Nagao 84, Sato 90, Sumita 90]. This type of system, called Example-Based Machine Translation, has stored a large amount of bilingual translation examples as a database, and translates input expressions by re­trieving an example most similar to the input from the database. There is no failure of output in this method because it selects the most similar example not the identical one.</p><p>However this example-based translation system needs a large-scale database of translation exam­ples, and it is difficult to collect an adequate amount of bilingual corpora. Even if it is possible, there is no means to divide the sentences of such corpora into fragments and link them automati­cally, and it costs us too much time and money to divide and link manually. Besides, this method can neither achieve precise meaning interpretation be­cause it selects equivalent translation directly from the input expression and leaves meaning interpre­tation out of consideration.</p><p>To overcome this problem, we have also proposed a new mechanism based on sentential examples in dictionary, which utilize the merits of both the translation by logical restrictions and the example-based method, by selecting the equivalent transla­tion which has the most similar example to the in­put expression [Doi 92]. This mechanism can guar­antee no failure in selecting an equivalent transla­tion, but the description of relations are still based only on direct structural dependencies.</p></subsection><subsection number="2.3" title="Statistics-Based Translation"><p>Several new methods especially of machine trans­lation have been proposed lately, which select a suitable equivalent translation using statistical or probabilistic information extracted from language text [Brown 90, Nomiyama 9l]. Because many ma­chine readable texts have been already collected nowadays, it is not difficult to extract statistical information of each expression in the texts semi-automatically. Moreover, the statistical informa­tion reflects the context in which each word occurs and implies the logical restrictions based on indi­rect structural dependencies.</p><p>Although we call the systems in a same word "statistics-based translation", statistical informa­tion used in the methods is diverse, such as trans­lation probability, connectivity of words, statistics for (co-)occurrcnce, etc. We make comments on the characteristics and the limits of these systems.</p><p>The first method uses fertility probabilities, translation probabilities and distortion probabili­ties [Brown 90]. Fertility means the number of the words in target language that the word of the source language produces, and distortion means the distance between the position of the word of the source language and the one of the target language.<page local="3"/> The method has been applied to an experimental translation system from French to English. How­ever, since these probabilities are extracted from a large amount of text pairs that are translations of each other, this method must be suffered from the same difficulties as example-based translation in collecting and analyzing an adequate amount of bilingual corpora, and it's very difficult to ap­ply this method to the languages whose linguistic; structures aren't similar each other, such as English and Japanese.</p><doubt alpha="45.7" length="94" tooSmall="False" monospace="0.0">Actes de COUNG-92, Nai^tes, 23-28 août 1992 5 2 6 Proc. of COLING-92, Nantes. Aug. 23-28, 1992</doubt><p>The second method uses the statistics for occur­rence in target language text [Nomiyama 91]. It is calculated in advance how frequently the each ex­pression occurs in the target language text, which needs only to belong the same field as the source language text belongs, but not to be a translated text of the source language text. If there are more than one possible equivalent translations, the most frequent translation is selected through this calcu­lated data. Moreover, this method can be applied to make good use of the conventional methods of selecting equivalent translations, for it employs the frequency data exclusively when logical restrictions cannot select one out of candidates.</p><p>However this method has one big problem. The high frequency of the expression in the target lan­guage text may not originate from the frequency of the expression in the source language text to be translated, because one target language expression does not correspond to only one source language expression in general.</p><p>Suppose the following sentence is a first example:</p><p>JAP:     <i>Sono   saibankan-wn kooto-to</i> market or <i>nekutai-o kattn.</i><i></i></p><doubt alpha="58.8" length="34" tooSmall="False" monospace="0.0">that    judge       subj. coat and</doubt><doubt alpha="100.0" length="5" tooSmall="False" monospace="0.0">court</doubt><doubt alpha="62.1" length="29" tooSmall="False" monospace="0.0">tie        obj. bought marker</doubt><p><b>jj.</b></p><p><b>ENG:    </b>The judge bought a coat and a tie.</p><p>Figure 1 indicates translation process through bilingual dictionary and the statistics for co­occurrence of each pair of expressions in both Japanese and English necessary to translate the sentence<footnote anchor="1"/>. The Japanese word <i>"kooto" </i>has two equivalent English translations: <footnote anchor="4"/> (over)coat<footnote anchor="1"/> and '(tennis) court*. We cannot decide which is eligible</p><p>lThe statistics for co-occurrence of expressions shown in the figures are given provisionally for understanding.</p><p>with only logical restrictions on the direct object of the Japanese verb "Jean", because we can buy both 'coat' and 'court'—the sentence <i>"Tenisu-kooto o kau" </i>'To buy a tennis court' is also quite accept­able. In this case, the statistics for co-occurrence in the target language English text denotes that the most frequent pair is 'court-judge', because the word 'court' also means a 'law court'. Then using only statistical data on the target language text misleads a wrong expression 'court' as the equiva­lent translation of <i>"kooto", </i>and the example sen­tence may be translated into 'The judge bought a court and a tie.'.</p><p>The second example is this sentence<footnote anchor="2"/>:</p><doubt alpha="65.7" length="35" tooSmall="False" monospace="0.0">j ai':Kotori-no    kago-ni uiizu- o</doubt><p>bird     of    cage dative     water obj.</p><p>or marker marker basket <i>iretii booru-o oit a.</i><i></i></p><p>filled bowl   obj. put or marker</p><doubt alpha="100.0" length="4" tooSmall="False" monospace="0.0">ball</doubt><p><b>Enc</b>;:    I put a bowl filled with water in the bird cage.</p><p>Translation process of this sentence and the statistics for co-occurrence are shown in Figure 2. Because the pair of 'basket' and 'ball' co-occurs most frequently in the target language, the sen­tence may be translated into 'I put, a ball filled with water in the bird basket.'.</p><p><b>3    Equivalent Translation Selection by Statistical Data on Dual Corpora of Source and Target Languages</b></p><p>Now we propose a new method to provide rea­sonable criteria for selecting a suitable equivalent translation or meaning using the simple statistical data extracted from source language text in addi­tion to the one from target language text. These source and target language texts don't have to be translations of each other. The proposed method gives us a way to select the expression with the highest frequency of the target language text that keeps high frequency of the source language text at the same time, so it overcomes the difficulty of the method using the frequency data on the target language text only, because it does not select the expression with the highest frequency of only the target language text.</p><p>The subject phrase "watasiji-wft" = T is omitted in this sentence.</p><doubt alpha="49.4" length="89" tooSmall="False" monospace="0.0">Actes deCOLING-92, Nantes, 23-28 août19925 2 7 Proc.ovCOLING-92, Nantes, Aug. 23-28, 1992</doubt><page local="4"/></subsection><subsection number="3.1" title="Using statistical data on source language text"><p>The method using only statistical data on the target language text may mislead a wrong equiv­alent translation, because in general each target language expression corresponds to more than one source language expression.</p><p>The equivalent translation selection with statis­tics for co-occurrence in the target language text when a source language expression <i>Sa </i>has <i>n </i>equiva­lent translations in target language Tu;(z — 1 ■ • ■ <i>n)</i><i> </i>is shown as this:</p><doubt alpha="66.7" length="3" tooSmall="False" monospace="0.0">Si,</doubt><doubt alpha="66.7" length="12" tooSmall="False" monospace="0.0">SCO(Toi,Tt&gt;)</doubt><doubt alpha="100.0" length="5" tooSmall="False" monospace="0.0">where</doubt><p><b>Sjl- </b>: source language expression</p><p>Tfci : <i>n </i>target language equivalent</p><p><i>(i — </i>1 ■ ■ ■ <i>n)      </i>translations of <b>SCO(E;,Ej)    </b>: statistics for co-occurrence of two expressions E,-,Ej</p><p>The method using only statistical data on the tar­get language text selects T„,* which maximizes the statistics for co-occurrence in the target language text <footnote anchor="3"/> as the equivalent translation of S„, where the partner of the co-occurrence <i>T^j </i>plays the part of the basis for the equivalent translation selection. The biggest problem of this method is that Ttj which depends both <i>b </i>and <i>j </i>is selected by only sta­tistical data on the target language text.</p><p>Our new method provides reasonable criteria for selecting the basis for the equivalent translation se­lection using the statistical data on the source lan­guage text. First the source language expression Sf, which maximizes the statistics for co-occurrence in the source language text <footnote anchor="4"/> is selected, then the equivalent translation Trt^ which maximizes the statistics for co-occurrence in the target language text <footnote anchor="5"/> is selected. The dependency relation in the source language is reflected in the translated text through this method. We call this method for equivalent translation and meaning selection DMAX Criteria (Double Maximize Criteria based on Dual Corpora).</p><p><b>Ä,2   Double Maximum Criteria based on Dual Corpora</b></p><p>The algorithm of this method is summarized as follows:</p><p>1. Prepare the source and target language texts (the target language text needs not to be a translated text of the source language text).</p><p>sToi|maxi(ij*SCO(T0i,Tw) *Ttti|m!ttijSCO(Tai,Tw)</p><footnote label="4">S fc |mjix*SCO(S a ,S t )</footnote><p>2. Accumulate the statistics for co-occurrence of every expression in both texts.</p><p>3. When a source language expression Sa has <i>n </i>equivalent translations in target language (a) Select St|maxtSCO(Sa,S4) (b) Select Tai I maxij SCO(Tai, <i>Tbj)</i></p></subsection><subsection number="3.3" title="Operation Example"><p>Figure 1-3 show operation examples. Figure 1 and 2 are examples of Japanese-English transla­tion. In Figure 1, with only statistical data on the target language text, 'court<footnote anchor="1"/> may be chosen as an equivalent translation of <i>"kooto" </i>because the pair of 'court-judge' co-occurs most frequently in the target language. However with DMAX Crite­ria, the equivalent translation of <i>"kooto" </i>is selected correctly.</p><p>• The expression which co-occurs with <i>"kooto" </i>most frequently in the source language is <i>"nekutas".</i></p><p>• The pair of the equivalent translation of <i>"kooto" </i>and the one of "neJcutai" which co-occurs most frequently in the target language</p><p>is 'coat-tie'.</p><p>• As a result, <i>"kooto" </i>is translated into 'coat'.</p><p>It is the same as shown in Figure 2. A pair of 'basket-ball' co-occurs most frequently in the target language. Dut using DMAX Criteria, giv­ing attention first to the most frequent pair in the source language text, <i>"kotoii-kago" </i>can gain the correct equivalent translation 'cage'. Next, a pair of "mizu-booru" decides 'bowl' as an equivalent translation of "boom". Finally, correct translation can be acquired in this way.</p><p>Figure 3 shows the translation process and the statistics for co-occurrence of another English-Japanese translation example.</p><p><b>ENG:    </b>The ceiling of the court was cleaned quite well.</p><p><b>j ai':     </b><i>Saibansho-no tcnjoo~wa</i> <i>kireini sonji-s&amp;reteita.</i><i> </i>quite well be cleaned</p><doubt alpha="63.2" length="38" tooSmall="False" monospace="0.0">court       of    ceiling subj. marker</doubt><p>In this case, the English words 'court' and 'clean' have two meanings respectively.</p><doubt alpha="71.4" length="7" tooSmall="False" monospace="0.0">'court'</doubt><p><b>saibansho </b><i>a room or building in which law cases can be heard and judged</i></p><p><b>kooto </b><i>(a part of) an area specially prepared and marked for various ball games, such as tennis</i><page local="5"/></p><doubt alpha="47.8" length="92" tooSmall="False" monospace="0.0">Actes de COLING-92, Nantes, 23-28 août19925 2 8 Proc. oe COLING-92, Nantes, Aug. 23-28, 1992</doubt><doubt alpha="62.5" length="8" tooSmall="False" monospace="0.0">* clean'</doubt><p><b>souji-snru </b><i>to clean rooms</i></p><p><b>kuriiningu-suru </b><i>to clean clothes with chemicals instead of water</i></p><p>A pair of <i>"kooto -kuriiniiigu" </i>co-occurs most fre­quently in the target language, so the sentence may be translated into "Kooto-no tenjoo-iia Jcireini <i>kuriiningu-sareteita. </i>". But using DMAX Criteria, 'ceiling' is selected as a basis for the equivalent translation selection of 'court', and ''saibanslio" is selected as an equivalent translation of 'court' by the comparison between statistics for co-occurrence on the pairs of "teiijoo-saibansiio" and <i>"tenjoo kooto".</i></p><p><b>4   Calculation of Linguistic Statistics for Semantic Interpretation</b></p><p>In language understanding systems or machine translation systems through semantic expressions, one suitable meaning must be selected out of the ones described in a dictionary according to an entry word. However in conventional systems the mean­ing selection mechanism isn't robust and cannot select the most suitable meaning only by logical restrictions described in the dictionaries. We pre­sented a new method for the equivalent transla­tion selection in the former chapter using statis­tical data on source language and target language through bilingual dictionary. To apply this method to meaning selection, it is necessary to calculate statistical data on the pairs of each meaning in ad­vance, but there is no means of calculating them automatically.</p><p>We have already developed an intcrlingua-based machine translation system whose interlin-gua named PIVOT doesn't depend on any par­ticular natural language [Muraki 86, Ichiyama 89, Okumura 91]. In its dictionary, as illus­trated in Figure 4., expressions in the source lan­guage are mapped onto some interlingua vocab­ularies (CONCEPTUAI^PRIMITIVE:CP), which are next mapped onto some equivalent translations. Then we propose a new method of computing lin­guistic statistics for occurrence of meanings auto­matically using this format of dictionary.</p><p>Suppose linguistic statistics on the pairs of ex­pressions in both source and target language texts have already been calculated. In case of transla­tion, when an expression occurs in the source language text, an equivalent translation T,^ is de­cided through the passage of Si =&gt;C;y =i&gt;T^f and as a result, CPC^j is also selected from the CPs corresponding to the expression S;. Therefore, the linguistic statistics on the pairs of CPs or meanings is nothing but coupling linguistic statistics on the pairs of corresponding expressions in the target lan­guage text. Thus, the linguistic statistics iî on the pairs of the meaning expressions in the dictionary can be obtained as the sum of the linguistic statis ­tics <i>u </i>on the pairs of target language expressions according to the following equation.</p><p>This linguistic statistics can be added to the dic­tionary in advance, and we can select the meaning in the same way as equivalent translation selection.</p></subsection></section><section number="5" title="Conclusion"><p>We proposed a new method DMAX Criteria (Double Maximize Criteria based on Dual Corpora) in this paper. It can select a suitable equivalent translation or meaning using the statistical data extracted from both source and target language corpora even when linguistic restrictions described in the dictionary or grammar cannot. The depen­dency relation in the source language is reflected in the translated text through bilingual dictionary. Moreover, the method has the following features;</p><p>1. It utilizes linguistic statistics as context infor­mation in addition to logical restrictions effec­tive for ambiguity resolution.</p><p>2. The source of the linguistic statistics is the dual corpora of source and target languages, not the bilingual corpora (the target language text doesn't have to be the translation of the source language text).</p><p>3. The linguistic statistics can be computed semi-automatically in advance.</p><p>4. The linguistic statistics on the pairs of mean­ing expressions are computed from the lin­guistic statistics in source and target language texts with the intcrlingua-based bilingual dic­tionary to resolve ambiguity in meaning inter­pretation,</p><p>Based oji this method, we have carried out an experiment on a limited-scale translation system, and confirmed effectiveness of the method. We are preparing further experiments on a large-scale dual corpora with PIVOT interlingua dictionary. Their result will be reported on another paper.</p></section><section number="6" title="Acknowledgments"><p>The authors wish to thank Mr. Masao WATAR1 for his continuous encouragement. The authors also thank the members of Media Technology Lab­oratory for their good suggestions.</p><doubt alpha="46.8" length="94" tooSmall="False" monospace="0.0">Actes de COLING-92, Nantes, 23-28 août 1992 5 2 9 Proc. of COLING-92, Nantes, Aug. 23-28, 1992</doubt><page local="6"/><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>i    </b><b>bilingual dictionary</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>^celling</b></p></td><td class="cell"><p><b>-tenjoo —</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>f </b><b>court </b><b>"c</b></p></td><td class="cell"><p><b><i>__saibansho ~~</i></b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b><i>"—kooto -</i></b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>f </b><b>clean*</b></p></td><td class="cell"><p><b><i>-kuriiningu</i></b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>j </b><b>coat --</b></p></td><td class="cell"><p><b><i>—kooto</i></b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></section><references><p>[Brown 90] P.F.Brown et al. "A Statistical Ap­proach to Machine Translation", Computa­tional linguistics , Vol.16, No.2, 1990</p><p>[Doi 92] S.Doi and K.Muraki "Robust Translation and Meaning Interpretation Mechanism based on Examples in Dictionary", <i>Proc. of 44th Annual Conference of IPSJ </i>, 1P-2, 1992 (in Japanese)</p><p>[Ichiyama 89] S.Ichiyama "Multi-lingual Machine Translation System," <i>Office Equipment und Products, </i>18-131, pp.46-48, August 1989</p><p>[Muraki 86] K.Muraki "VENUS: Two-phase Ma­chine Translation System," <i>Future Genera­tions Computer Systems, </i>2, pp.117-119, 1986</p><p>[Muraki 91] K.Muraki and S.Doi "Translation Ambiguity Resolution by using Text Corpora of Source and Target Languages", Proc. <i>of 5th Annual Conference of JSAI </i>, 11-7, 1991 (in Japanese)</p><p>[Nagao 84] M.Nagao "A Framework of a Mechani­cal Translation between Japanese and English by Analogy Principle", Artificial and Human Intelligence, <i>ed. A.Elithorn </i>and R.J3anerji, <i>North-Holland </i>, 1984</p><p>[Nomiyama 91] H.Nomiyama "Lexical Selection Mechanism Using Target Language Knowl­edge and Its Learning Ability", <i>IPSJ-WG , </i>NL86-8 , 1991 (in Japanese)</p><p>[Okumura 9l] A.Okumura, K.Muraki and S.Aka-mine "Multi-lingual Sentence Generation from the PIVOT interlingua," Proc. <i>of MT SUM­MIT III, </i>pp.67-71, July 1991</p><doubt alpha="66.7" length="78" tooSmall="False" monospace="0.0">[Sato 90] S.Sato and M.Nagao "Toward Memory-based Translation",COLING-90, 1990</doubt><p>[Sumita 90] E.Sumita, H.Iida and H.Kohyama "Example-based Approach in Machine Trans­lation", InfoJapan'flO , 1990</p><p><b>statistics for co-occurrence J in source language text </b><b>'(</b></p><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">f</doubt><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">V</doubt><doubt alpha="100.0" length="7" tooSmall="False" monospace="0.0">nekutai</doubt><doubt alpha="100.0" length="5" tooSmall="False" monospace="0.0">kooto</doubt><p><b><i>aaibankan</i></b> <b><i>saibansho</i></b></p><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">50</doubt><doubt alpha="0.0" length="1" tooSmall="False" monospace="0.0">^</doubt><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">10</doubt><doubt alpha="50.0" length="2" tooSmall="False" monospace="0.0">J_</doubt><doubt alpha="0.0" length="1" tooSmall="False" monospace="0.0">\</doubt><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">80</doubt><footnote label="9">*********************************ß</footnote><p><b>bilingual dictionary</b></p><p><b><i>-nekutai-</i></b> <b><i>-fsaibankan </i></b><b>—judge</b></p><doubt alpha="75.0" length="4" tooSmall="False" monospace="0.0">-tie</doubt><doubt alpha="60.0" length="10" tooSmall="False" monospace="0.0">-t-kooto ■</doubt><doubt alpha="66.7" length="21" tooSmall="False" monospace="0.0">..saibansho---court '</doubt><p><b>statistics for co-occurrence in target language text</b></p><doubt alpha="80.0" length="5" tooSmall="False" monospace="0.0">—rtie</doubt><doubt alpha="100.0" length="2" tooSmall="False" monospace="0.0">SO</doubt><doubt alpha="44.0" length="25" tooSmall="False" monospace="0.0">—^-coat -s.» -^-court-?0I</doubt><doubt alpha="33.3" length="6" tooSmall="False" monospace="0.0">/80j J</doubt><doubt alpha="50.0" length="10" tooSmall="False" monospace="0.0">-^-Judge-^</doubt><p><b>Figure 1 </b><b><i>"Sono saibankan-wa kooto-to nekutai-o katta." </i></b><b>'The judge bought a coat and a tie.'</b><page local="7"/></p><doubt alpha="48.9" length="90" tooSmall="False" monospace="0.0">Actes de COLING-92, Nantes, 23-28 août19925 3 0 Proc.orCOLING-92, Nantes, Aug. 23-28, 1992</doubt><doubt alpha="100.0" length="5" tooSmall="False" monospace="0.0">court</doubt><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">f</doubt><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">V</doubt><doubt alpha="50.0" length="2" tooSmall="False" monospace="0.0">J_</doubt><doubt alpha="0.0" length="1" tooSmall="False" monospace="0.0">\</doubt><p><b>bilingual dictionary</b></p><doubt alpha="100.0" length="6" tooSmall="False" monospace="0.0">kotori</doubt><doubt alpha="100.0" length="4" tooSmall="False" monospace="0.0">kago</doubt><doubt alpha="80.0" length="5" tooSmall="False" monospace="0.0">—boom</doubt><doubt alpha="80.0" length="5" tooSmall="False" monospace="0.0">-bird</doubt><p><b>cage 4-basketh ball </b><b>X</b></p><doubt alpha="66.7" length="6" tooSmall="False" monospace="0.0">bowl4-</doubt><doubt alpha="60.0" length="10" tooSmall="False" monospace="0.0">-uiater +-</doubt><p><b>^statistics for co-occurrence; {in target language text</b></p><doubt alpha="57.1" length="7" tooSmall="False" monospace="0.0">bird ~\</doubt><p><b>—Vcage </b><b><i>-r </i></b><b>basket </b><b>_j, </b><b>ball</b> <b>v mater '</b></p><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">r</doubt><doubt alpha="75.0" length="4" tooSmall="False" monospace="0.0">bow!</doubt><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">30</doubt><p><b>Figure 2 </b><b><i>"Kotori-no kago-ni mizu-o ireta booru-o oita."</i></b> <b>'I put a bowl filled with water in the bird cage.</b><b><footnote anchor="1"/></b></p><p><b>statistics for co-occurrence in source language text</b></p><p><b>celling -y-</b></p><doubt alpha="25.0" length="4" tooSmall="False" monospace="0.0">40J_</doubt><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">"\</doubt><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">20</doubt><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">60</doubt><p><b><i>I </i></b><b>statistics for co-occurrence </b><b>'•,</b><b> in target language text</b> <b><i>—^saibanah</i></b></p><doubt alpha="0.0" length="2" tooSmall="False" monospace="0.0">40</doubt><doubt alpha="50.0" length="12" tooSmall="False" monospace="0.0">—j*kooto.■20</doubt><p><b>-soujj.</b></p><p>-Jcuriixiingu ^</p><p><b>Fieure3 'The ceiling of the court was cleaned quite well.'</b> <b>"Saibansho </b><b><i>no tenjoo-wa kxreini </i></b><b>souji-sareteiti</b></p><doubt alpha="47.9" length="48" tooSmall="False" monospace="0.0">Actes de COLING-92, Nantes, 23-28 août 1992 5 31</doubt><doubt alpha="50.0" length="42" tooSmall="False" monospace="0.0">l'Roc.or COLING-92,Nantes, Auu.23-28, 1992</doubt><doubt alpha="0.0" length="1" tooSmall="False" monospace="0.0">2</doubt><p><b>clean coat</b></p></references></body></article>