<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="147"/><title>KUNLP system using Classification Information Model at SENSEVAL-2</title><author surname="Seo" givenname="Hee-Cheol"><org  name="Korea University" country="Korea" city="Seoul"/></author><author surname="Lee" givenname="Sang-Zoo"><org  name="Korea University" country="Korea" city="Seoul"/></author><author surname="Rim" givenname="Hae-Chang"><org  name="Korea University" country="Korea" city="Seoul"/></author><author surname="Lee" givenname="Ho"><org  name="Frei University" country="Germany" city="Berlin"/></author></firstpageheader><frontmatter><p>KUNLP system using Classification Information Model</p><p>at SENSEVAL-2</p><p>Hee-Cheol Seo, Sang-Zoo Lee, Hae-Chang Rim Ho Lee</p><p>Dept. of Computer Science and Engineering, Astronest Inc.</p><p>Korea University 135-090 3rd floor, Hanam BD</p><p>1, 5-ka, Anam-dong 157-18 Samsung-Dong</p><p>Seongbuk-Gu, Seoul, 136-701, Korea Kangnam-Gu, Seoul, Korea</p><p>{hcseo,zoo,rim}@nlp.korea.ac.kr leeho@astronest.com</p></frontmatter><abstract>The classification information model or CIM classi­fies instances by considering the discrimination abil­ity of their features, which was proven to be useful for word sense disambiguation at Senseval-1. But the CIM has a problem of information loss. KUNLP system at Senseval-2 uses a modified version of the CIM for word sense disambiguation. We used three types of features for word sense disambiguation: local, topical, and bigram context. Local and topical context are similar to Chodorow's context and refer to only unigram information. The window of a bigram context is similar to that of a local context but a bigram context refers to only bigram information. We participated in the English lexical sample task and the Korean lexical sample task, where our sys­tems ranked high. </abstract></header><body><section number="1" title="Introduction"><p>The classification information model (Ho, 1997) is the model that classifies instances by considering the discrimination ability of their features. In the CIM, a feature with high discrimination ability con­tributes to the classification more than one with low discrimination ability. Hence, we can omit the fea­ture selection procedure.</p><p>The CIM has a kind of information loss problem due to the assumption that a feature contributes to only one class. We devised a modified version of the CIM where a feature can contribute to all classes.</p><p>Word sense disambiguation task can be treated as a kind of classification process(Ho, 2000). When a classification technique is applied to word sense dis­ambiguation, an instance corresponds to a context containing a polysemous word and its class to the proper sense of the word, and one of its features to a piece of context information. As a classifica­tion problem, word sense disambiguation task can be solved by the CIM.</p><p>We used three types of features for word sense disambiguation: local, topical, and bigram context. Local and topical context are similar to Chodorow's context(Chodorow, 2000) and consist of only unigram information. A bigram context has a similar window to a local context but consists of only bigram information.</p></section><section number="2" title="KUNLP system"><p>To disambiguate senses, we did two phases: corpus preprocessing and sense disambiguation. Figure 1 shows the flow chart of our system.</p><doubt alpha="100.0" length="6" tooSmall="False" monospace="0.0">Corpus</doubt><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">ï</doubt><p>Sense-Tagged Corpus</p><figure caption="Figure 1: Flow chart of KUNLP system"></figure><subsection number="2.1" title="Corpus preprocessing"><p>At the corpus preprocessing phase, we tokenized a corpus and then tagged it with parts-of-speech using Brill's Tagger(Brill, 1994). The tokenizer just sepa­rates symbols from a word. For example, a sentence <i>"I'm straight, white, no longer middle class, anti-IRA, have </i>..." is tokenized to "J<i>'m stright , white , no longer middle class , anti - IRA , have </i>Un­like other symbols, an apostrophe is not separated from the following characters.</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Tokenizer</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p><i>I</i></p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>POS-Tagger</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Corpus Preprocessing</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Phrase Filter</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>i</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Sense Tagger</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>using</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Modified CIM</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Sense Disambiguation</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="2" global="148"/></subsection><subsection number="2.2" title="Phrase filtering"><p>At the phrase filtering phase, we filtered senses using the satellite feature, which is marked with <i>sat </i>tag in training and test corpus given by the task organizer. For example, in a sentence <i>This air of disengagement &lt; head sats="carry-over. 067:0"&gt; carried&lt; /head&gt; &lt;sat id=" carry-over. 067:0"&gt;over&lt;/sat&gt; to his apparent attitude toward his things, carried over </i>is a phrase and also a satellite feature.</p><p>Phrase filtering is applied to sense disambiguation as in Table 1</p><p>Table 1: phrase filtering and sense disambiguation</p><p>if the number of filtered senses = 1 then determine sense</p><p>else if the number of filtered senses &gt; 1 then execute sense-tagger with the filtered senses</p><p>else if the number of filtered senses <i>— </i>0 then execute sense-tagger with all senses</p><p>There are satellite features in the English lexical sample, but not in the Korean lexical sample. Hence, phrase filtering was applied only in the English lex­ical sample task.</p></subsection><subsection number="2.3" title="Classification Information Model (CIM)"><p>The CIM is a kind of classification model based on the entropy theory. Given an input instance, the CIM decides the proper class of the instance by con­sidering individual decisions made by each feature of the instance. In the model, the proper class of an instance,X, is determined by Equation 1.</p><doubt alpha="66.7" length="33" tooSmall="False" monospace="0.0">Class(X)d=arg maxRe\(classj,X)(1)</doubt><doubt alpha="100.0" length="6" tooSmall="False" monospace="0.0">classj</doubt><p>where <i>class j </i>is the j-th class and <i>Kel(classj, </i><i>X)</i><i> </i>is the relevance between the j-th class and the instance <i>X. </i>Here, if we assume that features are independent of each other, the relevance can be defined as in Equation 2.</p><doubt alpha="61.1" length="18" tooSmall="False" monospace="0.0">Re\(classj,X) =Xj'</doubt><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">(2)</doubt><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">1=1</doubt><p>where <i>m </i>is the size of the feature set, <b><i>X{ </i></b>is the value of the 2-th feature and <b><i>Wij </i></b>is the weight of the <i>i­</i>th feature for the j-th class. In Equation 2, <b><i>X{ </i></b>has a binary value (1 if the feature occurs within the window, 0 otherwise) and is defined in terms of classification information.</p><p>The classification information of a feature is com­posed of two components. One is the discrimination score (DS), which represents the discrimination abil­ity of classifying instances. The other is the most probable class (MPC), which represents the most closely related class to the feature. <b><i>Wij </i></b>is defined by using these two components as follows:</p><p>def <b><i>Win —</i></b></p><p>DS; if <i>class j </i>= MPC, 0 otherwise</p><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">(3)</doubt><p>In Equation 3, <i>DSi </i>and <i>MPd </i>represent the DS and MPC of the <i>i-th. </i>feature, respectively. In the CIM, DS and MPC are defined in terms of the con­ditional probability of a class given a feature, which is normalized by the corpus size. The normalized conditional probability is defined as follows:</p><p>def <b><i>Pji =</i></b></p><p><i>p(classj\fi</i></p><doubt alpha="100.0" length="5" tooSmall="False" monospace="0.0">Njclt</doubt><p><b><i>N(classj</i></b></p><doubt alpha="66.0" length="47" tooSmall="False" monospace="0.0">Tnr)(rlnmi.\f )N(class)2^k=lP\aassk\Ji)N(dassk)</doubt><p><i><u>p(fi\classj)</u></i></p><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">(4)</doubt><p>In Equation 4, <i>pji </i>is a normalized conditional probability, <i>N(classj) </i>is the number of instances belonging to the j-th class in the training data, <i>N(class) </i>is the average number of instances for each class and <i>n </i>is the number of classes. Given the normalized conditional probability distribution, DSs and MPCs are defined as follows:</p><doubt alpha="52.6" length="19" tooSmall="False" monospace="0.0">DSi   = log2n~H(pi)</doubt><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">n</doubt><doubt alpha="31.2" length="32" tooSmall="False" monospace="0.0">=     l°g2n+ ]CPUl°&amp;2Pji(5)3 = 1</doubt><doubt alpha="100.0" length="3" tooSmall="False" monospace="0.0">MPd</doubt><doubt alpha="100.0" length="3" tooSmall="False" monospace="0.0">def</doubt><p>arg max <i>pji</i></p><doubt alpha="85.7" length="7" tooSmall="False" monospace="0.0">class j</doubt><doubt alpha="57.7" length="26" tooSmall="False" monospace="0.0">=    arg maxp( fi\clas s j</doubt><doubt alpha="0.0" length="3" tooSmall="False" monospace="0.0">(6)</doubt><p>In Equation 5, <i>H </i><b><i>(pi) </i></b>is the entropy of the i-th feature over the normalized conditional probability distribution.</p></subsection><subsection number="2.4" title="Modifying CIM"><p>The CIM has a problem caused by using MPCs, which is information loss. For example, let us con­sider the situation in Table 2 and Table 3. Table 2 shows the normalized conditional probability distri­bution, DSs and MPCs of features in an instance. Table 3 shows the weights and the relevance values at the CIM using <b><i>Wij </i></b>and at the modified CIM us­ing <i>w^, </i>for the instance of Table 2. The feature <i>f\ </i>co-occurred with <i>class\ </i>and <b><i>class2 </i></b>and the MPC of <i>fi </i>is <i>classi </i>at Table 2.  In the CIM, this feature<page local="3" global="149"/></p><table caption="Table 2: A normalized conditional probability, DSs and MPCs of features of an instance"></table><p>Table 3: The weights and the relevance values at the CIM using and at the modified CIM using <i>w^, </i>for the instance of Table 2 contributes to only <i>classi.</i><i> </i>Actually the feature }\ can contribute to distinguishing <i>class2 </i>from <i>classs </i>if it consults the normalized conditional probability distribution. In the CIM, however, the feature can not distinguish them because their weights have the same value.</p><p>Another aspect of the problem is that the CIM fails to capture the minor contribution of features, which is crucial in the case where the sum of the minor contribution of features to a non-MPC class dominates that of the major contribution of fea­tures to MPC classes. For example, at Table 2, all features, /i, <i>f2, </i>and <i>fs,</i><i> </i>have different MPCs: <i>classi, classs </i>and <i>class^, </i>respectively, it is also ob­vious that they have some minor contribution to the <b><i>class2. </i></b>The CIM will classify the instance as <i>classi </i>because <i>Rel(class\,X) — </i>1.1187 is the maximum number among the <i>Rel(classj, </i><i>X).</i><i> </i>However, if we consider the minor contribution of all the features, we prefer <b><i>class2 </i></b>to <i>classi </i>because <b><i>class2 </i></b>intuitively gains the total contribution more than <i>clas$\.</i></p><p>A solution to the problem may be not to use MPCs, but to use a measure of contribution of a feature to a class which is proportional to the dis­crimination score of the feaure and the normalized conditional probability of the class given the feature. The modified CIM can be defined as follows:</p><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">m</doubt><doubt alpha="56.0" length="25" tooSmall="False" monospace="0.0">Rc\(classj,X)=XjWjj(7)1=1</doubt><p><b><i>Wij </i></b>d= <i>DSi </i>x <i>Pji </i>(8) As shown in Table 3, the <i>w\2 </i>is larger than <i>wis </i>(0.3356 &gt; 0) and the instance is classified not as <i>classi </i>but as <b><i>class2 </i></b>because <i>Rel</i><b><i>(class2</i></b><i>,X)</i><i> </i><i>=</i></p><p>1.0028 &gt; <i>Rel(class!,X) = </i>0.7831, which is based on the modified CIM.</p></subsection><subsection number="2.5" title="Feature Space"><p>We used three types of features for word sense dis­ambiguation: local, topical and bigram context. In the preliminary experiment, we have observed that, when the CIM considered all these three types of features, it mostly achieved the best result.</p><subsubsection number="2.5.1" title="Local context"><p>In a local context, there can be features of the fol­lowing templates for all words within its window:</p><p>• in the English lexical sample task</p><p><i>— word-position </i>: a word and its position <i>— word.</i><i>POS </i>: a word and its part-of-speech</p><p><i>— POS-position : </i>the part-of-speech and po­sition of a wrord• in the Korean lexical sample task</p><p><i>— morpheme-position : </i>a morpheme<footnote anchor="1"/> and its position.</p><p><i>— morpheme-POS </i>: a morpheme and its part-of-speech.</p><p><i>— POSjposition </i>: the part-of-speech and po­sition of a morpheme</p><p>In the English lexical sample task, <i>word </i>is a sur­face form and can be either one of open-class words whose POS is one of the noun, verb, adjective, and adverb: or one of closed-class words whose POS is</p><p>*A Korean sentence is composed of one or more <b><i>eojeols, </i></b>which are separated by spaces, and an <b><i>eojeol </i></b>consists of one or more morphemes.</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>feature</p></td><td class="cell"><p>normalized conditional probability(p^)</p></td><td class="cell"><p>DS</p></td><td class="cell"><p>MPC</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b><i>classi</i></b></p></td><td class="cell"><p><b><i>class2</i></b></p></td><td class="cell"><p><b><i>classs</i></b></p></td><td class="cell"><p><b><i>class4</i></b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>h</i></p></td><td class="cell"><p>0.7</p></td><td class="cell"><p>0.3</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>1.1187</p></td><td class="cell"><p><b><i>classi</i></b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>h</i></p></td><td class="cell"><p>0</p></td><td class="cell"><p>0.4</p></td><td class="cell"><p>0.6</p></td><td class="cell"><p>0</p></td><td class="cell"><p>1.0290</p></td><td class="cell"><p><b><i>classs</i></b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>h</i></p></td><td class="cell"><p>0</p></td><td class="cell"><p>0?4</p></td><td class="cell"><p>0.1</p></td><td class="cell"><p>0.5</p></td><td class="cell"><p>0.6390</p></td><td class="cell"><p><b><i>class±</i></b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>feature</p></td><td class="cell"><p>weight <i>(w^)</i></p></td><td class="cell"><p>weight (%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b><i>classi</i></b></p></td><td class="cell"><p><b><i>class2</i></b></p></td><td class="cell"><p><b><i>classs</i></b></p></td><td class="cell"><p><b><i>class4</i></b></p></td><td class="cell"><p><b><i>classi</i></b></p></td><td class="cell"><p><b><i>class2</i></b></p></td><td class="cell"><p><b><i>class?,</i></b></p></td><td class="cell"><p><b><i>class^</i></b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>h</i></p></td><td class="cell"><p>1.1187</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0.7831</p></td><td class="cell"><p>0.3356</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>h</i></p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>1.0290</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0.4116</p></td><td class="cell"><p>0.6174</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>h</i></p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0.6390</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0.2556</p></td><td class="cell"><p>0.0639</p></td><td class="cell"><p>0.3195</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>Rel</i><b><i>(classj, </i></b><i>X)</i></p></td><td class="cell"><p>1.1187</p></td><td class="cell"><p>0</p></td><td class="cell"><p>1.0290</p></td><td class="cell"><p>0.6390</p></td><td class="cell"><p>0.7831</p></td><td class="cell"><p>1.0028</p></td><td class="cell"><p>0.6813</p></td><td class="cell"><p>0.3195</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="4" global="150"/><p>• in the English lexical sample task</p><p>• in the Korean lexical sample task</p><p>one of the determiner, preposition, pronoun, and punctuation. The window size of ±3 words in the English lexical sample task and the window size from -2 to +3 word in the Korean lexical sample task were empirically chosen.</p><p>In the first phase of the experiments, we used just one complicated template, word.position.POS<i>(in </i>Korean <i>morpheme-position-POS), </i>which brought about data sparseness problem. So we split the tem­plate into three simpler templates.</p></subsubsection><subsubsection number="2.5.2" title="Topical context"><p>A topical context includes features of the following templates for all open-class words within its window:</p><p><i>- word </i>: an open-class word.</p><p><i>- morpheme </i>: an open-class morpheme.</p><p>The window size of ±1 sentences in the English lexical sample task and the window size of all sen­tences in the Korean lexical sample task were em­pirically chosen.</p></subsubsection><subsubsection number="2.5.3" title="Bigram context"><p>In a bigram context, there can be features of the fol­lowing templates for all word-pairs within its win­dow:</p><doubt alpha="64.0" length="50" tooSmall="False" monospace="0.0">- (wordi^wordj):  the z-th word and j-th word(i&gt;j)</doubt><doubt alpha="66.1" length="59" tooSmall="False" monospace="0.0">- (wordi,POSj):  the i-th word and j-th part-of-speech(i&gt;j)</doubt><p><i>- (eojeoli,eojeolj) </i>: the <i>i-th </i>eojeol and j-th eojeol <i>(i&gt;j)</i></p><p>Unlike local and topical contexts, bigram contexts are composed of only bigram information surround­ing the polysemous word. The window size of ±2 words in the English lexical sample task and the win­dow size from —2 to 4-3 word in the Korean lexical sample task were empirically chosen.</p></subsubsection></subsection></section><section number="3" title="Experimental Result"><p>The following tables show the results of our systems at Senseval-2 (Table 4). For the Korean lexical sample task at senseval-2, only fine-grained sense distinction was made.</p><p>Table 4: Results of KUNLP systems at Senseval-2</p></section><section number="4" title="Conclusion"><p>We have described the modified CIM used for word sense disambiguation at senseval-2. In the exper­iments, three types of features; local, topical, and bigram context, are used. Our system ranked as the highest at the Korean lexical sample task and as the topmost group at the English lexical sample task among the supervised models at Senseval-2. Consequently, the results back up the fact that the modified CIM and three types of features are useful for discriminating word senses.</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>task</p></td><td class="cell"><p>prec.</p></td><td class="cell"><p>recall</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>English Lexical Sample (fine g.)</p></td><td class="cell"><p>0.629</p></td><td class="cell"><p>0.629</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>English Lexical Sample (coarse g.)</p></td><td class="cell"><p>0.697</p></td><td class="cell"><p>0.697</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Korean Lexical Sample (fine g.)</p></td><td class="cell"><p>0.698</p></td><td class="cell"><p>0.74</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></section><references><p>Eric Brill 1994. Some advances in rule-based part of speech tagging. In <i>Proceedings of the Twelfth National Conference on Artificial Intel­ligence (ÀAAI</i><i>-94)-</i></p><p>Martin Chodorow, Claudia Leacock and George A. Miller 2000. A Topical/Local Classifier for Word Sense Identification. In <i>Computers and the Hu­manities </i><i>34' </i><i>115</i><i>-120.</i></p><p>Ho Lee, Dae-Ho Baek and Hae-Chang Rim 1997. Word Sense Disambiguation Based on The In­formation Theory. In <i>Proceedings of Research on Computational Linguisitcs Conference.</i></p><p>Ho Lee, Hae-Chang Rim and JungYun Seo 2000. Word Sense Disambiguation Using the Classifica­tion Information Model. In <i>Computers and the Humanities </i><i>34' 141-146-</i></p></references></body></article>