<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="292"/><title>CityU-DAC: Disambiguating Sentiment-Ambiguous Adjectives within Context</title><pubinfo>Proceedings of the 5th International Workshop on Semantic Evaluation, ACL 2010,pages 292-295, Uppsala, Sweden, 15-16 July 2010. ©2010 Association for Computational Linguistics</pubinfo><author surname="Lu" givenname="Bin"><org  name="City University of Hong Kong" country="Hong Kong" city="Kowloon"/></author><author surname="Tsou" givenname="Benjamin K."><org  name="University of Hong Kong" country="Hong Kong" city="Pokfulam"/></author></firstpageheader><frontmatter><p><b>CityU-DAC: Disambiguating Sentiment-Ambiguous Adjectives within</b></p><p><b>Context</b></p><p><b>Bin LU and Benjamin K. TSOU</b></p><p>Department of Chinese, Translation and Linguistics &amp; Language Information Sciences Research Centre City University of Hong Kong {lubin2010, rlbtsou}@gmail.com</p></frontmatter><abstract>This paper describes our system participating in task 18 of SemEval-2010, i.e. disambiguating Sentiment-Ambiguous Adjectives (SAAs). To disambiguating SAAs, we compare the machine learning-based and lexicon-based methods in our submissions: 1) Maximum entropy is used to train classifiers based on the annotated Chinese data from the NTCIR opinion analysis tasks, and the clause-level and sentence-level classifiers are compared; 2) For the lexicon-based method, we first classify the adjectives into two classes: intensifiers (i.e. adjectives intensifying the intensity of context) and suppressors (i.e. adjectives decreasing the intensity of context), and then use the polarity of context to get the SAAs' contextual polarity based on a sentiment lexicon. The results show that the performance of maximum entropy is not quite high due to little training data; on the other hand, the lexicon-based method could improve the precision by considering the polarity of context. </abstract></header><body><section number="1" title="Introduction"><p>In recent years, <i>sentiment analysis, </i>which mines opinions from information sources such as news, blogs, and product reviews, has drawn much attention in the NLP field (Hatzivassiloglou and McKeown, 1997; Pang et al., 2002; Turney, 2002; Hu and Liu, 2004; Pang and Lee, 2008). It has many applications such as social media monitoring, market research, and public relations.</p><p>Some adjectives are neutral in sentiment polarity out of context, but they could show positive, neutral or negative meaning within specific context. Such words can be called dynamic sentiment-ambiguous adjectives (SAAs). However, SAAs have not been intentionally tackled in the researches of sentiment analysis, and usually have been discarded or ignored by most previous work. Wu et al., (2008) presents an approach of combining collocation information and SVM to disambiguate SAAs, in which the collocation-based method was first used to disambiguate adjectives within the context of collocation (i.e. a sub-sentence marked by comma), and then the SVM algorithm was explored for those instances not covered by the collocation-based method. According to their experiments, their supervised algorithm achieves encouraging performance.</p><p>The task 18 at SemEval-2010 is intended to create a benchmark dataset for disambiguating SAAs. Given only 100 trial sentences, but not provided with any official training data, participants are required to tackle this problem data by unsupervised approaches or use their own training data. The task consists of 14 SAAs, which are all high-frequency words in Mandarin Chinese. They are y^jbig, /hjsmall, #|many, <i>;P </i>|few, rftjhigh, {Ëjlow, fl^jthick, #|thin, ^jdeep, ^shallow, J!|heavy, gjlight, lÊL^jhuge, J!;^ |grave. This task deals with Chinese SAAs, but the disambiguating techniques should be language-independent. Please refer to (Wu and Jin, 2010) for more descriptions of the task.</p><p>In our participating system, the annotated Chinese data from the NTCIR opinion analysis tasks is used as training data with the help of a combined sentiment lexicon. A machine learning-based method (namely maximum entropy) and the lexicon-based method are compared in our submissions. The results show that the performance of maximum entropy is not quite high due to little training data; on the other hand, the lexicon-based method could improve the precision by considering the context of SAAs.<page local="2" global="293"/> In Section 2, we briefly describe data preparation of sentiment lexicon and training data. Our approaches for disambiguating SAAs are given in Section 3. The experiment and results are presented in Section 4, followed by a conclusion in Section 5.</p></section><section number="2" title="Data Preparation"><subsection number="2.1" title="Sentiment Lexicon"><p>Several traditional Chinese resources of polar words/phrases are collected, including NTU Sentiment Dictionary<footnote anchor="1"/>, <i>The Lexicon of Chinese Positive Words </i>(Shi and Zhu, 2006), <i>The Lexicon of Chinese Negative Words </i>(Yang and Zhu, 2006) 0, and CityU's sentiment-bearing word/phrase list (Lu et al, 2008), which were manually marked in the political news data by trained annotators (Benjamin and Lu, 2008). Sentimentbearing items marked with the <i>SENTIMENT_KW</i> tag (SKPI), including only positive and negative items but not neutral ones, were also automatically extracted from the Chinese sample these polar item lexicons were combined, and the combined polar item lexicon consists of 13,437 positive items and 18,365 negative items, a total of 31,802 items.</p><doubt alpha="57.8" length="45" tooSmall="False" monospace="0.0">data of NTCIR-6 OAPT (Seki et al., 2007). All</doubt></subsection><subsection number="2.2" title="Training Data"><p>The training data is extracted from the Chinese sample and test data from the NTCIR opinion analysis task, including NTCIR-6 (Seki et al., 2007), NTCIR-7 (Seki et al., 2008) and NTCIR-8 (Seki et al., 2010). The NTCIR opinion analysis tasks provide an opportunity to evaluate the techniques used by different participants based on a common evaluation framework in Chinese (simplified and traditional), Japanese and English.</p><p>For data from NTCIR-6 and NTCIR-7, three annotators manually marked the polarity of each opinionated sentence, and the lenient polarity is used here as the gold standard (please refer to Seki et al., 2008 for explanation of lenient and strict standard). For each opinionated sentence from NTCIR-8, only two annotators marked and the strict polarity is used as the gold standard. The traditional Chinese sentences are transferred into simplified Chinese. In total, there are about 12K opinionated sentences annotated with polarity, out of which about 9K are marked as positive or negative, and others neutral. All the 9K sentences plus the 100 sentences from the trial data are used as the sentence-level training data.</p><footnote label="1"> http://nlg18.csie.ntu.edu.tw:8080/opinion/index.html</footnote><p>Meanwhile, we also try to get the clause-level training data since the context of collocation within sub-sentences are quite crucial for disambiguating SAAs according to Wu et al. (2008). From the 9K positive/ negative sentences above, we automatically extract the clause for each occurrence of SAAs.</p><p>Note the polarity for a whole sentence is not necessarily the same with that of the clause containing SAAs. Consider the sentence <i>&amp;</i></p><p><i>û</i><b><i>ù </i></b><i>tittf</i><i> </i><i>3</i><i> </i><i>MM</i><i> </i><i>#</i><i> </i><i>,</i><i> </i><i>#m</i><i> </i><i>MB</i><i> </i><i>tâ5</i><i> </i><i>3</i><b><i>c&amp;</i></b></p><p><i>(In the current large circumstance of the world, China and Russia support each other). </i>The polarity of the whole sentence is positive, while the clause <i>&amp;È</i><i>H</i><i>Ê</i><i>î</i><i>fffii</i><i>â</i><i>j?-j:M-M</i><i>tP </i><i>(In the current large circumstance of the world) </i>containing a SAA <i>3 </i><i>(large) </i>is neutral, and the polarity lies in the second part of the whole sentence, i.e. <i>f</i><i>B</i><i>S </i>(support each other). Thus, we manually checked the polarity of clauses containing SAAs. Due to time limitation, we only checked 465 clauses. Plus the clauses extracted from 100 trial sentences, the final clause-level training data consist of 565 positive/negative clauses containing SAAs.</p></subsection></section><section number="3" title="Our Approach for Disambiguating"><doubt alpha="100.0" length="4" tooSmall="False" monospace="0.0">SAAs</doubt><p>To disambiguating SAAs, we use the maximum entropy algorithm and the sentiment lexicon-based method, and also combine them together.</p><subsection number="3.1" title="The Maximum Entropy-based Method"><p>Maximum entropy classification (MaxEnt) is a technique which has proven effective in a number of natural language processing applications (Berger et al., 1996). Le Zhang's maximum entropy tool<footnote anchor="2"/> is used for classification.</p><p>The Chinese sentences are segmented into words using a production segmentation system. Unigrams of words are used as basic features for classification. Bigrams are also tried, but does not show improvement, and thus are not described in details here.</p></subsection><subsection number="3.2" title="The Lexicon-based Method"><p>For the lexicon-based method, we first classify the 14 adjectives into two classes: intensifiers and suppressors.<page local="3" global="294"/> Intensifiers refer to adjectives intensifying the intensity of context, including ^ |big, # |many, M |high, W |thick, W |deep, H |heavy, lÊL^huge, H^grave, while suppressors refer to adjectives decreasing the intensity of context, including /h|small, <i>;P </i>|few, {Ë|low, <i>M </i>|thin, ^shallow, g|light.</p><footnote label="2"> http://  homepages.inf.ed.ac.uk/lzhang10/maxent_toolkit.html</footnote><p>Meanwhile, the collocation nouns are also classified into two classes: positive and negative. Positive nouns include M |quality, |^ III |standard, 7R^P |level, |benefit, JsSogfc |achievement, etc. Negative nouns include <i>Bi</i><i>^J </i>|pressure, lîllSËî |disparity, [BJj|_ |problem, Jxlßä |risk, f ^pollution etc.</p><p>The hypothesis here is that intensifiers will receive the polarity of their collocations while suppressors will get the opposite polarity of their collocations. For example, JsSolsfc |achievement could be collocated with one of the following intensifiers: y^|big, #|many or M|high, and the adjectives just receive the polarity of JsSolsfc |achievement, which is positive. Meanwhile, <i>ff</i><i> </i>^pollution could be collocated with one of the following suppressors: /h|small, ^|few, {Ë|low, and the adjectives just receive the opposite polarity of ffl^pollution, which is also positive.</p><p>Based on this hypothesis, we could get the polarity of SAAs through theirs collocation nouns within the clauses containing SAAs. The context of SAAs is a sub-sentence marked by comma. The sentiment lexicon mentioned in Section 2.1 is used to find polarity of collocation nouns.</p></subsection><subsection number="3.3" title="Combining   Maximum   Entropy and Lexicon"><p>To combine the two methods above, the lexicon-based method is first used to disambiguate the sentiment of SAAs, and the context of collocation is a sub-sentence marked by comma. Then for those instances that are not covered by the lexicon-based method, the maximum entropy algorithm is explored.</p></subsection></section><section number="4" title="Experiment and Results"><p>The dataset contains two parts: some sentences were extracted from Chinese Gigaword (LDC corpus: LDC2005T14), and other sentences were gathered through the search engine like Google. Firstly, these sentences were automatically segmented and POS-tagged, and then the ambiguous adjectives were manually annotated with the correct sentiment polarity within the sentence context. Two annotators annotated the sentences double blindly, and the third annotator checks the annotation. All the data of 2,917 sentences is provided as the test set, and evaluation is performed in terms of micro accuracy and macro accuracy.</p><p>We submitted 4 runs: run 1 is based on the sentence-level MaxEnt classifier; run 2 on the clause-level MaxEnt classifier; run 3 is got by combining the lexicon-based method and the sentence-level MaxEnt classifier; and run 4 by combining the lexicon-based method and the clause-level MaxEnt classifier. The official scores for the 4 runs are shown in Table 2.</p><table caption="Table 2. Results of 4 Runs"></table><p>From Table 2, we can observe that:</p><p>1) Compared the highest scores achieved by other teams, the performance of maximum entropy (run 1 and 2) is not quite high due to little training data;</p><p>2) By integrating the lexicon-based method and maximum entropy (run 3 and 4), we improve the accuracy by considering the context of SAAs;</p><p>3) The sentence-level maximum entropy classifier shows better macro accuracy, and clause-level one better micro accuracy.</p><p>In addition to the official scores, we also evaluate the performance of the lexicon-based method alone. The micro and macro accuracy are respectively 0.847 and 0.835665, showing that the lexicon-based method is more accurate than the maximum entropy algorithm (run 1 and 2).</p><doubt alpha="47.7" length="44" tooSmall="False" monospace="0.0">But it only covers 1,436 (49%) of 2,917 test</doubt><p>instances.</p><p>Because the data from the NTCIR opinion analysis task is not specifically annotated for this task, and the manually checked clauses are less than 600, the performance of our system is not quite high compared to the highest performance achieved by other teams.</p></section><section number="5" title="Conclusion"><p>To disambiguating SAAs, we compare machine learning-based and lexicon-based methods in our submissions: 1) Maximum entropy is used to train classifiers based on the annotated Chinese data from the NTCIR opinion analysis tasks, and the clause-level and sentence-level classifiers are compared; 2) For the lexicon-based method, we first classify the adjectives into two classes:<page local="4" global="295"/> intensifiers (i.e. adjectives intensifying the intensity of context) and suppressors (i.e. adjectives decreasing the intensity of context), and then use the polarity of context to get the SAAs' contextual polarity. The results show that the performance of maximum entropy is not quite high due to little training data; on the other hand, the lexicon-based method could improve the precision by considering the context of</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Run</p></td><td class="cell"><p>Micro Acc. (%)</p></td><td class="cell"><p>Macro Acc. (%)</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>1</p></td><td class="cell"><p>61.98</p></td><td class="cell"><p>67.89</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>2</p></td><td class="cell"><p>62.63</p></td><td class="cell"><p>60.85</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>3</p></td><td class="cell"><p>71.55</p></td><td class="cell"><p>75.54</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>4</p></td><td class="cell"><p>72.47</p></td><td class="cell"><p>69.80</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>SAAs.</p></section><references><p>Adam L. Berger, Stephen A. Della Pietra, and Vincent J. Della Pietra. 1996. A maximum entropy approach to natural language processing. <i>Computational Linguistics, </i>22(1):39-71.</p><p>Vasileios Hatzivassiloglou and Kathleen McKeown. 1997. Predicting the Semantic Orientation of Adjectives. <i>Proceedings of ACL-97. </i>174-181.</p><p>Minqing Hu and Bing Liu. 2004. Mining Opinion Features in Customer Reviews. In <i>Proceedings of the 19th National Conference on Artificial Intelligence, </i>pp. 755-760.</p><p>Bin Lu, Benjamin K. Tsou and Oi Yee Kwong. 2008. Supervised Approaches and Ensemble Techniques for Chinese Opinion Analysis at NTCIR-7. In <i>Proceedings of the Seventh NTCIR Workshop (NTCIR-7).</i><i> </i>pp. 218-225. Tokyo, Japan.</p><p>Bo Pang and Lillian Lee. 2008. <i>Opinion mining and sentiment analysis, Foundations and Trends in Information Retrieval, </i>Now Publishers.</p><p>Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002. Thumbs up? Sentiment classification using machine learning techniques. In <i>Proceedings of EMNLP 2002, </i>pp.79-86.</p><p>Yohei Seki, David Kirk Evans, Lun-Wei Ku, Le Sun, Hsin-His Chen, Noriko Kando. 2007. Overview of</p><p>Opinion Analysis Pilot Task at NTCIR-6. <i>Proc. of</i> <i>the Seventh NTCIR Workshop.</i><i> </i>Japan. 2007.6.</p><p>Yohei Seki, David Kirk Evans, Lun-Wei Ku, Le Sun, Hsin-His Chen, Noriko Kando and Chin-Yew Lin. 2008. Overview of Multilingual Opinion Analysis Task at NTCIR-7. <i>Proc. of the Seventh NTCIR Workshop. </i>Japan. Dec. 2008.</p><p>Yohei Seki, Lun-Wei Ku, Le Sun, Hsin-His Chen, Noriko Kando. 2010. Overview of Multilingual Opinion Analysis Task at NTCIR-8. <i>Proc. of the</i> <i>Seventh NTCIR Workshop.</i><i> </i>Japan. June, 2010.</p><p>Jilin Shi and Yinggui Zhu. 2006. The Lexicon of Chinese Positive Words ( lSlip?]p?]jft). Sichuan Lexicon Press.</p><p>Benjamin K. Tsou and Bin Lu. 2008. A Political News Corpus in Chinese for Opinion Analysis. In <i>Proceedings of the Second International Workshop on Evaluating Information Access (EVIA2008).</i><i> </i>pp. 6-7. Tokyo, Japan.</p><p>Peter D. Turney. 2002. Thumbs up or thumbs down? Semantic orientation applied to unsupervised classification of reviews, In <i>Proceedings of ACL-02, </i>Philadelphia, Pennsylvania, 417-424.</p><p>Yunfang Wu, Miao Wang, Peng Jin and Shiwen Yu. 2008. Disambiguate sentiment ambiguous adjectives. In <i>Proceedings of IEEE International Conference on Natural Language Processing and Knowledge Engineering (NLP-KE'08).</i></p><p>Yunfang Wu, and Peng Jin. 2010. SemEval-2010 Task 18: Disambiguate sentiment ambiguous adjectives. In <i>Proceedings of SemEval-2010.</i></p><p>Ruifeng Xu, Kam-Fai Wong and Yunqing Xia. 2008. Coarse-Fine Opinion Mining - WIA in NTCIR-7 MOAT Task. In <i>Proceedings of the Seventh NTCIR</i> <i>Workshop (NTCIR-7).</i><i> </i>Tokyo, Japan, Dec. 16-19.</p><p>Ling Yang and Yinggui Zhu. 2006. The Lexicon of Chinese Negative Words (IJIip?]p?]jft). Sichuan Lexicon Press.</p></references></body></article>