<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="440"/><title>YSC-DSAA: An Approach to Disambiguate Sentiment Ambiguous Adjectives Based on SAAOL</title><pubinfo>î "MS Proceedings of the 5th International Workshop on Semantic Evaluation, ACL 2010,pages 440^43, Uppsala, Sweden, 15-16 July 2010. ©2010 Association for Computational Linguistics</pubinfo><author surname="Yang" givenname="Shi-Cai"><org  name="Brno University of Technology" country="Czech Republic" city="Brno"/></author><author surname="Liu" givenname="Mei-Juan"><org  name="Brno University of Technology" country="Czech Republic" city="Brno"/></author></firstpageheader><frontmatter><p><b>YSC-DSAA: An Approach to Disambiguate Sentiment Ambiguous</b></p><p><b>Adjectives Based On SAAOL</b></p><p><b>Shi-Cai Yang</b></p><p>Ningbo University of Technology Ningbo, Zhejiang, China nbcysc@12 6.com</p><p><b>Mei-Juan Liu</b></p><p>Zhejiang Ocean University Zhoushan, Zhejiang, China</p><p>azalea1212@126.com</p></frontmatter><abstract>In this paper, we describe the system we de­veloped for the SemEval-2010 task of disam-biguating sentiment ambiguous adjectives (hereinafter referred to SAA). Our system cre­ated a new word library named SAA-Oriented Library consisting of positive words, negative words, negative words related to SAA, posi­tive words related to SAA, and inverse words, etc. Based on the syntactic parsing, we ana­lyzed the relationship between SAA and the keywords and handled other special processes by extracting such words in the relevant sen­tences to disambiguate sentiment ambiguous adjectives. Our micro average accuracy is 0.942, which puts our system in the first place. </abstract></header><body><section number="1" title="Introduction"><p>We participated in disambiguating sentiment ambiguous adjectives task of SemEval-2010(Wu and Jin, 2010).</p><p>Together 14 sentiment ambiguous adjectives are chosen by the task organizers, which are all high-frequency words in Mandarin Chinese. They are: y^|big, /h|small, #|many, ^|few, rft |high, {Ë |low, W |thick, <b><i>M</i></b><b><i> </i></b>|thin, <b>$H </b>|deep, B |shallow, S |heavy, $5 |light, |ÊL;^|huge, <b><i>SJ\ </i></b>|grave. These adjectives are neutral out of con­text, but when they co-occur with some target nouns, positive or negative emotion will be evoked. The task is designed to automatically determine the semantic orientation of these sen­timent ambiguous adjectives within context: positive or negative (Wu and Jin, 2010). For in­stance, "if<b>rfê</b>-rt|the price is high" indicates negative meaning, while ^|the quality is high" has positive connotation.</p><p>Considering the grammar system of contem­porary Chinese, a word is one of the most basic linguistic granularities consisting of a sentence. Therefore, as for the sentiment classification of a sentence, the sentiment tendency of a sentence can be identified on the basis of that of a word. Wiebe et al. (2004) proposed that whether a sen­tence is subjective or objective should be dis­criminated according to the adjectives in it. On the basis of <i>General Inquirer Dictionary, A Learner's Dictionary of Positive and Negative Words, HowNet, A Dictionary of Positive Words and A Dictionary of Negative Words </i>etc., Wang et al.(2009) built a word library for Chinese sen­timent words to discriminate the sentiment cate­gory of a sentence using the weighted linear combination method.</p><p>Unlike the previous researches which have not taken SAA into consideration specially in dis­criminating the sentiment tendency of a sentence, in the SemEval-2010 task of disambiguating sen­timent ambiguous adjectives, systems have to predict the sentiment tendency of these fourteen adjectives within specific context.</p><p>From the view of linguistics, first we devel­oped a SAA-oriented keyword library, then ana­lyzed the relationship between the keywords in the clauses and SAA, and classified its positive or negative meaning of SAA by extracting the clauses related to SAA in the sentence.</p></section><section number="2" title="SAAOL"><p>We create a SAA-oriented library marked as SAAOL which is made up of positive and nega­tive words irrelevant to context, negative words related to SAA (NSAA), positive words related to SAA (PSAA), and inverse words. The above five categories of words are called keywords for short in the paper.</p><p>Positive and negative words irrelevant to con­text refer to the traditional positive or negative words which are gathered from <i>The Dictionary</i> <i>of Positive </i>Words(Shi, 2005), <i>The Dictionary of Negative </i>Words(Yang, 2005), HowNet <footnote anchor="1"/> and other network resources, such as Terms of Ad­verse Drug Reaction, Codes of Diseases and Symptoms, etc.<page local="2" global="441"/></p><p>Distinguishing from the traditional positive and negative words, NSAA and PSAA in our SAAOL refer to those positive and negative words which are related to SAA, yet not classi­fied into the positive and negative words irrele­vant to context mentioned above.</p><p>We divide SAA into two categories: A cate­gory and B category listed in Table 1.</p><p>We identify whether a word belongs to NSAA or not on the following principle: any words when used with A category are negative; con­versely, when used with B category, they are positive.</p><p>For example, in the following clauses, "}'É'MFÏ|oil prices are high" , "Sf{3:ÎË^|the responsibility is important", "{î^4IIÎË| the task is very heavy", "If^Ä^H^the workload is very large" , "^ifMoil prices", "fHî|responsibility", "ft #|task", "I^S|workload" are NSAA.</p><p>Correspondingly, we identify whether a word belongs to PSAA or not on the following princi­ple: any words when used with A category are positive; however, when used with B category, they are negative. In the clauses, much food",</p><p>efficiency is extremely low" , "#S^0$iI] [interest rate on deposit is high", "food", "W| efficiency", " |interest rate on deposit" are PSAA.</p><p>In general, when two negative words are used together, the sentiment tendency that they show is negative. For instances, " |incidence of diabetes", " ^î|]Ë§J?c|virus infec­tion", "i^^5jji^|destruction of wars". However, in certain cases, some words play a part in elimi­nating negative meaning when used with nega­tive words, for example, " Jx. |anti-" , " <b><i>fflM </i></b>|restrain", " |avoid", " K |resist", " |reduce", "F$ipS|fell", "M4&gt;|decrease", "fèflj |control", " Jt*|cost", " Suppose", " <b><i>TM</i></b><b><i> </i></b>|decrease", " # |non-", " |not". These special words are called inverse words in our SAAOL.</p><footnote label="1"> http://www.keenage.com .</footnote><p>In the following instances, "MfefëS|reduce the injury", " »JffiJH curb inflation", "S$4 |anti-war", the words "ffif||injury", infla­tion", and " ^^|war" themselves are all nega­tive. When used with the inverse words" M$r |reduce ", curb", "Jx|anti-", they express positive meaning instead.</p><p>On the basis of the above collected word li­brary, we discriminate manually the positive and negative meaning, PSAA, NSAA, and inverse words in 50,000 Chinese words according to <i>Richard Xiao's Top 50,000 Chinese Word Fre­quency List, </i>which collects the frequency of the top 50000 Chinese words covered in the just published frequency dictionary of Mandarin Chinese based on a balanced corpus of ca. 50 million words. The list is available at http://www.lancs.ac.uk/fass/projects /corpus/data/top50000 Chinese words. zip.</p><p>Based on HowNet lexical semantic similarity computing(Liu, 2002), Yang and Wu(2009) se­lected the new positive and negative bench-mark words to identify the sentiment tendency by adopting the improved benchmark words and the modified method of computing similarity be­tween words and benchmark words. Their ac­curacy rate arrived at 98.94%.</p><p>In light of the errors of manual calibration, we extended the keywords in SAAOL by applying Yang and Wu's (2009) method and added syn­onymic and antonymous words in it. Eventually we proofread and revised manually the new ex­tended keywords.</p></section><section number="3" title="Our method"><p>According to the structural characteristics of the sentence, the sentence can be divided into simple sentences and complex sentences. A simple sen­tence consists of a single clause which contains a subject and a predicate and stands alone as its own sentence. However, a complex sentence is the one which is linked by conjunctions or consists of at least two or more clauses without any conjunctions in it.<page local="3" global="442"/></p><table caption="Table 1: SAA Classification Table" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>A category</p></td><td class="cell"><p>B category</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>^|big</p></td><td class="cell"><p>/h|small</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>#|many</p></td><td class="cell"><p>4&gt;|few</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>rft|high</p></td><td class="cell"><p>{Ë|low</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>J5|thick</p></td><td class="cell"><p>i||thin</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>$l|deep</p></td><td class="cell"><p>^§|shallow</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>JB.|heavy</p></td><td class="cell"><p>flight</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>MA|huge</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>MÀ|grave</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>A complicated sentence in structure is divided into several clauses in accordance with punctua­tions, such as a full stop, or a exclamatory mark, or a comma, or a semicolon, etc. We analyze the syntax of the clause by extracting the clause in­cluding SAA and the adjacent one. We extract SAAOL keywords in the selected clauses, and then analyze the grammatical relationship be­tween the keywords and SAA.</p><p>Wang et al's research of extraction technology based on the dependency relation of Chinese sen­timental elements indicated that the dependency analyzer designed by Stanford University had not showed a high rate of accuracy. And the wrong dependency relation will interfere with the sub­sequent parsing process seriously (Wang, et al., 2009).</p><p>Taking the above factors into consideration, we have not analyzed the dependency relation at present. Through studying abundant instances, we specialize in the structural relationship be­tween the keywords and SAA to extract the rela­tion patterns which have a higher occurrence fre­quency. In the meantime, inverse words are proc­essed particularly. Eventually we supplemented modification of the inaccuracy of automatic segmented words and some special adverbs, such as fe| prejudiced, Ü|excessive, ;fe|too.</p><p>To sum up, based on the word library SAAOL and structural analysis, SAA classification pro­cedures are as follows:</p></section><section number="4" title="Evaluation"><p>In disambiguating sentiment ambiguous adjec­tives task of SemEval-2010, there are 2917 in­stances in test data for 14 Chinese sentiment am­biguous adjectives. According to the official re­sult of the task, our micro average accuracy is 0.942, which puts our system in the first position among the participants.</p><p>Depending upon the answers from organizers of the task, we notice that errors occur mainly in the following cases.</p><p>Firstly, there is a key word related to SAA, but it has no such key word in our SAAOL.</p><p>For instance,</p><p>Âtt^lWÈll pf fê/B$m&lt;head&gt;M &lt;/head&gt;t&gt;H | Why is the usage rate of pf so high in my computer?</p><p>"pf^ffi$|The usage rate of pf" should be NSAA, but it does not exist in our SAAOL.</p><p>Secondly, the sentence itself is too compli­cated to be analyzed effectively in our system so far.</p><p>Thirdly, as the imperfection of SAAOL itself, there are some inevitable mistakes in it. For instance,</p><p>&amp;fë &lt;head&gt; X</p><p>&lt;/head&gt; | The diver's feat is extremely difficult.</p><p>It is generally known that if the bigger the dif­ficulty of the dive is, the better the diver's per­formance will be, both of which are of propor­tional relation. However, generally speaking, the degree of difficulty is negative. For this reason, we made a mistake in such instance.</p><p>Step 1 Extract unidentified clauses in­cluding SAA;</p><p>Step 2 Extract the keywords in SAAOL from the clause;</p><p>Step 3 Label the sentiment tendency of each sentiment word by using SAAOL;</p><p>Step 4 Discriminate the positive or nega­tive meaning of a sentence in accordance with the different relationships. If there are no keywords in the sentence, perform step 5; otherwise, discrimination is over.</p><p>Step 5 Extract the clauses next to SAA, and identify them according to Steps 2-4. If there are no extractable clauses, mark them as SAA which will be recognized. A is for the positives, and B for the negatives.</p></section><section number="5" title="Conclusions"><p>In this paper, we describe the approach taken by our systems which participated in the disambigu-ating sentiment ambiguous adjectives task of</p><p>SemEval-2010.</p><p>We created a new word library named SAAOL. Through gathering words from relative dictionaries, HowNet, and other network re­sources, we discriminated manually the positive and negative meaning, PSAA, NSAA, and in­verse words in 50,000 Chinese words according to <i>Richard Xiao's Top 50,000 Chinese Word Frequency List. </i>And then we extended the key­words in SAAOL by applying Yang's (2009) method and added synonymic and antonymous words in it. Eventually the new extended key­words were proofread and revised manually.</p><p>Based on SAAOL and structural analysis, we describe a procedure to disambiguate sentiment ambiguous adjectives.<page local="4" global="443"/> Evaluation results show that this approach achieves good performance in the task.</p></section><references><p>Qun Liu, JianSu Li. 2002. Calculation of semantic similarity of words based on the HowNet. <i>The Third Chinese Lexical Semantics Workshop. </i>Tai Bei.</p><p>Jilin Shi, Yinggui Zhu. 2005. <i>A Dictionary of Posi­tive Words. </i>Lexicographical Publishing House, Chengdu, Sichuan.</p><p>Su Ge Wang, An Na Yang, De Yu Li. 2009. Research on sentence sentiment classification based on Chi­nese sentiment word table. <i>Computer Engineer­ing and Applications, </i>45(24):153-155.</p><p>Qian Wang, TingTing He, et al. 2009. Research on dependency Tree-Based Chinese sentimental ele­ments extraction, <i>Advances of Computational Linguistics in China, </i>624-629.</p><p>Janyce Wiebe, Theresa Wilson, et al. 2004. Learning subjective language. <i>Computational Linguistics, </i>30(3): 277-308.</p><p>Yunfang Wu, Peng Jin. SemEval-2010 task 18: Dis-ambiguating sentiment ambiguous adjectives.</p><p>Yu bing Yang, Xian wei Wu. Improved lexical se­mantic tendentiousness recognition computing. 2009. <i>Computer Engineering and Applications,</i> 45(21): 91-93.</p><p>Ling Yang, Yinggui Zhu. 2005. <i>A Dictionary of Negative Words. </i>Lexicographical Publishing House, Chengdu, Sichuan.</p></references></body></article>