<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="59"/><title>SemEval-2007 Task 12: Turkish Lexical Sample Task</title><pubinfo>Proceedings of the 4th International Workshop on Semantic Evaluations (SemEval-2007),pages 59-63, Prague, June 2007. ©2007 Association for Computational Linguistics</pubinfo><author surname="Orhan" givenname="Zeynep"><org  name="Fatih University" country="Turkey" city="Istanbul"/></author><author surname="Çelik" givenname="Emine"><org  name="Fatih University" country="Turkey" city="Istanbul"/></author><author surname="Neslihan" givenname="Demirgüç"><org  name="Fatih University" country="Turkey" city="Istanbul"/></author></firstpageheader><frontmatter><p><b>SemEval-2007 Task 12: Turkish Lexical Sample Task</b></p><p><b>Zeynep Orhan</b></p><p>Department of Computer Engineering, Fatih University 34500, Bûyûkçekmece, Istanbul, Turkey</p><p>zorhan@fatih.edu.tr</p><p><b>Emine Çelik</b></p><p>eminemm@gmail.com</p><p><b>Neslihan Demirgüç</b></p><p>Department of Computer Engineering, Fatih University 34500, Büyü^ekmece, Istanbul, Turkey</p><p>nesli_han@hotmail.com</p></frontmatter><abstract>This paper presents the task definition, re­sources, and the single participant system for Task 12: Turkish Lexical Sample Task (TLST), which was organized in the Se-mEval-2007 evaluation exercise. The methodology followed for developing the specific linguistic resources necessary for the task has been described in this context. A language-specific feature set was defined for Turkish. TLST consists of three pieces of data: The dictionary, the training data, and the evaluation data. Finally, a single system that utilizes a simple statistical method was submitted for the task and evaluated. </abstract></header><body><section number="1" title="Introduction"><p>Effective parameters for word sense disambigua­tion (WSD) may vary for different languages and word types. Although, some parameters are com­mon in many languages, some others may be lan­guage specific. Turkish is an interesting language that deserves being examined semantically. Turk­ish is based upon suffixation, which differentiates it sharply from the majority of European languages, and many others. Like all Turkic languages, Turk­ish is agglutinative, that is, grammatical functions are indicated by adding various suffixes to stems. Turkish has a SOV (Subject-Object-Verb) sentence structure but other orders are possible under certain discourse situations. As a SOV language where objects precede the verb, Turkish has postpositions rather than prepositions, and relative clauses that precede the verb. Turkish, as a widely-spoken lan­guage, is appropriate for semantic researches.</p><p>TLST utilizes some resources that are explained in Section 2-5. In Section 6 evaluation of the sys­tem is provided. In section 7 some concluding re­marks and future work are discussed.</p></section><section number="2" title="Corpus"><p>Lesser studied languages, such as Turkish suffer from the lack of wide coverage electronic re­sources or other language processing tools like on­tologies, dictionaries, morphological analyzers, parsers etc. There are some projects for providing data for NLP applications in Turkish like METU Corpus Project (Oflazer et al., 2003). It has two parts, the main corpus and the treebank that con­sists of parsed, morphologically analyzed and dis-ambiguated sentences selected from the main cor­pus, respectively. The sentences are given in XML format and provide many syntactic features that can be helpful for WSD. This corpus and treebank can be used for academic purposes by contract.</p><p>The texts in main corpus have been taken from different types of Turkish written texts published in 1990 and afterwards. It has about two million words. It includes 999 written texts taken from 201 books, 87 papers and news from 3 different Turk­ish daily newspapers. XML and Text Encoding Initiative (TEI) style annotation have been used. The distribution of the texts in the Treebank is similar to the main corpus. There are 6930 sen­tences in this Treebank. These sentences have been parsed, morphologically analyzed and disambigu-ated. In Turkish, a word can have more than one analysis, so having disambiguated texts is very important.</p><page local="2" global="60"/><doubt alpha="54.3" length="46" tooSmall="False" monospace="0.0">&lt;?xml version="1.0" encoding="windows-1254" ?&gt;</doubt><doubt alpha="60.0" length="20" tooSmall="False" monospace="0.0">-&lt;Set sentences="1"&gt;</doubt><doubt alpha="27.3" length="11" tooSmall="False" monospace="0.0">-&lt;S No="1"&gt;</doubt><doubt alpha="45.8" length="59" tooSmall="False" monospace="0.0">&lt;W IX="1"LEM="" MORPH="" IG="[(1,"soguk+Adj")(2,"Adv+Ly")j"</doubt><doubt alpha="55.6" length="99" tooSmall="False" monospace="0.0">REL="[2,1,(MODIFIER)j"&gt;Sogukça&lt;/W&gt; &lt;W IX="2"LEM="" MORPH="" IG="[(1,"yanitla+Verb+Pos+Past+A1sg")j"</doubt><doubt alpha="44.6" length="101" tooSmall="False" monospace="0.0">REL="[3,1,(SENTENCE)j "&gt;yanitladim&lt;/W&gt; &lt;W IX="3"LEM="" MORPH="" IG="[(1,".+Punc")j"REL="[,( )]"&gt;.&lt;/W&gt;</doubt><doubt alpha="25.0" length="4" tooSmall="False" monospace="0.0">&lt;/S&gt;</doubt><doubt alpha="42.9" length="7" tooSmall="False" monospace="0.0">_&lt;/Set&gt;</doubt><figure caption="Figure 1: XML file structure of the Treebank"></figure><table caption="Table 1: Target words in the SEMEVAL-1 Turkish Lexical Sample task" class="main" frame="box" rules="all" border="0" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p><b>Main English</b></p></td><td class="cell"><p><b>#</b></p></td><td class="cell"><p></p></td><td class="cell"><p><b>Train</b></p></td><td class="cell"><p><b>Test</b></p></td><td class="cell"><p><b>Total #of</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Words</b></p></td><td class="cell"><p><b>translation</b></p></td><td class="cell"><p><b>Senses</b></p></td><td class="cell"><p><b>MFS</b></p></td><td class="cell"><p><b>size</b></p></td><td class="cell"><p><b>size</b></p></td><td class="cell"><p><b>instances</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Nouns</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ara</p></td><td class="cell"><p>distance, break, interval, look for</p></td><td class="cell"><p>7</p></td><td class="cell"><p>53</p></td><td class="cell"><p>192</p></td><td class="cell"><p>63</p></td><td class="cell"><p>255</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>baç</p></td><td class="cell"><p>head, leader, beginning, top, main, principal</p></td><td class="cell"><p>5</p></td><td class="cell"><p>34</p></td><td class="cell"><p>68</p></td><td class="cell"><p>22</p></td><td class="cell"><p>90</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>el</p></td><td class="cell"><p>hand, stranger, country</p></td><td class="cell"><p>3</p></td><td class="cell"><p>75</p></td><td class="cell"><p>113</p></td><td class="cell"><p>38</p></td><td class="cell"><p>151</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>göz</p></td><td class="cell"><p>eye, glance, division, drawer</p></td><td class="cell"><p>3</p></td><td class="cell"><p>48</p></td><td class="cell"><p>92</p></td><td class="cell"><p>27</p></td><td class="cell"><p>119</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>kiz</p></td><td class="cell"><p>girl, virgin, daughter, get hot, get angry</p></td><td class="cell"><p>2</p></td><td class="cell"><p>72</p></td><td class="cell"><p>96</p></td><td class="cell"><p>21</p></td><td class="cell"><p>117</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>ön</p></td><td class="cell"><p>front, foreground, face, breast, prior, preliminary anterior</p></td><td class="cell"><p>5</p></td><td class="cell"><p>21</p></td><td class="cell"><p>72</p></td><td class="cell"><p>23</p></td><td class="cell"><p>95</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>sira</p></td><td class="cell"><p>queue, order, sequence, turn, regularity, occasion desk</p></td><td class="cell"><p>7</p></td><td class="cell"><p>30</p></td><td class="cell"><p>85</p></td><td class="cell"><p>28</p></td><td class="cell"><p>113</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>üst</p></td><td class="cell"><p>upper side, outside, clothing</p></td><td class="cell"><p>7</p></td><td class="cell"><p>20</p></td><td class="cell"><p>69</p></td><td class="cell"><p>23</p></td><td class="cell"><p>92</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>yan</p></td><td class="cell"><p>side, direction, auxiliary, askew, burn, be on fire be alight</p></td><td class="cell"><p>5</p></td><td class="cell"><p>21</p></td><td class="cell"><p>65</p></td><td class="cell"><p>31</p></td><td class="cell"><p>96</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>yol</p></td><td class="cell"><p>way, road, path, method, manner, means</p></td><td class="cell"><p>6</p></td><td class="cell"><p>17</p></td><td class="cell"><p>68</p></td><td class="cell"><p>29</p></td><td class="cell"><p>97</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Average</b></p></td><td class="cell"><p></p></td><td class="cell"><p><b>5</b></p></td><td class="cell"><p><b>39</b></p></td><td class="cell"><p><b>92</b></p></td><td class="cell"><p><b>31</b></p></td><td class="cell"><p><b>123</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Verbs</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>al</p></td><td class="cell"><p>take, get, red</p></td><td class="cell"><p>24</p></td><td class="cell"><p>180</p></td><td class="cell"><p>963</p></td><td class="cell"><p>125</p></td><td class="cell"><p>1088</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>bak</p></td><td class="cell"><p>look, fac, examine</p></td><td class="cell"><p>4</p></td><td class="cell"><p>136</p></td><td class="cell"><p>207</p></td><td class="cell"><p>85</p></td><td class="cell"><p>292</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>çaliç</p></td><td class="cell"><p>work, study, start</p></td><td class="cell"><p>4</p></td><td class="cell"><p>33</p></td><td class="cell"><p>103</p></td><td class="cell"><p>61</p></td><td class="cell"><p>164</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>çik</p></td><td class="cell"><p>climb, leave, increase</p></td><td class="cell"><p>6</p></td><td class="cell"><p>45</p></td><td class="cell"><p>138</p></td><td class="cell"><p>87</p></td><td class="cell"><p>225</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>geç</p></td><td class="cell"><p>pass,happen, late</p></td><td class="cell"><p>11</p></td><td class="cell"><p>51</p></td><td class="cell"><p>164</p></td><td class="cell"><p>90</p></td><td class="cell"><p>254</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>gel</p></td><td class="cell"><p>come, arrive, fit, seem</p></td><td class="cell"><p>20</p></td><td class="cell"><p>154</p></td><td class="cell"><p>346</p></td><td class="cell"><p>215</p></td><td class="cell"><p>561</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>gir</p></td><td class="cell"><p>enter, fit, begin, penetrate</p></td><td class="cell"><p>6</p></td><td class="cell"><p>88</p></td><td class="cell"><p>163</p></td><td class="cell"><p>84</p></td><td class="cell"><p>247</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>git</p></td><td class="cell"><p>go, leave, last, be over, pass</p></td><td class="cell"><p>13</p></td><td class="cell"><p>130</p></td><td class="cell"><p>214</p></td><td class="cell"><p>120</p></td><td class="cell"><p>334</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>gör</p></td><td class="cell"><p>see, understand, consider</p></td><td class="cell"><p>5</p></td><td class="cell"><p>155</p></td><td class="cell"><p>206</p></td><td class="cell"><p>68</p></td><td class="cell"><p>274</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>konuç</p></td><td class="cell"><p>talk, speak</p></td><td class="cell"><p>6</p></td><td class="cell"><p>42</p></td><td class="cell"><p>129</p></td><td class="cell"><p>63</p></td><td class="cell"><p>192</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Average</b></p></td><td class="cell"><p></p></td><td class="cell"><p><b>9.9</b></p></td><td class="cell"><p><b>101.4</b></p></td><td class="cell"><p><b>263.3</b></p></td><td class="cell"><p><b>99.8</b></p></td><td class="cell"><p><b>363.1</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Others</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>büyük</p></td><td class="cell"><p>big, extensive, important, chief, great, elder</p></td><td class="cell"><p>6</p></td><td class="cell"><p>34</p></td><td class="cell"><p>97</p></td><td class="cell"><p>26</p></td><td class="cell"><p>123</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>dogru</p></td><td class="cell"><p>straight, true, accurate, proper, fair, line towards, around</p></td><td class="cell"><p>6</p></td><td class="cell"><p>29</p></td><td class="cell"><p>81</p></td><td class="cell"><p>38</p></td><td class="cell"><p>119</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>kûçûk</p></td><td class="cell"><p>little, small, young, insignificant, kid</p></td><td class="cell"><p>4</p></td><td class="cell"><p>14</p></td><td class="cell"><p>45</p></td><td class="cell"><p>14</p></td><td class="cell"><p>59</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>öyle</p></td><td class="cell"><p>such, so, that</p></td><td class="cell"><p>4</p></td><td class="cell"><p>20</p></td><td class="cell"><p>51</p></td><td class="cell"><p>23</p></td><td class="cell"><p>74</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>son</p></td><td class="cell"><p>last, recent, final</p></td><td class="cell"><p>2</p></td><td class="cell"><p>76</p></td><td class="cell"><p>86</p></td><td class="cell"><p>18</p></td><td class="cell"><p>104</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>tek</p></td><td class="cell"><p>single, unique, alone</p></td><td class="cell"><p>2</p></td><td class="cell"><p>38</p></td><td class="cell"><p>40</p></td><td class="cell"><p>10</p></td><td class="cell"><p>50</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Average</b></p></td><td class="cell"><p></p></td><td class="cell"><p><b>4</b></p></td><td class="cell"><p><b>35.2</b></p></td><td class="cell"><p><b>66.7</b></p></td><td class="cell"><p><b>21.5</b></p></td><td class="cell"><p><b>88.2</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="3" global="61"/><p>Frequencies of the words have been found as it is necessary to select appropriate ambiguous words for WSD. There are 5356 different root words and 627 of these words have 15 or more occurrences, and the rest have less.</p><p>The XML files contains tagging information in the word (morphological analysis) and sentence level as a parse tree as shown in Figure 1. In the word level, inflectional forms are provided. And in the sentence level relations among words are given. The S tag is for sentence and W tag is for the word. IX is used for index of the word in the sentence, LEM is left as blank and lemma is given in the MORPH tag as a part of it with the morpho­logical analysis of the word. REL is for parsing information. It consists of three parts, two numbers and a relation. For example REL="[2, 1, (MODI­FIER)]" means this word is modifying the first in­flectional group of the second word in the sen­tence. The structure of the treebank data was de­signed by METU. Initially lemmas were decided to be provided as a tag by itself, however, lemmas are left as blank. This does not mean that lemmas are not available in the treebank; the lemmas are given as a part of "IG" tag. Programs are available for extracting this information for the time being. All participants can get these programs and thereby the lemmas easily and instantly.</p><p>The sense tags were not included in the treebank and had to be added manually. Sense tagging has been checked in order to obtain gold standard data. Initial tagging process has been finished by a sin­gle tagger and controlled. Two other native speaker in the team tagged and controlled the examples. That is, this step was completed by three taggers. Problematic cases were handled by a commission and the decision was finalized when about 90% agreement has been reached.</p></section><section number="3" title="Dictionary"><p>The dictionary is the one that is published by TDK<footnote anchor="1"/> (Turkish Language Foundation) and it is open to public via internet. This dictionary lists the senses along with their definitions and example sentences that are provided for some senses. The dictionary is used only for sense tagging and enumeration of the senses for standardization. No specific information other than the sense numbers is taken from the dictionary; therefore there is no need for linguistic processing of the dictionary.</p></section><section number="4" title="Training and Evaluation Data"><p>In Table 1 statistical information about the final training and testing sets of TLST is summarized. The data have been provided for 3 words in the trial set and 26 words in the final training and test­ing sets (10 nouns, 10 verbs and 6 other POS for the rest of POS including adjectives and adverbs). It has been tagged about 100 examples per word, but the number of samples is incremented or dec­remented depending on the number of senses that specific word has. For a few words, however, fewer examples exist due to the sparse distribution of the data. Some ambiguous words had fewer ex­amples in the corpus, therefore they were either eliminated or some other examples drawn from external resources were added in the same format. On the average, the selected words have 6.7 senses, verbs, however, have more. Approximately 70% of the examples for each word were delivered as training data, whereas approximately 30% was reserved as evaluation data. The distribution of the senses in training and evaluation data has been kept proportional. The sets are given as plain text files for each word under each POS. The samples for the words that can belong to more than one POS are listed under the majority class. POS is provided for each sample.</p><p>We have extracted example sentences of the tar­get word(s) and some features from the XML files. Then tab delimited text files including structural and sense tag information are obtained. In these files each line has contextual information that are thought to be effective (Orhan and Altan, 2006; Orhan and Altan, 2005) in Turkish WSD about the target words. In the upper level for each of them XML file id, sentence number and the order of the ambiguous word are kept as a unique key for that specific target. In the sentence level, three catego­ries of information, namely the features related to the previous words, target word itself and the sub­sequent words in the context are provided.</p><footnote label="1"> http://tdk.org.tr/tdksozluk/sozara.htm</footnote><page local="4" global="62"/><p>In the treebank relational structure, there can be more than one word in the previous context related to the target, however there is only a single word in the subsequent one. Therefore the data for all words in the previous context is provided sepa­rately. The features that are employed for previous and the subsequent words are the same and they are the root word, POS(corrected), tags for ontol­ogy level 1, level 2 and level 3, POS, inflected POS, case marker, possessor and relation. How­ever for the target word only the root word, POS, inflected POS, case marker, possessor and relation are taken into consideration. Fine and coarsegrained (FG and CG respectively) sense numbers and the sentence that has the ambiguous word have been added as the last three feature. FG senses are the ones that are decided to be the exact senses. CG senses are given as a set that are thought to be possible alternatives in addition to the FG sense. Table 2 demonstrates the whole list of features provided in a single line of data files along with an example. The "?" in the features shows the missing values. This is actually corresponding to the fea­tures that do not exist or can not be obtained from the treebank due to some problematic cases. The line that corresponds to this entry will be the fol­lowing line (as tab delimited):<page local="5" global="63"/></p><table caption="Table 2: Features and example" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Feature</b></p></td><td class="cell"><p><b>Example</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>File id</p></td><td class="cell"><p>00002213148.xml</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Sentence number</p></td><td class="cell"><p>9</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Order</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word root/lemma</p></td><td class="cell"><p>tap</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word POS(corrected)</p></td><td class="cell"><p>verb</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word onthology level1</p></td><td class="cell"><p>abstraction</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word onthology level2</p></td><td class="cell"><p>attribute</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word onthology level3</p></td><td class="cell"><p>emotion</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word POS</p></td><td class="cell"><p>verb</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word POS(derivation)</p></td><td class="cell"><p>adv</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word case marker</p></td><td class="cell"><p>?</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word possessor</p></td><td class="cell"><p>fl</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Previous related word-target word relation</p></td><td class="cell"><p>modifier</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Target word root/lemma</p></td><td class="cell"><p>sev</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Target word POS</p></td><td class="cell"><p>verb</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Target word POS(derivation)</p></td><td class="cell"><p>noun</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Target word case marker</p></td><td class="cell"><p>abl</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Target word possessor</p></td><td class="cell"><p>tr</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Target word-subsequent word relation</p></td><td class="cell"><p>object</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word root/lemma</p></td><td class="cell"><p>sikil</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word POS(corrected)</p></td><td class="cell"><p>verb</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word onthology level1</p></td><td class="cell"><p>abstraction</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word onthology level2</p></td><td class="cell"><p>attribute</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word onthology level3</p></td><td class="cell"><p>emotion</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word POS</p></td><td class="cell"><p>verb</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word POS(derivation)</p></td><td class="cell"><p>verb</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word case marker</p></td><td class="cell"><p>?</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word possessor</p></td><td class="cell"><p>fl</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Subsequent related word-target word relation</p></td><td class="cell"><p>sentence</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Fine-grained sense number</p></td><td class="cell"><p>2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Coarse-grained sense number</p></td><td class="cell"><p>2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>#ne tuhaf §ey ; degil mi ?</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>iyi olmamdan ; onu taparcasina</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Sentence</p></td><td class="cell"><p>sevmemden sikildi .#</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><doubt alpha="52.5" length="40" tooSmall="False" monospace="0.0">00002213148.xml 9 0 tap verb abstraction</doubt><p>attribute emotion verb adv ? fl modifier sev verb noun abl tr object sikil verb abstraction attribute emotion verb verb ? fl sentence 2 2 #ne tuhaf §ey ; degil mi ?iyi olmamdan ; onu taparcasina sevmemden sikildi .#</p></section><section number="5" title="Ontology"><p>A small scale ontology for the target words and their context was constructed. The Turkish Word­Net developed at Sabanci University<footnote anchor="2"/> is somehow insufficient. Only the verbs have some levels of relations similar to English WordNet. The nouns, adjectives, adverbs and other words that are fre­quently used in Turkish and in the context of the ambiguous words were not included. This is not a suitable resource for fulfilling the requirements of TLST and an ontology specific to this task was required. The ontology covers the examples that are selected and has three levels of relations that are supposed to be effective in the disambiguation process. We tried to be consistent with the Word­Net tags; additionally we constructed the ontology not only for nouns and verbs but for all the words that are in the context of the ambiguous words se­lected. Additionally we tried to strengthen the rela­tion among the context words by using the same tags for all POS in the ontology. This is somehow deviating from WordNet methodology, since each word category has its own set of classification in it.</p></section><section number="6" title="Evaluation"><p>WSD is a new area of research in Turkish. The sense tagged data provided in TLST are the first resources for this specific domain in Turkish. Due to the limited and brand new resources available and the time restrictions the participation was less. We submitted a very simple system that utilizes statistical information. It is similar to the Naïve Bayes approach. The features in the training data was used individually and the probababilities of the senses are calculated. Then in the test phase the probabilities of each sense is calculated with the given features and the three highest-scored senses are selected as the answer. The average precision and recall values for each word category are given in Table 3. The values are not so high, as it can be expected. The size of the training data is limited, but the size is the highest possible under these cir­cumstances, but it should be incremented in the near future. The number of senses is high and pro­viding enough instances is difficult. The data and the methodology for WSD will be improved by the experience obtained in SemEval evaluation exer­cise.</p><footnote label="2"> http://www.hlst.sabanciuniv.edu/TL/</footnote><p>The evaluation is done only for FG and CG senses. For FG senses no partial points are as­signed and 1 point is assigned for a correct match. On the other hand, the CG senses are evaluated partially. If the answer tags are matching with any of the answer tags they are given points.</p></section><section number="7" title="Conclusion"><p>In TLST we have prepared the first resources for WSD researches in Turkish. Therefore it has sig­nificance in Turkish WSD studies. Although the resources and methodology have some deficien­cies, a valuable effort was invested during the de­velopment of them. The resources and the method­ology for Turkish WSD will be improved by the experience obtained in SemEval and will be open to public in the very near future from http://www.fatih.edu.tr/~zorhan/senseval/senseval.htm.</p><table caption="Table 3: Average Precision and Recall values" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Words</p></td><td class="cell"><p></p></td><td class="cell"><p>FG</p></td><td class="cell"><p></p></td><td class="cell"><p>CG</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>P</p></td><td class="cell"><p>R</p></td><td class="cell"><p>P</p></td><td class="cell"><p>R</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Nouns</p></td><td class="cell"><p>0,15</p></td><td class="cell"><p>0,50</p></td><td class="cell"><p>0,65</p></td><td class="cell"><p>0,43</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Verbs</p></td><td class="cell"><p>0,10</p></td><td class="cell"><p>0,38</p></td><td class="cell"><p>0,56</p></td><td class="cell"><p>0,50</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Others</p></td><td class="cell"><p>0,13</p></td><td class="cell"><p>0,50</p></td><td class="cell"><p>0,57</p></td><td class="cell"><p>0,44</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Average</p></td><td class="cell"><p><b>0,13</b></p></td><td class="cell"><p><b>0,46</b></p></td><td class="cell"><p><b>0,59</b></p></td><td class="cell"><p><b>0,46</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></section><references><doubt alpha="65.6" length="131" tooSmall="False" monospace="0.0">Orhan, Z. and Altan, Z. 2006.Impact of Feature Selec­tion for Corpus-Based WSD in Turkish,LNAI, Springer-Verlag, Vol. 4293: 868-878</doubt><p>Orhan Z. and Altan Z. 2005. <i>Effective Features for Dis­ambiguation of Turkish Verbs, </i>IEC'05, Prague, Czech Republic: 182-186</p><p>Oflazer, K., Say, B., Tur, D. Z. H. and Tur, G. 2003. <i>Building A Turkish Treebank, </i>Invited Chapter In Building And Exploiting Syntactically-Annotated Corpora, Anne Abeille Editor, Kluwer Academic Publishers.</p></references></body></article>