<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>WNTERM: Enriching the MCR with a Terminological Dictionary</title><author surname="Pociello" givenname="Eli"><org  name="Elhuyar R&amp;D" city="Basque Country"/></author><author surname="Gurrutxaga" givenname="Antton"><org  name="Elhuyar R&amp;D" city="Basque Country"/></author><author surname="Agirre" givenname="Eneko"><org  name="Elhuyar R&amp;D" city="Basque Country"/></author><author surname="Aldezabal" givenname="Izaskun"><org  name="Elhuyar R&amp;D" city="Basque Country"/></author><author surname="Rigau" givenname="German"><org  name="Elhuyar R&amp;D" city="Basque Country"/></author></firstpageheader><frontmatter><p><b>WNTERM: Enriching the MCR with a terminological dictionary</b></p><p><b>Eli Pociello, Antton Gurrutxaga</b></p><p>Elhuyar R&amp;D</p><p>Zelai Haundi kalea, 3. Osinalde Industrialdea, 20170 Usurbil. Basque Country E-mail: {agurrutxaga, eli}@elhuyar.com</p><p><b>Eneko Agirre, Izaskun Aldezabal, German Rigau</b></p><p>IXA NLP Research Group 649 pk. 20.080 - Donostia. Basque Country E-mail: {e.agirre, izaskun.aldezabal, german.rigau}@ehu.es</p></frontmatter><abstract>In this paper we describe the methodology and the first steps for the creation of WNTERM (from WordNet and Terminology), a specialized lexicon produced from the merger of the EuroWordNet-based Multilingual Central Repository (MCR) and the Basic Encyclopaedic Dictionary of Science and Technology (BDST). As an example, the ecology domain has been used. The final result is a multilingual (Basque and English) light-weight domain ontology, including taxonomic and other semantic relations among its concepts, which is tightly connected to other wordnets. </abstract></header><body><section number="1." title="Introduction"><p>In the last seven years the IXA research group has been working on the Basque WordNet (Agirre et al., 2006). The focus have been on the representation of common vocabulary (we have incorporated around 27,000 words), but now, we have turned our attention to specialized language.</p><p>Due to the unstoppable development of specialized language and terminology, it is becoming increasingly difficult to capture and organize terminological information. Domain ontologies are helpful to handle this kind of information, and have been shown to be useful in knowledge representation, management and exchange. For instance, there are several representative ontologies in the domains of e-commerce (UNSPSC1,NAICS2), medicine (GALEN3, UMLS), engineering -EngMath (Gruber &amp; Olsen, 1994), PhysSys (Borst, 1997) -, enterprise -Enterprise Ontology (Uschold et al., 1998) -, and knowledge management -KA (Decker et al. 1999) . Moreover, in the NLP community domain ontologies are being used to develop and evaluate different computational systems and applications (Navigli et al. 2003; Sagri et al., 2003; Stamou et al., 2002; Roventini &amp; Marinelli, 2003).</p><p>The aim of WNTERM is to create a light-weight ontology belonging to the science and technology domains for Basque and English. WNTERM (i) stores the domain terminology, (ii) fixes relations among all the domain terms, and (iii) is connected to the Basque and English wordnets through the Multilingual Central Repository (MCR) (Atserias et al., 2004).</p><p>In order to have WNTERM linked to the wordnets currently available, we decided to structure it following the MCR framework, which we also used to build the Basque WordNet). The MCR model provides English-Basque equivalence links across the English and Basque wordnets, and it also offers the possibility of enriching both   the   Basque   WordNet   and WNTERM simultaneously.</p><p>Domain ontology terms are imported into WNTERM from the MCR itself (more precisely, from the Basque and English wordnets) and the in-house <i>Basic Encyclopaedic Dictionary of Science and Technology </i>(BDST) data-base4. Taking this sources into account, there are additional objectives which we want to address with WNTERM: (i) to extend the Basque WordNet with terminology, (ii) to organize hierarchically all the terms of the BDST according to the MCR architecture, in order to provide to lexicographers more specific information about a term in their navigation through the data-base, and (iii) to link all the sources in the project: the MCR, the BDST and the domain ontology (WNTERM).</p><p>The final result of this project will be a domain ontology on science and technology, which will show the taxonomic and semantic relations among all its domain concepts, and which is also linked to the MCR and the BDST. In this paper we focus on the overall methodology for the construction of WNTERM, which will be illustrated on the ecology domain, a subset of science and technology.</p><p>This paper is organized as follows. The resources used in this project are introduced in Section 2 (MCR and the Basque WordNet) and Section 3 (BDST). In Section 4, we present the methodology followed to create a domain ontology, illustrated on the ecology domain. Finally, future work is presented in Section 5.</p><footnote label="1"> http://www.unspsc.org</footnote><footnote label="2"> http://www.naics.com</footnote><footnote label="3"> http://opengalen.org</footnote><footnote label="4"> http ://www.zientzia.net/hiztegia/index. asp</footnote><page local="2"/></section><section number="2." title="The Multilingual Central Repository"><p>The Multilingual Central Repository (MCR)5follows the model proposed by the EuroWordNet project6.EuroWordNet (Vossen, 1998) is a multilingual semantic lexicon with wordnets for several European languages, which are structured as the Princeton WordNet (Fellbaum, 1998).</p><p>It groups each languages' words into sets of synonyms called synsets, and records various semantic relations (such as hypernymy, hyponymy, meronymy, holonymy) between these synonym sets forming a hierarchy. Each of these synsets corresponds to a lexical concept and many have a textual gloss which often provides an explanation of what this concept represents. The MCR (Atserias et al., 2004) is a result of the 5th Framework Meaning project7(Rigau et al., 2003). The MCR integrates in the same EuroWordNet framework wordnets from five different languages, including Spanish, Italian, Catalan and Basque (together with six English WordNet versions).</p><p>The wordnets are currently linked via an Inter-Lingual-Index (ILI) allowing the connection from words in one language to translation equivalent words in any of the other languages. In that way, the MCR constitutes a natural multilingual large-scale linguistic resource for a number of semantic processes that need large amount of multilingual knowledge to be effective tools. For instance, the English synset <i>{party, political_party} </i>is linked through the ILI to the Basque synset <i>{partidu, partidu_politiko, alderdi_politiko, alderdi}. </i>The MCR also integrates the latest version of the WordNet Domains (Magnini &amp; Cavagliä, 2000), new versions of the Base Concepts and the Top Concept Ontology (Älvez et al., 2008), and the SUMO ontology (Niles &amp; Pease, 2001). The current version of the MCR contains 934,771 semantic relations between synsets, most of them acquired by automatic means. This represents almost four times larger than the Princeton WordNet (235,402 unique semantic relations in WordNet 3.0).</p><p>Although these resources have been derived using different WordNet versions, using the technology for the automatic alignment of wordnets (Daude et al., 2001), most of these resources have been integrated in the MCR maintaining the compatibility among all the knowledge resources which use a particular WordNet version as a sense repository.</p><p>MCR was developed based on WordNet 1.6 version. However, for this project we have moved the MCR to the latest WordNet version (3.0), because WordNet 3.0 has a higher amount of terminology (than the 1.6 version). The result of the automatic mapping was manually corrected. We will now present WordNet domains and the Basque WordNet.</p><footnote label="5"> http://adimen.si.ehu.es/cgi-bin/wei5/public/wei.consult.perl</footnote><footnote label="6"> http://www.illc.uva.nl/EuroWordNet/</footnote><footnote label="7"> http://www.lsi.upc.edu/~nlp/meaning</footnote><subsection number="2.1" title="WordNet Domains"><p>WordNet Domains8(Magnini &amp; Cavaglia, 2000) is a lexical resource where the synsets have been annotated semi automatically with one or more domain labels. These domain labels are organized hierarchically. These labels group meanings in terms of topics or scripts, e.g. <i>Transport, Sports, Medicine, Gastronomy. </i>which were partially derived from the Dewey Decimal Classification. The version we used in these experiments is a hierarchy of 171 Domain Labels associated to WordNet 1.6. Information brought by Domain Labels is complementary to what is already in WordNet. First of all Domain Labels may include synsets of different syntactic categories: for instance <i>Medicine </i>groups together senses from nouns, such as <i>doctor </i>and <i>hospital, </i>and from verbs such as <i>to operate. </i>Second, a Domain Label may also contain senses from different WordNet subhierarchies. For example, <i>Sport </i>contains senses such as <i>athlete </i>deriving from <i>person, game equipment </i>from <i>artifact, sport </i>from <i>act, </i>and <i>playing </i>field from <i>location.</i></p></subsection><subsection number="2.2" title="The Basque WordNet"><p>The Basque WordNet9 was developed by the IXA research group10developed within the framework of multilingual counterparts EuroWordNet and the MCR. The Basque WordNet has been constructed with the expand approach (Vossen, 1998), which means that the English synsets have been enriched with Basque variants. Besides, we also incorporate new synsets that exist for Basque but not for English. Due to EuroWordNet and the MCR frameworks, the Basque WordNet is already linked to the Spanish, Catalan, English and Italian wordnets, and it can also be linked to any other wordnet tightly linked to the English WordNet.</p><doubt alpha="63.6" length="44" tooSmall="False" monospace="0.0">WordNet (Fellbaum, 1998)11,as well as on its</doubt><p>Up to now, the Basque WordNet has been focused on general vocabulary leaving aside specialized language and terminology, and one of the goals of WNTERM is to enrich the Basque WordNet with terminological information. It currently contains 26,999 headwords, 33,302 synsets and 50,841 senses.</p><p><b>3.   The Elhuyar Basic Dictionary of Science and Technology (BDST)</b></p><p>The BDST is an specialized dictionary published on line by Elhuyar Foundation12. The BDST is designed as a terminological dictionary; thus, each concept is represented in a terminological record, which includes all the information relating to that concept: the terms (descriptors) that convey the concept, the definition, the domain(s), etc. One concept can be related to several domains. The BDST includes concepts of several areas of Science and Technology.<page local="3"/> The current amount of concepts is 15,627 belonging to the following knowledge areas or domains. For instance, there are 353 concepts in the ecology domain.</p><footnote label="8"> http://wndomains.itc.it/wordnetdomains.html</footnote><footnote label="9"> ixa2 .si.ehu. es/mcr/wei.html</footnote><footnote label="1">0  http://ixa.si.ehu.es</footnote><footnote label="1">1  http://wordnet.princeton.edu/</footnote><footnote label="12"> www.elhuyar.org</footnote><p>The BDST includes terms in four languages: English, Spanish, French, and Basque. There is no semantic relation between concepts.</p><p>The BDST is an ongoing project. The objectives for 2008 are to include 10,000 new concepts in the dictionary. One of the most relevant term sources is the Corpus of Science and Technology (Alegria et al., 2007), a 7,6 million words corpus tagged by hand with morphosyntactic information13. Erauzterm (Alegria et al., 2004) was used to extract automatically the terms from the corpus. Further developments of the BDST are directly related to the results of WNTERM, as far as the establishment of relations between concepts is a promising task that would enrich the dictionary and enhance its value for users.</p></subsection></section><section number="4." title="A methodology for the construction of WNTERM"><p>In this section, we will describe the steps for the construction of WNTERM. We first compare the two resources, with special attention to the relation between the domain labels used in each resource, and an automatic analysis of the different cases found. In the following steps we select the concepts to be added to WNTERM and, the terms to be added. We finally structure the concepts in a hierarchical structure. We will illustrate the steps with data from the ecology domain.</p><subsection number="4.1" title="Comparison of both sources"><p>The first step has been the comparison between both sources (the MCR and the BDST) in order to measure the amount of terms in each resource (and their overlap) taking into account the domain information across both resources. For this comparison an automatic procedure has been used, and the result have been qualitatively analyzed.</p><p>In order to carry out this comparison, the first step has been the manual mapping of the domain labels of both resources.</p><subsubsection number="4.1.1" title="Manual mapping of domain labels"><p>As we have mentioned in Sections 2 and 3, concepts of the MCR and the BDST have a domain label. However, each of these sources has a different domain classification. In the MCR there are 171 domains labels which are organized hierarchically (for instance, <i>cinema, radio, post, tv, telegraphy </i>and <i>telephony </i>are subdomains of <i>telecommunications). </i>BDST has 34 flat domain labels. As domain information is required to measure the overlap of concepts between both resources, we have manually mapped the 171 domain labels of the MCR to the 34 domain labels of the BDST. From the BDST domains 18 have been mapped to a single MCR domain, and 16 have been related to one or more domain labels of the MCR (e.g. <i>telecommunications </i>to <i>cinema, radio, post, tv, telegraphy, telephony). </i>Regarding MCR domains, 64 have been mapped to a single BDST domain, but 20 have been mapped to more than one (e.g. <i>topography </i>to <i>geography, geology </i>and <i>town planning </i>to <i>architecture, building). </i>87 of the MCR domains have not been mapped to any BDST domain because of their high degree of specification <i>(fencing, paranormal, numismatics), </i>or because the BDST have not been enriched with these fields yet <i>(religion, sport, gastronomy).</i></p></subsubsection><subsubsection number="4.1.2" title="Automatic comparison and the resulting casuistry"><p>We performed an automatic analysis of the relation between the MCR and the BDST. We checked for overlaps of terms in each resource (both in English and Basque). If the term is found in both resources, we analyze whether the domain labels can be mapped (according to the domain mapping we described in the previous sections).</p><p>The results of this comparison have been tagged using the codes described in Table 2.</p><footnote label="1">3  www.ztcorpusa.net</footnote><table caption="Table 1: Current figures and domain classification of the BDST concepts." class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Aeronautics</p></td><td class="cell"><p>333</p></td><td class="cell"><p>Geography</p></td><td class="cell"><p>124</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Agriculture</p></td><td class="cell"><p>474</p></td><td class="cell"><p>Geology</p></td><td class="cell"><p>661</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Anatomy</p></td><td class="cell"><p>705</p></td><td class="cell"><p>Mathematics</p></td><td class="cell"><p>603</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Anthropology</p></td><td class="cell"><p>20</p></td><td class="cell"><p>Medicine</p></td><td class="cell"><p>2,207</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Architecture</p></td><td class="cell"><p>112</p></td><td class="cell"><p>Metallurgy</p></td><td class="cell"><p>356</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Astronomy</p></td><td class="cell"><p>196</p></td><td class="cell"><p>Meteorology</p></td><td class="cell"><p>435</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Astronautics</p></td><td class="cell"><p>31</p></td><td class="cell"><p>Mycology</p></td><td class="cell"><p>138</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Biochemistry</p></td><td class="cell"><p>698</p></td><td class="cell"><p>Microbiology</p></td><td class="cell"><p>109</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Biology</p></td><td class="cell"><p>711</p></td><td class="cell"><p>Mineralogy</p></td><td class="cell"><p>260</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Botanic</p></td><td class="cell"><p>1,606</p></td><td class="cell"><p>Motoring</p></td><td class="cell"><p>81</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Building Industry</p></td><td class="cell"><p>212</p></td><td class="cell"><p>Palaeontology</p></td><td class="cell"><p>36</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Chemistry</p></td><td class="cell"><p>1,490</p></td><td class="cell"><p>Photography</p></td><td class="cell"><p>51</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Computer Science</p></td><td class="cell"><p>392</p></td><td class="cell"><p>Physics</p></td><td class="cell"><p>827</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Ecology</p></td><td class="cell"><p>353</p></td><td class="cell"><p>Technology</p></td><td class="cell"><p>790</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Electricity</p></td><td class="cell"><p>220</p></td><td class="cell"><p>Telecommunications 127</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Electronics</p></td><td class="cell"><p>55</p></td><td class="cell"><p>Zoology</p></td><td class="cell"><p>2,152</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>General</p></td><td class="cell"><p>40</p></td><td class="cell"><p>Not classified</p></td><td class="cell"><p>267</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Genetics</p></td><td class="cell"><p>156</p></td><td class="cell"><p><b>TOTAL</b></p></td><td class="cell"><p><b>17,028</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="4"/></subsubsection><subsubsection number="4.1.3" title="First conclusions of the comparison"><p>Table 3 shows the results of the automatic comparison. There are 51,469 terms overlapped in English and 26,379 in Basque. Therefore, as we expected, the overlap between the Basque WordNet and the BDST is smaller. As we have already mentioned, this is due to the fact that the Basque WordNet has focused on common vocabulary.</p><p><b><u>MCR-BDST overlap</u></b></p><p>BDST.</p><p>Note that we applied a strict string match among terms. This means that often the same term is written in similar but different ways in both resources. For instance, the BDST has <i>videotape, vanilla plant, goldfish </i>and <i>desertization </i>while the MCR contains very similar terms like <i>video tape, vanilla </i>and <i>gold-fish, </i>or a synonym like <i>desertification. </i>The numbers in Table 3 correspond to this strict term match, and thus, they are an underestimation of the real overlap between the MCR and the BDST. In the next section, we will see that a manual review of the 0-1 and 1-0 pairs allows to detect this false mismatches.</p><p>Table 3 also shows that most of the terms overlapped have different domain labels (45,876 in English and 23,664 in Basque). These terms must be previously checked by hand because some of them may not refer to the same concept. For instance, the term <i>horn </i>has been marked in the automatic comparison as 1-2, that is, this term is both in the MCR and in the BDST, but they have different domains. In the MCR the term <i>horn </i>can belong to the domain of <i>anatomy </i>and in the BDST belongs to <i>zoology, </i>both referring to the same concept <i>(one of the bony outgrowths on the heads of certain ungulates). </i>On the other hand, there is another <i>horn </i>in the MCR belonging to the <i>transport </i>domain <i>(a</i><i> device on an automobile for making a warning noise) </i>which has nothing to do with the BDST term from <i>zoology. </i>Therefore, once more, these numbers are an underestimation of the real overlap between the MCR and the BDST.</p><p><b>_<u>Domain</u>_<u>English terms Basque terms</u></b></p><p>Table 4: Number of English and Basque terms missing in the MCR with comparison to the BDST, organized by BDST donains and ordered according to the number of English missing terms.</p><table caption="Table 2: Codes used in the automatic procedure." class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Codes</b></p></td><td class="cell"><p><b>Description</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 0</b></p></td><td class="cell"><p>The term is in the MCR but it is not in the BDST</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 1</b></p></td><td class="cell"><p>The term is both in the MCR and the BDST</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 2</b></p></td><td class="cell"><p>The term is both in the MCR and the BDST, but they have different domains</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 3</b></p></td><td class="cell"><p>The term is only in the MCR, but there is one synonym of this term in the BDST and they have the same domain label</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 4</b></p></td><td class="cell"><p>The term is only in the MCR, and there is one synonym of this term in the BDST but they have not the same domain label</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>0 - 1</b></p></td><td class="cell"><p>The term is in the BDST but it is not in the MCR</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 1</b></p></td><td class="cell"><p>The term is both in the BDST and the MCR</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>2 - 1</b></p></td><td class="cell"><p>The term is both in the BDST and the MCR, but they have different domains</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>3 - 1</b></p></td><td class="cell"><p>The term is only in the BDST, but there is one synonym of this term in the MCR and they have the same domain label</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>4 - 1</b></p></td><td class="cell"><p>The term is only in the BDST, and there is one synonym of this term in the MCR but they have not the same domain label</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Medicine</p></td><td class="cell"><p>1,111</p></td><td class="cell"><p>1,592</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Chemistry</p></td><td class="cell"><p>1,085</p></td><td class="cell"><p>994</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Zoology</p></td><td class="cell"><p>465</p></td><td class="cell"><p>1,141</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Biochemistry</p></td><td class="cell"><p>459</p></td><td class="cell"><p>228</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Botanics</p></td><td class="cell"><p>374</p></td><td class="cell"><p>775</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Geology</p></td><td class="cell"><p>332</p></td><td class="cell"><p>417</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Computer Science</p></td><td class="cell"><p>329</p></td><td class="cell"><p>368</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Meteorology</p></td><td class="cell"><p>311</p></td><td class="cell"><p>366</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Metallurgy</p></td><td class="cell"><p>310</p></td><td class="cell"><p>321</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Biology</p></td><td class="cell"><p>308</p></td><td class="cell"><p>482</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Technology</p></td><td class="cell"><p>287</p></td><td class="cell"><p>386</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Physics</p></td><td class="cell"><p>274</p></td><td class="cell"><p>460</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Aeronautics</p></td><td class="cell"><p>266</p></td><td class="cell"><p>291</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Anatomy</p></td><td class="cell"><p>244</p></td><td class="cell"><p>468</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Ecology</p></td><td class="cell"><p>211</p></td><td class="cell"><p>245</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Mathematics</p></td><td class="cell"><p>203</p></td><td class="cell"><p>325</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Agriculture</p></td><td class="cell"><p>197</p></td><td class="cell"><p>177</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Telecommunications</p></td><td class="cell"><p>138</p></td><td class="cell"><p>135</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Mineralogy</p></td><td class="cell"><p>129</p></td><td class="cell"><p>206</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Genetics</p></td><td class="cell"><p>97</p></td><td class="cell"><p>45</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Electricity</p></td><td class="cell"><p>74</p></td><td class="cell"><p>123</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Building Industry</p></td><td class="cell"><p>62</p></td><td class="cell"><p>97</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Astronautics</p></td><td class="cell"><p>37</p></td><td class="cell"><p>95</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Motoring</p></td><td class="cell"><p>36</p></td><td class="cell"><p>40</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Microbiology</p></td><td class="cell"><p>33</p></td><td class="cell"><p>81</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Architecture</p></td><td class="cell"><p>29</p></td><td class="cell"><p>51</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Mycology</p></td><td class="cell"><p>25</p></td><td class="cell"><p>90</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Geography</p></td><td class="cell"><p>24</p></td><td class="cell"><p>31</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Photography</p></td><td class="cell"><p>15</p></td><td class="cell"><p>26</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Palaeontology</p></td><td class="cell"><p>9</p></td><td class="cell"><p>23</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>General</p></td><td class="cell"><p>6</p></td><td class="cell"><p>7</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>TOTAL</b></p></td><td class="cell"><p><b>7,480</b></p></td><td class="cell"><p><b>10,086</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 3: Overlap between terms of the MCR and the" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Overlap type</b></p></td><td class="cell"><p><b>English terms</b></p></td><td class="cell"><p><b>Basque terms</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 1</b></p></td><td class="cell"><p>2,613</p></td><td class="cell"><p>1,853</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 2</b></p></td><td class="cell"><p>18,591</p></td><td class="cell"><p>12,190</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 3</b></p></td><td class="cell"><p>2,850</p></td><td class="cell"><p>803</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 4</b></p></td><td class="cell"><p>16,572</p></td><td class="cell"><p>4,447</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>2 - 1</b></p></td><td class="cell"><p>9,228</p></td><td class="cell"><p>6,040</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>3 - 1</b></p></td><td class="cell"><p>130</p></td><td class="cell"><p>59</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>4 - 1</b></p></td><td class="cell"><p>1,485</p></td><td class="cell"><p>987</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>TOTAL</b></p></td><td class="cell"><p><b>51,469</b></p></td><td class="cell"><p><b>26,379</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="5"/><p>The comparison between both resources showed that the MCR contains many English words which are missing from the BDST (80,119 in English and 7,541 thousand in Basque), and that terms missing from the MCR which are covered in the BDST is smaller (7,480 in English). This is not surprising, as the BDST focuses on just science and technology related terms. In any case, besides the construction of WNTERM itself, the cross-enrichment of both resources is significant, as both of them benefit from the information contained in the other.</p><p>In the next subsections we will illustrate the methodology to build WNTERM on the terms from a single domain. Table 4 shows that it's in the medicine domain where the BDST could benefit the MCR most. However, the chosen domain is that of ecology, mainly because it is related to the project KYOTO14, where we take part and which is focused on environmental issues.</p></subsubsection></subsection><subsection number="4.2" title="Selection of concepts"><p>In this step, we select the concepts that are going to be added to WNTERM. We focus on English (because the Basque WordNet is a subset of the English WordNet, and the Basque and English BDST are equivalent), and on the ecology domain. The idea is to take as basis the concepts (synsets) already in the MCR, and add all of them to WNTERM. If new concepts exist in the BDST, they are also added to WNTERM and manually linked to their corresponding superclass (hypernym). All the concepts in the MCR tagged with the ecology domain label, irrespective of their presence in the BDST, have been copied to WNTERM. In other words, those terms of the MCR tagged with the codes 1-0, 1-1, 1-3 and 3-1 have been directly added to WNTERM. Note that terms tagged with 1-2 and 1-4 have been ignored for the time being, because, as we said in section 4.1.3, these terms, having different domain labels from those in the BDST, must be previously checked by hand to confirm whether they refer to the same concept. We then have turned our attention to those concepts in the BDST not present in the MCR (those marked as 0-1). We first have examined term matching problems (see examples in section 4.1.3). If we think two terms are the same, and if the concept behind the term was not already in WNTERM, the relevant concept is copied to WNTERM. If the term corresponds to a concept which was not in the MCR, this has been copied to WNTERM, and a link to its hypernym in the Interlingual Index of the MCR added. The result of this additions can be seen in Figure 1, where concepts in bold correspond to those coming from ths BDST and those in italics to their hypernyms.</p><p>WNTERM.</p><p>Table 5 shows the final amount of concepts included in WNTERM. As it can be seen most of the concepts are nominal (222). Adjectives have been also added (40) but due to the fact they can also function as nouns (e.g. <i>mutant(s), perennial(s)), </i>we have given special attention to their inclusion in WNTERM. For the time being, we have focused on the <i>Elhuyar Basque Dictionary </i>to decide their category. However, we would like to study the possibility of introducing these concepts twice (both as nominal and as adjectival) and relating them with the MCR <i>XPOS_EQ_near_synonymy </i>relation. This relation is used when a meaning matches multiple ILI-records simultaneously (in this case, a noun and an adjective), or when there is some doubt about the precise mapping.</p></subsection><subsection number="4.3" title="Selection of terms"><p>The third step corresponds to the terms. Both English and Basque terms have been automatically added from the MCR and the BDST to their corresponding concepts of WNTERM. As we have already detected the relevant concepts in the previous section, we just look up the terms for those concepts in the two resources, and copy them to WNTERM.</p><p>h<u>ttp://www.kvoto-proiect.eu</u></p><table caption="Table 6: Figures of ecology domain terms in the"></table><table caption="Table 6 shows the amount of terms added from the MCR and the BDST in WNTERM."></table><p>Finally, in the case of Basque, some gaps have been automatically detected, and therefore, Basque variants have been added by hand.</p></subsection><subsection number="4.4" title="Structure the domain ontology"><p>Finally, the hypernymy links have been used to structure the hierarchy of concepts in the domain ontology, i.e. we have used the hypernymy tree of WordNet to hierarchically organize the terms, plus the hypernym links for the new concepts that we added by hand in the previous section. Irrelevant splits in the hierarchy can be automatically detected and deleted (Vossen, 2001).</p><table caption="Table 5: Figures of the ecology domain concepts in the" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Term values</b></p></td><td class="cell"><p><b>WNTERM concepts</b></p></td><td class="cell"><p><b>Noun</b></p></td><td class="cell"><p><b>Adj.</b></p></td><td class="cell"><p><b>Verb</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 0</b></p></td><td class="cell"><p>8</p></td><td class="cell"><p>7</p></td><td class="cell"><p>1</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 1</b></p></td><td class="cell"><p>2</p></td><td class="cell"><p>2</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 3</b></p></td><td class="cell"><p>2</p></td><td class="cell"><p>2</p></td><td class="cell"><p>0</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>3 - 1</b></p></td><td class="cell"><p>94</p></td><td class="cell"><p>89</p></td><td class="cell"><p>5</p></td><td class="cell"><p>0</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Revised 0-1</b></p></td><td class="cell"><p>116</p></td><td class="cell"><p>80</p></td><td class="cell"><p>34</p></td><td class="cell"><p>2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>TOTAL</b></p></td><td class="cell"><p>222</p></td><td class="cell"><p>180</p></td><td class="cell"><p>40</p></td><td class="cell"><p>2</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Term values</b></p></td><td class="cell"><p><b>English</b></p></td><td class="cell"><p><b>Basque</b></p></td><td class="cell"><p><b>English</b></p></td><td class="cell"><p><b>Basque</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 0</b></p></td><td class="cell"><p>22</p></td><td class="cell"><p>4</p></td><td class="cell"><p>-</p></td><td class="cell"><p>-</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Revised 0 - 1</b></p></td><td class="cell"><p>-</p></td><td class="cell"><p>-</p></td><td class="cell"><p>160</p></td><td class="cell"><p>143</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 1</b></p></td><td class="cell"><p>2</p></td><td class="cell"><p>3</p></td><td class="cell"><p>2</p></td><td class="cell"><p>4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>1 - 3</b></p></td><td class="cell"><p>7</p></td><td class="cell"><p>3</p></td><td class="cell"><p>-</p></td><td class="cell"><p>-</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>3 - 1</b></p></td><td class="cell"><p>-</p></td><td class="cell"><p>-</p></td><td class="cell"><p>3</p></td><td class="cell"><p>3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>TOTAL</b></p></td><td class="cell"><p><b>31</b></p></td><td class="cell"><p><b>10</b></p></td><td class="cell"><p><b>165</b></p></td><td class="cell"><p><b>150</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="6"/><p>Figure 1: Subhierarchy of WNTERM, with terms added from the BDST (in bold) and their hypernym links in the MCR (in italics).</p><p>Later on, and in case the domain experts deem it necessary, the hierarchy can be changed to suit the particularity of the domain. Alternatively to WordNet hierarchy, SUMO or Top Ontology (both of which are linked to WordNet synsets in the MCR) can also be used, depending on the user needs.</p><p>As a result, a new wordnet (more precisely, an ecological domain ontology) has been constructed with its own set of concepts and relations, which is linked to the rest of wordnets via the ILI.</p></subsection></section><section number="5." title="Conclusions and future work"><p>In this paper we have presented the construction methodology of WNTERM, a science and technology light-weight ontology based on the MCR architecture which has been built with the combination of the Basque and English wordnets and the BDST. The construction methodology has been illustrated on the ecology domain. The final result is a multilingual (Basque and English) light-weight domain ontology, including taxonomic and other semantic relations among its concepts, which is tightly connected to the English WordNet 3.0 and other wordnets. The study also reveals that both wordnets and the BDST are mutually enriched in the process. In the future, we would like to reorganize the hierarchy coming from the MCR, that is, to design the terminological hierarchical organization according to the criteria of domain experts (lexicographers working on the BDST). We also want to develop a semi-automatic framework based on domain corpora (Vossen, 2001), in order to speed up the construction of other domain ontologies. WNTERM will be used as the evaluation benchmark for the automatic procedures.</p></section><section title="Acknowledgements"><p>We would like to thank Mari Susperregi and Pili Lizaso who have collaborated in this research. Eli Pociello has a grant from the Provincial Council of Gipuzkoa. Part of this work has been funded by the European Commission (KYOTO, ICT-2007-211423) and the Education Ministry (KNOW, TIN2006-15049-C03-01).</p></section><references><p>Agirre, E., Aldezabal, I., Etxeberria, J., Izagirre, E., Mendizabal, K., Quintian, M. &amp; Pociello, E. (2006). A methodology for the joint development of the Basque WordNet and Semcor. In <i>Proceedings of the 5th International Conference on Language Resources and Evaluations (LREC). </i>Genoa (Italy). Alegria, I., Areta, N., Artola, X., Diaz de Ilarraza, A., Ezeiza, N., Gurrutxaga, A., Leturia, I. &amp; Sologaistoa, A. (2007). ZT Corpus: Annotation and tools for Basque corpora. In <i>Corpus Linguistics 2007. </i>Birmingham (England). Alegria, I., Gurrutxaga, A., Lizaso, P., Saralegi, X.,</p><doubt alpha="57.4" length="47" tooSmall="False" monospace="0.0">Ugartetxea, S. &amp; Urizar, R. (2004). A Xml-Based</doubt><p>Term Extraction Tool for Basque. In <i>LREC2004: 4th International Conference On Language Resources And Evaluation. </i>Lisbon (Portugal). Alvez, J., Atserias, J. , Carrera, J., Climent, S., Oliver, A. &amp; Rigau, G. (2008). Consistent annotation of EuroWordNet with the Top Concept Ontology. In <i>Proceedings of The 4th Global Wordnet Association Conference. </i>Szeged (Hungary). Atserias, J., Villarejo, L., Rigau, G., Agirre, E., Carroll,<page local="7"/></p><doubt alpha="56.5" length="46" tooSmall="False" monospace="0.0">J., Magnini, B. &amp; Vossen,P.(2004). The MEANING</doubt><p>Multilingual Central Repository. In <i>Proceedings of the 2nd Global WordNet Conference. </i>Brno (Czech Republic).</p><p>Borst, W.N. (1997). Construction of Engineering Ontologies. Centre for Telemätica and Information Technology, University of Tweety. Enschede, The Netherlands.</p><p>Daude, J., Padrö, L. &amp; Rigau, G. (2001). A Complete WordNet1.5 to WordNet1.6 Mapping. In <i>Proceedings of NAACL'01 Workshop on WordNet and Other Lexical ressources. </i>Pittsburgh (Pennsylvania).</p><p>Decker, S., Erdmann, M., Fensel, D. &amp; Suder, R. (1999). Ontobroker: Ontology Based Access to Distributed and Semi-Structured Information. In <i>Proceedings of Semantic Issues in Multimedia Systems </i>(DS-8). Rotorua (New Zealand).</p><p>Fellbaum, C. (1998). <i>WordNet: An electronic Lexical Database. </i>The MIT Press, Cambridge, Massachusetts. London (England).</p><p>Gruber, T.R., &amp; Olsen, F. (1994). An ontology for Engineering Mathematics. In <i>Proceedings of the Fourth International Conference on Principles of Knowledge Representation and Reasoning. </i>Bonn (Germany).</p><p>Magnini, B. &amp; Cavagliä, G. (2000). Integrating subject field codes into WordNet. In <i>Proceedings of LREC, </i>pages 1413-1418. Athens (Greece).</p><p>Navigli, R., Velardi, P. &amp; Gangemi, A. (2003). Ontology Learning and Its Application to Automated Terminology Translation. In <i>IEEE Intelligent Systems, </i>18 (1), pp. 22-31.</p><doubt alpha="61.1" length="54" tooSmall="False" monospace="0.0">Niles, I. &amp; Pease, A. (2001). Towards a standard upper</doubt><p>ontology. In <i>Proceedings of the 2nd International Conference on Formal Ontology in Information Systems, </i>pp. 17-19.</p><doubt alpha="52.1" length="48" tooSmall="False" monospace="0.0">Rigau, G., Agirre, E. &amp; Atserias, J. (2003). The</doubt><p>MEANING project. In <i>Proceedings of the XIX Congreso de la Sociedad Espanola para el Procesamiento del Lenguaje Natural (SEPLN). </i>Alcala de Henares (Madrid). Roventini, A., Marinelli, R. (2004). Extending the Italian WordNet with the Specialized Language of the Maritime Domain. In <i>Proceedings of the 2nd Global WordNet Conference. </i>Brno (Czech Republic).</p><doubt alpha="59.6" length="99" tooSmall="False" monospace="0.0">Sagri, M. T., Tiscornia, D. &amp; Bertagna, F. (2004). Jur-WordNet. In Sojka, P. et al. (Eds.) InSecond</doubt><p><i>International WordNet Conference. </i>Brno (Czech Republic).</p><p>Stamou, S., Ntoulas, A., Hoppenbrouwers, J., Saiz-</p><p>Noeda,  M.   &amp;   Christodoulakis,   D. (2002).</p><p>EUROTERM: Extending EWN using the expand and merge model. In <i>Proceedings of the 1st Global WordNet Conference. </i>Mysore (India). Uschold, M., King, M., Moralee, S. &amp; Zorgios, Y. (1998) The Enterprise Ontology, Knowledge Engineering Review, 13(1). 31-89.</p><p>Vossen, P. (1998). <i>EuroWordNet: A Multilingual Database with Lexical Semantic Networks. </i>Kluwer Academic Publishers. Vossen, P. (2001) Extending, Trimming and Fusing WordNet for Technical Documents. In <i>Proceedings of the NAACL Workshop on Extending Wordnet. </i>Pittsburgh (Pennsylvania).</p></references></body></article>