<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>For a Repository of NLP Tools</title><author surname="Chaudiron" givenname="Stéphane"><org  name="Ministère de la Recherche &amp; Université de Paris" country="France"/></author><author surname="Choukri" givenname="Khalid"><org  name="Ministère de la Recherche &amp; Université de Paris" country="France"/></author><author surname="Mance" givenname="Audrey"><org  name="Ministère de la Recherche &amp; Université de Paris" country="France"/></author><author surname="Mapelli" givenname="Valérie"><org  name="Ministère de la Recherche &amp; Université de Paris" country="France"/></author></firstpageheader><frontmatter><p>F<b>or a repository of NLP tools</b></p><p><b>Stéphane Chaudiron*, Khalid Choukri+, Audrey Mance+, Valérie Mapelli+</b></p><p>*Ministère de la Recherche &amp; Université de Paris 10 - CRIS 200, avenue de la République 92001 Nanterre cedex, France stephane.chaudiron@u-paris10.fr +ELRA/ELDA 55-57, rue Brillat-Savarin, 75013 Paris, France {choukri, mance, mapelli}@elda.fr</p></frontmatter><abstract>In this paper, we assume that the perspective which consists of identifying the NLP supply according to its different uses gives a general and efficient framework to understand the existing technological and industrial offer in a user-oriented approach. The main feature of this approach is to analyse how a specific technical product is really used by the users and not only to highlight how the developers expect the product to be used. To achieve this goal with NLP products, we first need to have a clear and quasi-exhaustive picture of the technical and industrial supply. During the 1998-1999 period, the European Language Resources Association (ELRA) conducted a study funded by the French Ministry of Research and Higher Education to produce a directory of language engineering tools and resources for French. In this paper, we present the main results of the study. The first part gives some information on the methodology adopted to conduct the study, the second part presents the main characteristics of the classification and the third part gives an overview of the applications which have been identified. achieve this goal with NLP products, we first need to have a clear and complete inventory of the technical and industrial supply. We can therefore define the term "context of use" as referring to the social and individual appropriation of a technical object, Perriault (1989) and Harvey (1995) give some examples of this approach. Concerning NLP applications, this methodological approach gives results which are the first step for a future user-oriented evaluation. Inthis paper, we present the main results of the study<footnote anchor="1"/> according to a classification based on the contexts of use. The first part gives some information on the methodology adopted to conduct the study; the second part presents the main characteristics of the classification and the third part gives an overview of the applications which have been identified. These studies are mainly conducted with an underlying philosophy which aims at filling the gap between research and industry, increasing economic efficiency within the European Community and identifying the measures required at government or supra-government levels to achieve these potential economic benefits. Some of these studies try to identify the different perspectives, the user demand, the industrial supply and the technological offer, but none of them addresses the problem in terms of <i>contexts of use. </i>During the 1998­1999 period, the European Language Resources Association (ELRA) conducted a study funded by the French Ministry of Research and Higher Education to produce a directory of language engineering tools and resources for French. Within this study, we assume that such a perspective, which consists in identifying the NLP supply according to its different uses, gives a general and efficient framework to better understand the existing technological and industrial offer in a user-oriented approach, even if we devote some sections to list some of the tools and the corresponding typology identified during the survey. The main feature of the user-oriented approach is to analyse how a specific technical product is really used by the users and not only to point out how the developers expect the product to be used. But, to </abstract></header><body><section number="2." title="Methodological features of the study"><p>The objective of the ELRA study was to identify the information processing tools that were elaborated by industry or research laboratories for the French language. Possible extensions to other languages are under discussion with some partners. In order to achieve this study, ELRA followed the following steps:</p><p>- Definition of a classification for the tool typology; - Identification of potential producers and tools; - Designing a tool description form; - Drafting a questionnaire; - Carrying out the survey; - Analysis of the information collected;</p><p>- Structuring the final information set as a database.</p><subsection number="2.1" title="Typology"><p>This first task consisted of listing the different categories of tools for the Natural Language Processing (NLP) field, also trying to determine the relations between these categories. The main focus point was to define whether a tool indicated as an NLP tool did or did not include a language component.<page local="2"/> Therefore, we considered to be NLP tools either computerised tools which process language (e.g. analysis or translation systems) or tools that use language knowledge to process information.</p><footnote label="1">A complete published version is under press and will be available soon.</footnote><p>In this study, a distinction had to be made between "language resources", such as corpora (speech or text), electronic dictionaries, glossaries, grammars, etc., and natural language processing tools, which allow to analyse, generate, understand, evaluate, extract, translate, etc. all kinds of information.</p><p>Finally, we established a list allowing to distinguish the different categories of NLP tools, including a special part concerning language resources. We chose not to make a hierarchical list of tools since it was too difficult to settle on the relationship between the tools. Therefore, we opted for a linear presentation of the tool categories. The main top categories that could be distinguished are the following: language resources, language analysis, automatic generation, automatic translation, automatic summarisation, language understanding systems, terminology management, speech processing, information management and retrieval, computer-aided authoring tools, optical character recognition, computer-aided learning, system and resource building, NLP systems evaluation.</p></subsection><subsection number="2.2" title="Prospect list"><p>Preparing the list of prospects was not an easy task. They were namely extracted from the ELRA contact database, which consists of more than 1200 contacts all over the world and over 300 for France; the list of organisations members of the Aupelf-Uref, now known as AUF (Agence Universitaire de la Francophonie) (1998); a few directories were provided by the French Ministry of Research and Higher Education; <i>The Language Engineering Directory </i>(Hearn, 1996) was also used.</p><p>More information was also extracted directly from several Web sites.</p><p>The use of these different sources allowed us to collect a good amount of information about contact persons, available tools, organisation's profile, etc. That also helped to carry out a targeted survey.</p></subsection><subsection number="2.3" title="Questionnaire"><p>In order to contact the different players in the NLP field, we decided to draft a specific questionnaire. This questionnaire consisted of five parts:</p><p>1. Tool identification: this part includes the minimum information required to identify a tool, in particular the tool name, type of tool (is the tool a language resource, an application, or a software?), tool category (according to the typology that we defined and that was given as an annex to the questionnaire), usage (potential users), availability, language(s).</p><p>2. Provider identification: this part includes all necessary information concerning the provider, i.e. organisation's name, contact, address, etc.</p><p>3. Technical and commercial information: this section requires information about the medium, the size of the data, the workstation, documentation available, constraints for distribution, etc.</p><p>4. Detailed description: this blank section (limited to a maximum of 3 complete pages) helps provide extra technical and linguistic information.</p><p>5. Free description: a blank field was added to help the prospects add more details, such as related information, bibliography, information sources, other existing tools that they were aware of.</p></subsection><subsection number="2.4" title="Survey"><p>Once the contact list had been completed, we decided to directly contact each player that we had identified by sending them the description form. We only completed the contact section so that the contacted person could freely fill in each part of the form.</p></subsection><subsection number="2.5" title="A catalogue of tools"><p>As soon as the description forms were collected (either under Word format or as hardcopies), all information was gathered in a specific database under Access-97. Beyond the description forms that we managed to collect, we decided to add other tools that were found during the study. This directory of language processing tools for the French language may be completed later on with a list of tools from French speaking countries that will be carried out by the FRANCIL network (AUF). Other studies might also be carried out in the framework of the European HLT programmes with the co-operation of the DFKI (Deutsches Forschungszentrum für Künstliche Intelligenz, University of Saarbrücken).</p></subsection><subsection number="2.6" title="Similar studies"><p>Several studies were carried out in the field of language processing tools. These can be divided into two different areas:</p><subsubsection number="2.6.1" title="French speaking organisations' studies"><p>Other interesting information can be found at four main French speaking organisations: the OFIL (Office Français des Industries de la Langue, France) which published the <i>Guide des produits et services d'ingénierie linguistique </i>(language engineering products and services guide); the DGLF (Délégation Générale à la Langue Française, France) has a Web site that includes a directory of French firms and research centres in the language engineering field (http://www.culture.fr); the RIOFIL (Réseau International des Observatoires Francophones des Industries de la Langue, Québec) which enquired about the language resources offer and the needs from the natural language processing players' point of view; the FRANCIL network (Réseau FRANCophone de l'Ingénierie de la Langue, AUF) which is working to expand the ELRA tool directory, that focused on French organisations, to the French speaking countries.</p></subsubsection><subsubsection number="2.6.2" title="European studies"><p>The European Language resources Distribution Agency (ELDA) offers on its Web site (http://www.elda.fr) a catalogue of language resources including different types of language resources: speech and written corpora, monolingual and multilingual lexica and terminology databases.  Language and<page local="3"/></p><p>Technology (Spain) constituted a language engineering directory, which consists of a thousand language tools and resources and over 600 language engineering organisations, on behalf of the European Community within the  framework  of the  MLIS programme (MultilinguaL Information Society - CE- DGXIII). This directory is also available on the Web (http://www2.echo.lu/mlis/fr/direct/home.html). The DFKI is currently working on a tool directory, namely the <i>Natural Language Software Registry </i>(http://corp-200.dfki.uni-sb.de/lt/registry).</p></subsubsection><subsubsection number="3.1.1" title="Basic tools"><p>This first class concerns the basic components which implement very precise and limited linguistic function as lemmatisation tools, morpho-syntactic and/or semantic analysers, terminology extractors, etc. The linguistic transformations may use complex technologies but these modules do not create any informational added-value, or very little if any.</p><p>In this sense, the output of the system is not the result of information processing but is the result of a linguistic data processing. These components are developer-oriented and are used by software engineers.</p></subsubsection><subsubsection number="3.1.2" title="Linguistic agents"><p><i>Linguistic agents </i>are off-the-shelf software or modules which achieve complex informational tasks (help to text-editing, text translation, query translation, filtering,.). Off-the-shelf products are user-oriented whereas modules are developer-oriented.</p><p>By their own or integrated in an information processing platform, they not only achieve linguistic data processing, but a process of information management. For example, translation is not only considered as a single process of homothetic transfer between phrases but fits into a context of information exchange. Information retrieval or information filtering and routing may use linguistic technologies and, if they do, create a high added-value in the context of information processing.</p><p>At the present time, linguistic agents are the most important technological and commercial supply source.</p></subsubsection><subsubsection number="3.1.3" title="Integrated applications"><p>The third class of applications concerns what we call <i>integrated applications. </i>These applications do not only achieve specific linguistic and informational functions as linguistic agents but process more complex tasks for information processing, or even knowledge management. It concerns for example a process of strategic information intelligence in a multilingual environment using several linguistic agents to perform the different tasks at the various steps : search engine, data extraction, translation tool, information filtering and routing for example. Integrated applications process the "information content" using language knowledge.</p></subsubsection></subsection></section><section number="3." title="Three Classes of NLP applications"><p>In this paper, we use the term "application" as a generic term designing all kinds of automatic processing, including NLP techniques or technologies. A NLP product aims at extracting information from linguistic data (as input of the system). The output of the system may be either the result of a linguistic transformation (translation, filtering, information retrieval,...), or a state transformation of a complex system (NL interface with a security system for nuclear central, voice control of a fighter plane,...).</p><p>Considering the use of NLP for professional information managing, we can distinguish three different classes of products and services. The low level, or <i>basic tools </i>which produce weak added-value to the information processing ; the <i>linguistic agents, </i>or lingaware, which both address linguistic and informational tasks and the <i>integrated applications </i>which use linguistic technologies at different steps of the information processing.</p></section><section number="4." title="Study outcome: a directory of language processing tools"><subsection number="4.1" title="Some figures"><p>Most of the contacts were from commercial organisations (about 62%) compared with research laboratories (38% of responses). Only 83 out of 253 organisations answered by returning the questionnaire duly completed (about 33% response rate), of which there were 47 commercial organisations and 36 research organisations. A total of 161 questionnaires were collected, out of which 143 that described tools and 18 that described language resources. Among the commercial organisations that answered the survey, we could distinguish between SMEs, large French groups, French subsidiaries of international groups, from different activity fields, such as computational linguistics, research and development, information management, speech technologies, software distribution. As for the academic organisations, those were mainly from the written field.</p><p>The table below gives the number of tools identified within the study according to their top category. A total of 261 tools and resources were identified within the survey (240 tools and 21 language resources). In addition, ELRA collected information on tools for which no questionnaire was completed. These appear in our survey under the label "Tools identified by ELRA". Moreover, some tools belong to several domains of the table but are counted as one single tool in our directory (such as COATIS which is both a terminology consolidation tool and a semantic analyser). Table 1 summarises the findings and results in a total of 283 tools.</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Tool category</b></p></td><td class="cell"><p><b>Completed forms</b></p></td><td class="cell"><p><b>Tools identified by ELRA</b></p></td><td class="cell"><p><b>Total</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>1. Language resources</p></td><td class="cell"><p>19</p></td><td class="cell"><p>3</p></td><td class="cell"><p>22</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>2. Language analysis</p></td><td class="cell"><p>19</p></td><td class="cell"><p>9</p></td><td class="cell"><p>28</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>3. Automatic generation</p></td><td class="cell"><p>4</p></td><td class="cell"><p>5</p></td><td class="cell"><p>9</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>4. Machine translation</p></td><td class="cell"><p>12</p></td><td class="cell"><p>10</p></td><td class="cell"><p>22</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>5. Automatic summarisation</p></td><td class="cell"><p>2</p></td><td class="cell"><p>2</p></td><td class="cell"><p>4</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>6. Language understanding</p></td><td class="cell"><p>3</p></td><td class="cell"><p>0</p></td><td class="cell"><p>3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>7. Terminology management</p></td><td class="cell"><p>16</p></td><td class="cell"><p>4</p></td><td class="cell"><p>20</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>8. Speech processing</p></td><td class="cell"><p>23</p></td><td class="cell"><p>18</p></td><td class="cell"><p>41</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>9. Information management and</p></td><td class="cell"><p>39</p></td><td class="cell"><p>6</p></td><td class="cell"><p>45</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="4"/><p><i>(Systran </i>or <i>Reverse Pro </i>(Softissimo), <i>Power Translator® </i>or <i>iTranslator </i>(Lernout &amp; Hauspie), <i>LIDIA </i>and <i>C-Star II </i>(GETA), <i>TACT </i>(CRLT, Franche-Comté)); (ii) computer-aided translation systems, such as <i>An-</i> <i>Nakel Al-Arabi </i>and <i>European Translator </i>(CIMOS), <i>TRANSIT </i>(Star), <i>Translator's Workbench </i>(Trados).</p><subsubsection number="4.2.4" title="Automatic summarisation systems"><p>Today, no summarisation system is reliable enough to answer general needs. Existing systems are designed to process homogeneous corpora for very precise tasks but cannot deal with heterogeneous corpora.</p><p>Known summarisation systems are <i>SAFIR </i>(Cams-Lalic and EDF), <i>Ciceron </i>(CORA), <i>SummarizerTM </i>(Inxight - Xerox), <i>RAFI </i>(Landisco, Nancy).</p><p>In order to offer an as exhaustive directory as possible, ELRA completed the information sent by the contacted persons with some extra information mainly gathered from the Web.</p></subsubsection></subsection><subsection number="4.2" title="List of tools"><p>ELRA's study led to the creation of a tool directory where tools were classified according to the typology defined within the survey. Examples of identified tools, ranked according to their top category, are given below.</p><subsubsection number="4.2.1" title="Language analysers"><p>Analysers are basic modules for most language processing systems. The identified tools were ranked according to different levels of analysis. Among existing morphological and morpho-syntactical analysers we can quote the <i>CRISTAL morphological</i> <i>analyser </i>(GRESEC, Grenoble 3), the <i>AMFLEX</i> <i>lemmatiser </i>(IRIT, Toulouse 3), analysers from CIMOS and Xerox, <i>MAUD </i>(LORIA, Nancy), <i>Labelgram </i>(CRLT), <i>EtiWeb </i>(LIA, Avignon). As for syntactical analysers, these come mainly from CORA, Xerox, LIMSI. Identified semantic or pragmatic analysers come also from LIMSI and Xerox and, notably, the Tropes semantic analysis software from Acetic.</p></subsubsection><subsubsection number="4.2.2" title="Automatic generation systems"><p>These systems allow the production of textual data in a natural language form. They are mainly used in translation and summarisation systems. We opted for two categories of generation tools: morphological generation and text generation. Examples of morphological generation tools are <i>Lemma le fléchisseur </i>(CORA), an inflected form generation tool (IGM, Marne-la-Vallée), and conjugation systems from CIMOS. Autonomous text generation systems are few and are often integrated into other systems like translation systems.  These are <i>Flaubert </i>(CORA), <i>CRISTAL generator </i>(GRESEC).</p></subsubsection><subsubsection number="4.2.3" title="Machine translation systems"><p>Machine translation can be used in various applications: translation of technical documents, multilingual information processing, information retrieval, etc. These systems can be classified into two main categories: (i) machine translation tools, that automatically generate a text in a target language</p></subsubsection><subsubsection number="4.2.5" title="Language understanding systems"><p>Understanding systems are composed of different NLP modules, such as analysers and generators. They are used as a basis to man-machine dialogue systems, translation, summarisation, knowledge retrieval systems, etc. and have been developed for either written or spoken data. For instance, <i>ILLICO </i>(LIM, Marseille), a <i>prototype for NLP multi-agent system </i>(GRESEC,</p><p>Grenoble), and <i>ItiSACT </i>(IASC, ENST Bretagne) were designed for written tasks, whereas <i>Openvox SLS </i>(Vecsys), <i>DictaMed </i>(GREYC, Caen), <i>C-Star II </i>(GETA, Grenoble 1) were designed for speech understanding, recognition and synthesis.</p></subsubsection><subsubsection number="4.2.6" title="Computer-aided authoring systems"><p>Computer-aided authoring systems can be classified into two different categories, computer-aided checking and computer-aided authoring systems. Many spell and grammar checkers are available on the market. Among them can be found <i>Voltaire </i>(CORA), <i>Hugo Plus </i>(Softissimo), <i>Pro Lexis </i>(Editions Diagonal), <i>Sans-Faute/Grammaire </i>(Bcdl - Hexacom), <i>Cordial </i>and <i>Lexical </i>(Synapse), <i>Vortex </i>(IRIT, Toulouse 3), <i>OrthoNet </i>(CILF), <i>ADICO® médical </i>(iES). As for computer-aided writing systems, these are in particular <i>Vitipi </i>(IRIT,</p><p>Toulouse 3), <i>Unitype </i>(Softissimo), <i>NTK.FOCUS</i> (Nemesia), <i>Euro-Letter Professional </i>(distributed by Apsydoc).</p></subsubsection><subsubsection number="4.2.7" title="Speech processing systems"><p>We decided to rank speech processing tools in four main categories: (i) analysis systems, for example <i>SNORRI  </i>(LORIA,  Nancy),   <i>WaveEdit </i>(GEOD,</p><p>Grenoble 1), <i>Phonedit </i>(LPL, SQLab); (ii) recognition systems, namely <i>CK10.5 </i>(Parrot SA), <i>DragonDictate V3 </i>and <i>Dragon Naturally Speaking </i>(Dragon), <i>ViaVoice </i>(IBM), <i>Voice Xpress Profesional </i>(Lernout &amp; Hauspie); (iii) synthesis systems, including <i>Syntaix </i>(LPL, Provence), <i>Lia_phon </i>(LIA, Avignon), <i>KALI </i>(Elsap, Caen), <i>ELAN Text to Speech </i>(Elan Informatique); and (iv) dialogue systems, such as <i>Openvox SLS </i>(Vecsys), <i>DictaMed </i>(GREYC, Caen), <i>C-Star II </i>(GETA,</p><p>Grenoble 1). Well-known tools from LIMSI are not described in this first release of the inventory.</p></subsubsection><subsubsection number="4.2.8" title="Computer-aided language learning (CALL) systems"><table caption="Table 1. Number of identified tools" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>retrieval</p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>10. Computer-aided authoring</p></td><td class="cell"><p>9</p></td><td class="cell"><p>8</p></td><td class="cell"><p>17</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>11.        Optical character recognition</p></td><td class="cell"><p>7</p></td><td class="cell"><p>2</p></td><td class="cell"><p>9</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>12.  Computer-aided language learning (CALL)</p></td><td class="cell"><p>12</p></td><td class="cell"><p>26</p></td><td class="cell"><p>38</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>13.    System    and resource building</p></td><td class="cell"><p>15</p></td><td class="cell"><p>3</p></td><td class="cell"><p>18</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>14. NLP systems evaluation</p></td><td class="cell"><p>1</p></td><td class="cell"><p>0</p></td><td class="cell"><p>1</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>15. Other tools</p></td><td class="cell"><p>4</p></td><td class="cell"><p>2</p></td><td class="cell"><p>6</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>TOTAL</p></td><td class="cell"><p>185</p></td><td class="cell"><p>98</p></td><td class="cell"><p>283</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="5"/><p>These systems are used for a variety of learning tasks and are consequently numerous. Some of them use written and speech processing modules. They were classified into five different categories: teaching French as a foreign language, French language learning, professional training, computer-aided communication, authoring systems. Among them, we can quote <i>TeLL me More </i>and <i>Talk to me </i>(Auralog), <i>Alexia </i>(LIB, Franche-Comté), <i>Alfy </i>or <i>Orthogram </i>(Chrysis), <i>Copie Double </i>(Goto),    <i>French    Connexions    </i>(Vektor), <i>Medmed</i> <i>multimédia </i>(CRIM, INALCO), <i>ALEx </i>(IASC, ENST</p><p>Bretagne), <i>Kombe </i>(Prologia), <i>Speaker Auteur </i>(Neuroconcept), <i>Amical </i>(LRL, Clermont 2).</p></subsubsection><subsubsection number="4.2.9" title="Terminology management systems"><p>Automatic processing of terminology (terms and semantic relations between terms) can have various possible applications, like elaboration or enrichment of knowledge bases, information and knowledge retrieval, text indexing, translation of technical texts, etc. Some tools can allow terminology data acquisition from existing databases: <i>LEXTER </i>and <i>COATIS </i>(EDF), <i>ACABIT </i>(IRIN, Nantes), <i>RCFilter </i>(Linguanet), <i>IOTA </i>(MRIM, Grenoble), <i>Seek </i>(IDIST-CREDO, Lille 3).</p><p>Other tools are dealing with knowledge and semantic relations extraction: <i>LexiTrack </i>and <i>LexiBuild </i>(LexiQuest), <i>Prométhée </i>(IRIN, Nantes), <i>STK </i>(GREYC, Caen). There also are more complete tools such as <i>SPI-Graphe </i>(CEA), <i>Dixit </i>(Terminotics), <i>MultiTerm'95 Plus! </i>(Trados), <i>Termstar </i>(Star), <i>Ztermino </i>(Lilla, Nice), <i>ERACLES </i>(INIST) and <i>Lexp </i>(LCI).</p></subsubsection><subsubsection number="4.2.10" title="Optical Character Recognition systems"><p>An increasing need is expressed to computerise information (in particular forms of old manuscripts). Scanners and optical character recognition software allow to transform these hardcopies into electronic documents. Two types of OCR were identified: typed character recognition where identified tools are <i>ICR Suite Pro© </i>(SWT), <i>EasyReader Elite </i>(Mimetics), <i>IrisPen </i>and <i>ReadIris </i>(IRIS) <i>TextBridge Pro 98 </i>(ScanSoft), <i>OmniPage Pro 9.</i><i>0 </i>(Caere) and hand writing recognition from which only one tool was identified, namely <i>Solare </i>(LERISS, Paris 12).</p></subsubsection><subsubsection number="4.2.11" title="Information management and retrieval"><p>Information processing is of a paramount importance in terms of strategy and economy. More precisely, having a quick access to a pertinent information, especially intranet or internet information is required by all professional sectors. Different types of tools were identified in this field: indexing tools, information retrieval and filtering, natural language query systems (including Internet search engines), text mining. Many products are offered to the users, from which  <i>Search'97  </i>(Verity),  <i>Intuition   </i>and <i>Darwin</i> (CORA), <i>Alchemy </i>(Bva), <i>eXtense </i>(Echo), <i>ANAGRAM</i> (Triel), <i>DigOut4U </i>and <i>Class4U </i>(Arisem), <i>SPIRIT</i> (Tgid), <i>LexiQuest </i>(LexiQuest), <i>WINDEX </i>(Multimédia</p><p>Solutions), <i>ECILA </i>(Ecila), <i>Lokace </i>(Forlog), <i>Voila </i>(France Telecom), <i>Halpin </i>(GEO, Grenoble 1), <i>Ulysse </i>(GREYC, Caen), <i>Gargantua </i>(Siatel), <i>Docubase Entreprise </i>(Docubase Systems), <i>CinDoc </i>(Cincom).</p></subsubsection></subsection></section><section number="5." title="Maturity of language tools"><p>The present study was completed by another survey conducted by ELDA aiming at analysing the maturity of language engineering tools within the French market. "Maturity" is a vague notion, since a language technology can be considered as mature or not depending on its application and on the targeted users. It is precisely the users and the way technology meets their needs that will define a tool or an application as mature or not. Some systems are not widely used, either because they have unsatisfactory performance or their application, if any, do not fulfil users' requirements.</p><p>However, faced with the huge amount of information available, users will need Natural Language systems and interfaces to access data in a more intuitive way. They will also need tools to structure and manage information, and to disseminate it.</p><p>Among the tools identified during the survey that would be adequate for successful technology transfers, we may quote:</p><p>- Text generation systems, which could be associated with translation, OCR or speech recognition systems, to facilitate the dissemination of multilingual information.</p><p>- Machine translation systems, associated with Internet search engines, will enable people to access information in any language. For professional translators, the development of controlled languages should improve the performance of MT systems.</p><p>- Text summarisation systems could be integrated in information management systems or search engines in order to help users select information.</p><p>- Language understanding systems should enable users to do searches in a more intuitive way, without being obliged to constantly rephrase their requests.</p><p>- The integration of spell/grammar checkers into OCR and machine translation systems should improve their performance.</p><p>- Speech recognition and voice synthesis systems could be combined with oral translation systems to make multilingual access to information possible.</p><p>- OCR systems should improve their performance by becoming "hybrid", that is processing both hand writing and typed characters. frequent as data input mode, data output being made by voice synthesis.</p></section><section number="6." title="Conclusion"><p>This survey allowed to identify a large number of NLP tools. Not all of them are or will be available for technology transfer that may turn them into useful and (best-)selling products. Many are still prototypes or research components. Nevertheless, a clear panorama may help streamline the efforts required to do so. No specific assessment has been conducted to evaluate the performance of these tools neither in terms of technology nor in terms of usefulness and usability.</p><p>An evaluation paradigm is still necessary to measure how mature such tools are before incorporating them into information processing systems. Here we should insist that "non-mature" NLP tools may be sufficient for a large number of applications but the level of maturity and performance should be measured and explained to the application integrators.</p><page local="6"/></section><section number="7." title="Acknowledgements"><p>We would like to thank the French Ministry of Research and Higher Education that has allowed ELDA to carry out this valuable survey work. Our gratitude is also extended to Emilie Marquois and Métiyé Meydan for their contribution to the survey.</p></section><references><p>Bossard Consultants, <i>Industrie de la langue : marché et</i></p><p><i>perspectives à 5 ans, </i>décembre 1988. Euromap,   <i>The  Euromap  Report :   Challenge and</i> <i>Opportunity for Europe's Information Society, </i>CCE -</p><p>DG  XIII,   Telematics   Applications Programme,</p><p>september 1998. Ink International and ECC, <i>Language industries survey,</i> 1989.</p><p>Mlis, <i>Language Engineering Directory, </i>CCE - DG XIII, Multilingual Information Society Prgramme, http://www2.echo.lu/mlis/en/direct/home.html, updated 07/23/1999.</p><p>Ofil,  <i>Guide  des produits et services d'ingénierie</i></p><p><i>linguistique, </i>OFIL, 1994. Ovum,  Engelien,  B.,  McBryde,  Ronnie, <i>Natural</i> <i>Language Markets :</i><i> Commercial Strategies, </i>London,</p><doubt alpha="56.7" length="60" tooSmall="False" monospace="0.0">OVUM, 1991. Ovum,   Lewin,   D.,   Lockwood,   Rose,Language</doubt><doubt alpha="64.9" length="94" tooSmall="False" monospace="0.0">Engineering 2000,London, OVUM, 1993. Owil, Biérin, E., Moulin, A., Pichault, F.,Les Industries</doubt><doubt alpha="66.7" length="48" tooSmall="False" monospace="0.0">de  la  langue :  un  marché  en  devenir,Liège,</doubt><p>Observatoire Wallon des Industriels de la Langue, 1990.</p><p>Perriault, J., <i>La Logique de l'usage : essai sur les machines à communiquer, </i>Paris, Flammarion, 1989.</p><p>Harvey, P.-L., <i>Cyberespace et communautique, </i>Laval, PUL, 1995.</p><p>Hearn, Paul M., <i>The Language Engineering Directory, A resource Guide to Language Engineering Organisations, </i>Products and Services, Madrid, Language and Technology SL, 1996.</p></references></body></article>