<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>Annotation Driven Concordancing: the PAX Toolkit.</title><author surname="Trippel" givenname="Thorsten"><org  name="Department of Linguistics and Literary Studies Bielefeld University Postf" country="Germany"/></author><author surname="Gibbon" givenname="Dafydd"><org  name="Bielefeld ttrippel" country="Germany"/></author></firstpageheader><frontmatter><p><b>Annotation Driven Concordancing: the PAX Toolkit</b></p><p><b>Thorsten Trippel, Dafydd Gibbon</b></p><p>Department of Linguistics and Literary Studies Bielefeld University Postf. 100131 33501 Bielefeld Germany</p><p>ttrippel, gibbon @spectrum.uni-bielefeld.de <b>Abstract</b></p></frontmatter><abstract>We describe PAX, "Portable Audio Concordance System", a proof-of-concept prototype of a multipurpose, multilingual audio concor­dance toolkit. The primary goal is to support efficient grammar and lexicon construction in the documentation of unwritten languages; languages currently included are Ega, Anyi, and Koulango (Ivory Coast), additional samples in German and English. The approach combines methods from corpus linguistics, annotation theory and practice, phonetics and lexicography. </abstract></header><body><section number="1." title="Objectives"><p>Finding occurences of selected utterances in multi­modal corpora for multimodal lexica is the objective ofthe <i>Portable Audio Concordance System </i>(PAX)<footnote anchor="1"/>.</p><p>Modern dictionaries these days claim to be corpus based, for example lexica from the COBUILD project (see (Sinclair, 1987)). This is in the sense that</p><p>1. the order of different meanings corresponds to the fre­quency in a defined corpus</p><p>2. the examples for the use of different words are taken from <i>real world </i>data, i.e. corpora.</p><p>This presupposes a sufficiently preprocessed (marked-up) textual source. For written texts there are a number of corpora used for this purpose such as the <i>British National Corpus </i>(British National Corpus, 2001) for English or the copora available via (COSMAS, 2002) for German.</p><p>However, these corpora contain written texts, and there are concordances for lexical analysis of written texts, which are well known (see for example (van Eynde and Gibbon, 2000)), but no adequate concordancing tools for spoken language exist. The concordancing task for spoken lan­guage is difficult: units are less well identified, access to both transcription text and speech signal is required, and standard aids like word statistics need to be supplemented by visualised transformations of the speech signal.</p><p>We demonstrate an enhanced <i>KeyWord in Context </i>(KWIC) concordance, based on a search space as defined by the annotation graph (Bird and Liberman, 2001), rep­resenting the transcription, and a search, which includes a variety of complex criteria. The XML formalism is based on the TASX formatas described by (Milde and Gut, 2001).</p><p>The position of a concordance in a concordance based multimodal lexicon system is described by Figure 1. Start­ing from the annotation of a multimodal source a lexicon is generated that falls back onto the annotation via the con­cordance for exemplified usages and possibly for evaluation of generated lexicon entries. The annotation itself is used</p><p>'The acronym is derived from PACS by merging the final let­ters.</p><figure caption="Figure 1: Concordances in a multimodal Lexicon system"></figure><p>by the concordance as data input as well, and the annota­tion can be refined within fixed environments by using a concordance. Hence there is a bidirectional connection be­tween the concordance and annotation.</p><subsection number="1.1." title="Methodology"><p>The PAX concordance design is based on a function where is a set of annotated signals partitioned for different languages, KWIC is the keyword in context concordance, and is a set of signal output renderings (audio, waveform, , spectrogram). The corpus consists of digital signals, which were annotated at different levels.</p><doubt alpha="54.5" length="33" tooSmall="False" monospace="0.0">/ : CORPUS ^&lt; KWIC, S IGT BAN S &gt;</doubt><p>The process of concordance generation consists of four functions, namely:</p><p>1. lexicon generation function for spoken lan­guage dictionaries</p><p><i>fiexSL </i><i>■ Corpus </i><i>-&gt;</i><i> Lexicon</i></p></subsection></section><section number="2." title="annotation generation function"><p><i>fanno </i><i>■ Signal —&gt; Annotation</i><page local="2"/></p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1568</doubt></section><section number="3." title="annotation access function"><p><b><i>fîêxSL </i></b>: <i>Lexicon </i><i>—&gt;•</i><i> Annotation</i></p></section><section number="4." title="J 'anno signal access function"><p><i>fanno </i><i>'■</i><i> Annotation </i><i>—&gt;•</i><i> Signal</i></p><p>The normalisation preprocessing function is omitted here; it would be necessary to have this function in order to han­dle annotations produced by different annotators and differ­ent annotation tools and conventions. It would denote the transformation of a source annotation format into a format that can be accessed by and . However, the functions above presuppose a normalised data format.</p><p>The simplest of this lexicon functions can be seen as a generated wordlist with the latent property of each <i>word </i>being <i>in the corpus. </i>The formal lexicon model respects but is not restricted to other lexicon models based on form, meaning or use; especially there are no semantic limits as for example discussed in (Sharoff, 2002).</p><subsection number="1.2." title="Concordance Use in Stand-off Annotation"><p>In the process of first approaching and annotating data it is not uncommon to omit features that do not seem appro­priate to tag or that are irrelevant for the present research task. In a later stage of reusing the corpus and the annota­tions, other features might become relevant (without mod­ifying the original, see (McKelvie and Thompson, 1997)) and consequently need to be checked in correspondence with the original —which means in the context of spoken language with the recorded signal.</p><p>By using existing annotations it should be easy to ac­cess only the relevant parts of the signal for reannotation or further — not necessarily automatized — detailed analysis. In this context it is meant to use the concordance for the preselection of data.</p><p>All that is necessary would be the inverse corpus based lexion function:</p><p><i>fiexSL </i><i>■ Lexicon </i><i>—&gt;•</i><i> Annotation </i>and a function</p><p>An audio concordance needs to provide these two func­tions. If possible further functionality should be added, such as interfaces to phonetic analysis software for auto­matic processing of selected parts of signals and basic cor­pus statistical features, such as word counting and word frequency analysis, also <i>Type-Token-Ratio </i>(TTR) can be added.</p></subsection><subsection number="1.3." title="Funtional Requirement Specification"><p>The functionality of the concordance system sets cer­tain technical and functional requirements for the imple­mentation. Additionally it should respect our general re­quirements for PAX:</p><p><b>Standardisation: </b>For coding standard requirements are used in order to avoid incompatibilities and to promote exchangeability. Proprietory — in the sense of not openly accessible and usable — formats are strongly discouraged.</p><p><b>Interoperability: </b>Support for major platforms such as Unix/Linux, Windows, Macintosh should be aimed at, because these platforms are frequently used by field workers and linguists.</p><p><b>Multifunctionality: </b>The system should be extensible to different requirements and new functionality as needed by the community.</p><p><b>Low-cost/low-end, online/offline: </b>As the target user group includes institutions and persons in areas with­out access to the newest IT infrastructure, the system should be independent of recent software versions or high-end hardware in order to ensure usability under local conditions of this kind. As networking devices might not be available in fieldwork locations, offline functionality is needed</p><p><b>Network Access: </b>For training purposes and for the consis­tency of data, network access is to be provided.</p></subsection></section><section number="2." title="Design"><p>Our design strategy is to use a KWIC (KeyWord In Context) approach, taking an annotated database of speech signals as input, with standard typewriter-friendly SAMPA transcription, XML annotation formats, and a suite of for­mat converters to cope with data input from different cor­pora, or annotators with different software and hardware platforms. Conceptually the approach is not very differ­ent from statistical training procedures used in spoken lan­guage technology, but the requirements are very different in detail.</p><figure caption="Figure 2 shows a detailed overview on the PAX concor-dancing systems design."></figure><p>The PAX architecture is modular, with wordlist extrac­tion and KWIC concordance construction modules (Perl), and signal extraction and processing modules (Java pro­grams and Praat scripts). These modules feed three inde­pendent user interfaces:</p><p>The system consists ofthree basic modules:</p><p><b>Data acquisition module: </b>The data acquisition module calculates corpus information, based on available cor­pora in specified locations. Among other information a wordlist is collected, a list of available annotation tiers and subcorpora. The search procedure uses this information as a basis for further processing, involving the generation of a static (predefined and accessible) and dynamic (on the fly generated) concordances.</p><p><b>Corpus consultation module: </b>The user selects search cri­teria from the information provided by the data acqui­sition module and defines an output filter to specify the size of the context (e.g. the number of words or char­acters left and right of the keyword occurrence). The output of the module is the KWIC concordance. Each<page local="3"/></p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1569</doubt><p>Time stamp</p><doubt alpha="100.0" length="6" tooSmall="False" monospace="0.0">Signal</doubt><p>Request for filters - audio selection - representation of - oscillogram - spectrogram</p><p>Signal selection (time stamp)</p><p>Signal 7 Analysis</p><p>KWIC context is supplemented with a further selec­tion panel specifying a set of choices for access to the signal: a range within the context; waveform or spec­trogramme or pitch track; selection of further output in the same window or a new window.</p><p>The consultation module provides some basic sta­tistical information as well. Based on the distribu­tion function measures of variability are given (Oakes, 1998), such as mean, range and median, standard devi­ation, type/token ratio (for simple tokens). The num­ber ratio of matches is also given. However, the in­terpretation of these measures is left to the user as it is dependent on the data types whether a measures is relevant or not.</p><p><b>Signal analysis: </b>The KWIC output contains selection pan­els for further processing on the signal level; segments are selected on the level of time-stamps assigned to the selected context. These segments are passed to signal processing for further analysis.</p></section><section number="3." title="Implementation"><p>Input data are time-aligned in SAMPA standard ASCII IPA coding (modified for efficient tone language coding), using Praat, Transcriber or esps/waves . These opera­tional formats are converted to an XML annotation graph format (Bird and Liberman, 2001) the TASX DTD (Milde and Gut, 2001), retaining SAMPA coding. TASX-XML and SAMPA fulfil the archive exchange and low-end avail­ability requirements. The TASX DTD is available at: coli.lili.uni-bielefeld.de/~milde/tasx/</p><p>The wordlist extraction, KWIC concordance construc­tion modules and format converters are in Perl, signal ex­traction and processing modules are in Praat scripts and Java. These modules feed three independent GUI modules, see section 3.2..</p><p>This hybrid implementation strategy fulfils the low-cost, low-end, on/offline and interoperability requirements.</p><subsection number="3.1." title="Pragmatic Choice of Programming Language"><p>PAX is implemented with a hybrid component structure, and the programming languages for the components were selected on a pragmatic basis:</p><p><b>Perl: </b>the modules itself are implemented in Perl, as it is available for many computer systems (including UNIX/Linux systems, MS-Windows, MacOS) and re­source friendly. Additionally it provides a rich source of regular expression capabilities and many inter­face format libraries, covering command line access, graphical user interfaces and CGI-access.</p><p>Perl allows access to other system components as well and can be used to call other modules from within the program.<page local="4"/></p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1570</doubt><table caption="Figure 2: Design ofPAX,based on TASX corpus format" class="main" frame="box" rules="all" border="1" regular="True"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Output selection</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>e.g. oscillograms</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>spectrograms, <i>y/</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p><b>Java: </b>large-scale signal processing is impossible with tra­ditional scripting languages such as Perl. For signal processing we selected Java as a suitable language that is system independent. The problem of Java being rel­atively resource unfriendly could be neglected for the present implementation as Java plays a minor role in the toolkit and does not result in a bottle-neck in per­formance. On the other hand no comparable system independence could be accomplished by using other programming languages.</p><p><b>Praat: </b>for signal processing the routines of the Praat pro­gramme for enhancing phonetic productivity are used. These are easy to incorporate since Praat provides a scripting language and interface that can be invoked by other programs. Praat is also available for all major platforms.</p><p><b>R: </b>the statistical functions provided by the R-statistics package are connected to the concordance for effec­tive availability of various statistic functions. R was chosen because it is freely available under GNU pub­lic licence and because it is available for all major plat­forms (Gentleman and Ihaka, 1997 2002).</p></subsection><subsection number="3.2." title="Select the User interface"><p>The PAX implementation exists with different user in­terfaces, all built upon the same core algorithms; they cor­respond to the design specification, including wordlist ex­traction and KWIC concordance modules in Perl, signal ex­traction and processing. They are:</p><p>1. a CGI application with HTML and WAV output for use with a web server</p><p>2. a TK application for offline use and without a local www server</p><p>3. a low end command line based access, also used as an interface for other programs.</p></subsection><subsection number="3.3." title="Interface Structure"><p>The interfaces are clearly structured for easy access.</p><p><b>Enter the concordance system </b>by selecting a language. This directs the system to the designated location, holding a corpus for a specific language in the TASX format.</p><p><b>Selecting a keyword in a subcorpus and tier </b>is possi­ble in the next interface. At the present stage the keyword, subcorpus and tier can be selected independently from each other so the result is only related to the selected subcorpus and tiers. Additionally the user can select the environment by a specifying a numerical values of words before or be­hind the keyword.</p><p><b>Resulting contexts </b>are presented with numberous addi­tional options on the signal analysis containing the word and a further specified context. Additional statistical infor­mation is produced and presented here as well in a simple table form. A sample interface of results is shown in Figure 3.</p><doubt alpha="86.7" length="30" tooSmall="True" monospace="0.0">Statistic outcome of the query</doubt><doubt alpha="56.9" length="58" tooSmall="True" monospace="0.0">Search for                    Language Number of hits Rate</doubt><doubt alpha="15.4" length="39" tooSmall="True" monospace="0.0">ab&lt;u                          EGA 7 000</doubt><doubt alpha="87.8" length="49" tooSmall="True" monospace="0.0">Further statistical analysis on the selected tier</doubt><doubt alpha="72.0" length="93" tooSmall="True" monospace="0.0">Total tokens    Total types     Type/token ratio     Range Mean     Median Standard-Deviation</doubt><doubt alpha="0.0" length="103" tooSmall="True" monospace="0.0">2927              717                 0 24            1 - 116 4 082287        2                9 909799</doubt><p>Figure 3: Keywords in context with additional options for further signal analysis; corpus statistics at the bottom ofthe page <b>Analytic functionality </b>is achieved by the use of Praat, producing appropriate images and value tables.</p></subsection></section><section number="4." title="Evaluation"><p>The toolkit was evaluated following EAGLES guide­lines as defined in (Gibbon et al., 1997), with</p></section><section number="1." title="inhouse testing for correctness ofresults,"><p>2. in-project testing with respect to substantive and er-gonomic user requirements,</p><p>3. extension to quite different corpora, including the VerbMobil German speech database and a German-English language acquisition corpus.</p></section><section number="5." title="Summary and Further Development"><p>The PAX tool was developed specifically as part of an environment for efficiently analysing spoken language, in particular unwritten languages, including African tone lan­guage annotations with tone markup. The specifications for the tool were established on the basis of experience in previous work on encyclopedia modelling for African lan­guages funded by the Deutscher Akademischer Austausch­dienst, work on the efficient analysis of endangered lan­guages funded by the Volkswagen Foundation, and work on the construction of multimodal lexica funded by the Deutsche Forschungsgemeinschaft. At the present time the PAX application contains corpora from five languages, three West African tone languages (Anyi, Ega, Koulango) and two European languages (English, German).</p><p>Current work is directed towards tagging enhancement, modules for further spoken language corpus analysis, and time-aligned multimodal data.</p><p>The next step will include the use ofthe concordance in a multimodal lexicon system of spoken language data. The concordance will be used for creating real-world examples for lexicon entries.</p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1571</doubt><page local="5"/></section><references><p>Steven Bird and Mark Liberman. 2001. A formal frame­work for linguistic annotation. <i>Speech Communication, </i>(33 (1,2)):23-60.</p><p>British National Corpus. 2001. British national corpus. CD-Rom.</p><p>COSMAS. 2002. Corpus storage, maintenance and access</p><p>system. http://corpora.ids-mannheim.de/cosmas/. Robert Gentleman and Ross Ihaka, 1997 - 2002. <i>The</i> <i>R Project for Statistical Computing.</i><i> </i>http://www.r-project.org.</p><p>Dafydd Gibbon, Roger Moore, and Richard Winski, edi­tors. 1997. <i>Handbook of Standards and Resources for Spoken Language Systems. </i>Mouton de Gruyter, Berlin.</p><p>David McKelvie and Henry S. Thompson. 1997. Hyper­link semantics for standoff markup of read-only docu­ments. In <i>Proceedings of SGML Europe'97, </i>Barcelona, May.</p><p>Jan-Torsten Milde and Ulrike Gut. 2001. The tasx-engine: an xml-based corpus database for time aligned language data. In <i>Proceedings of the IRCS Workshop on Linguistic Databases, </i>Philadelphia. University of Pennsylvania.</p><p>Michael P. Oakes. 1998. <i>Statistics for Corpus Linguis­tics. </i>Edinburgh Textbooks in Empirical Linguistics. Ed­inburgh University Press, Edinburgh.</p><p>Serge Sharoff. 2002. Meaning as use: exploitation of aligned corpora for the contrastive study of lexical se­mantics. In <i>Proceedings of LREC 2002, </i>volume this vol­ume, Las Palmas.</p><p>J. M. Sinclair, editor. 1987. <i>Looking up. </i>Collins ELT, Lon­don.</p><p>Frank van Eynde and Dafydd Gibbon. 2000. <i>Lexicon De­velopmentfor Speech and Language Processing. </i>Kluwer Academic Publishers, Dordrecht.</p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">1572</doubt></references></body></article>