<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<TEI xmlns="http://www.tei-c.org/ns/1.0" xmlns:ns2="http://www.tei-c.org/ns/Examples">
    <teiHeader>
        <fileDesc>
            <titleStmt>
                <title>Word Knowledge Acquisition, Lexicon Construction and Dictionary Compilation</title>
            </titleStmt>
        </fileDesc>
    </teiHeader>
    <text>
        <front>
            <div type="abs">
                <head>Abstract</head>
                <p>wl|ich make it possil)le to suplfly infornmtion concerning amenability to diathesis alternations o ~tq to avoid We describe an approach to semiautomatic lexexpanding distinct entries for related uses of the same icon development from |nachine readal)le dictioverb. This practice woldd allow us to develop an I,KB naries with specific reference to verbal diatlleses, from dictionary databases which offers a more co|nenvisaging ways in which tile results obtained can plate and linguistically relined repository of lexical inbe used to guide word classification in the conformation l, hall the source databases. Such an \],Kll strnction of dictionary datal)ases. wouhl be used to generate lexical components for NI,P systems, and couhl also be integrated into a lexicographer's workstation to guide word classification. 1 Introduction The acquisition and representation of lexical knowl2 The ACQUILEX Lexicon Development edge from machine-readable dictionaries and text corEnvironnmnt pora have increasingly become major concerns in Computational Lexicography/Lexicology. While this trend Our points of departure are tile tools for lexical acquiwas essentially set by the need to mmximize costsition and knowledge representation (lew~loped iL~ part effectiveness in building large scale Lexical Knowledge of the ACQUII,I'3X project ('The Acquisition of Lcxieal Bases for NLP (LKBs), there is a clear sense in which Knowledge for NLP Systems'). the construction of such knowledge I)ascs also caters to The ACQUILI'~X l,exicon l)evelopment Environthe demand for better dictionaries. Currently available men| uses typed graph unilication with inheritance dictionaries and thesauri provide an undoubtedly rich as its lexical representation htnguage (for details, see source of lexical information, but often omit or neglect Copestake (1992), Sanfiliplm &amp; l'oznafiski (1992), and to make explicit salient syntactic and semantic properpal)ers by Copestake, de Paiva and Sanfilippo in ties of word entries. For exa|nplc, it is well known that Briscoe el al. (1993)). It; allows the user to define the same verb sense can appear in a wtriety of snl)catan inheritance hierarchy of types with associated reegorization frames which can be related to one at|other strictions expressed in terms of attril)ute-wdue \[)airs through valency alternations (diatheses). Some dictioas shown in Fig 1, and to create lexicons where such naries provide subcategorization i formation by means types are used to create lexical templates which encode of grammar codes, as shown below for the &quot;sail&quot; sense. word-se,|se specific information ex{.racte.d from MRI)s of the verb dock in LI)OCE -- Longman's Dictionary st, ch as the one in Fig 2. (Bold lowerc~me is used for of Contemporary English (Procter, 1978). types, caps for attributes, and boxes enclosing types indicate total omission of attribute-vahm pairs, l)etails (1) a,,,:k &quot;| , \[Tl;m: (&quot;0\] .... The codes \[T1;10:(at)\] indicate that the vcrl) can bc either transitive or intransitive with the possible a(Idition of all oblique colnpienlent introduced by l.he preposition at: (2) a. \[T1 (at)\]: Kim docked his ship (at Clasgow) b. \[IO (at)l: The ship docked (at Glasgow) Unfortunately, an indication of diatheses which relate the various occurrences of tt,e verb to one another is rarely provided. Consequently, if we were to use the grammar code information found in M)OCE to create verb entries in an I,I(B by automatic conversion we would construct four seemingly vnrelated entries for the verb dock (see §3). Inadequacies of this kind may be redressed through semiantomatie techniques *The researcl, relmrted in this paper was carried out within the ACQUILFX project. Iatn indebted to Ted Briscoe, Ann Col)estake and Pete Whitek)ck for helpful comments.</p>
            </div>
        </front>
        <body>
            <div>
                <p>COIICel'llillg, |lie OllcOdillg o\[ vet'\]) Sylll, aX an(I seulantics can be found in Sanlilil)po (1993).)</p>
                <p>Feature Structure (I&quot;S) descriptions of word senses such as tilt.' one in Fig 2 are created semiautomatically through a program which converts syntactie an(I lex-sig, verb-slgn strict-intraua-Mgn sign .-. • . . ° verl,-Mgn CAT = ~lex-etd\] S~:M = ~ \[;'igure I: q'ype IIierarchy &amp; Constraints (fragment). \] T rule . h!xllrld-l'uh.. ... . . . L 'Npu'r= ~ J J • 8trict-intrans-nign Olq.Tn ~ ttwiln &quot; RESULT~trlct'intralla'cat= ~ J 1</p>
                <p>\[ Ilp-sllgn \] ACTIVE = \[ SEM = \[\] str|ct-intrana-~am IND = 1I\] PRED = and</p>
                <p>\[ / INnverb'f°rlnula= \[\] \] AROl = \[ PRED = \[~l~whnl~l_l</p>
                <p>tAROt = \[\]</p>
                <p>agt-formula</p>
                <p>IND = \[~proce~a ARG2 : \[\] PRED = llgt-cettt~-lnov-s||antt</p>
                <p>ARG1 = \[\]</p>
                <p>ARO2 e-anhnate CAT = SEM = Figure 2: LKB Entry for swim (simplified). semantic specifications encoded in MRDs into LKB types. For example, the choice of LKB types used in the characterization of the verb swim above was induced from the syntactic and semantic codes found in LDOCE and tim Longman Lexicon of Contemporary English (LLOCE, McArtllur 1980). In LI)OCE, the first sense of the verb swim is marked ,'us a strict intransitive verb (\[I0\]) whose subject is animate ((box .... 0)); in LLOCE, the same verb sense is semantically classified as a movement verb with manner of motion specified (M19): (3) swim 1 (1)</p>
                <p>LDOCE \[I0\] (box .... 0)...</p>
                <p>LLOCE M19 - Particular ways of moving The MRD-to-LKB equivalences induced by the conversion algorithm are as shown in (4) where agt-eausemove-manner indicates that the subject participant relation implies self-induced movement with manner specified. (4) \[10\] (box .... 0) M19 \[I0\], M19 3 Verbal Diatheses</p>
                <p>Acquisition In the example discussed above, MI~D-to-LKB conversion is relatively straightforward: asingle LKB entry is created for swim since a single grammar code is found in the MRD sources used. Where a verb-sense entry gives more than one grammar code, however, the question arises whether or not each grammar code should be mapped into a distinct LKB entry. For example, the codes given in LDOCE for the verb dock (see (1)) could potentially be used to derive four LKB verb entries: --4&quot; ~ -* --~ str|ct-|ntranu-slgn teaT: ACTIVE:SZM: ^aa2= ~-an\[lllate\] \[cxr:ACTWr.:SEU:Vm)= \[tgt-eHtlSe-lnOVe-l||H II II(~l&quot; l tS~M:tun = l&quot; ...... l and Lexieal (s) LKB TYPF, striet-trans-sign obl-trans-slgn</p>
                <p>EXAMP LI,;</p>
                <p>Kim docked the boat</p>
                <p>Kim docked the boat</p>
                <p>at Southampton striet-intrans-sign The boat docked obl-intrmxs-sign The boat docked at</p>
                <p>Southampton a. b. c. d. Notice, however, that in tllis case the creation of four distinct LKB entries is unnecessary insofar ,as the use of the verb exemplified in (5b) contains enongh information to derive the remaining uses of the verb through lexical rules which progressively reduce the verb's valency by dropping the subject and/or prepositional argument(s). Such a step would be linguistically motivated in that it establishes a clear link between alternative uses of the same verb sense. Moreover, compact representation of verb use extensions is desirable from an engineering perspective its it reduces the size of ttm lexicon, allowing verb use expansion to be delayed till parsing time. This practice can be made to facilitate the resolution of lexical ambiguity by enforcing selective application of lexical rules (Copestake &amp; Briscoe, 1994).</p>
                <p>Compact representation of verb use extensions due to valency alternations requires that a note of all applicable lexical rules be made in each kernel entry. In choosing ol)l-trans-slgn as the LKB type for dock, for example, specifications would be added saying that the verb is amenabh~&quot; to the causative-inchoative alternation relating agentive and agentless uses ((5a,b) vs. (5c,d)), and the path alternation pertaining to the omission of the prepositional argument ((5a,c) vs. (5b,d)). In addition, the path alternation would ha~(e to be specified as to whether it preserves amenability to a telic interpretation (accomplishment or achievemen|) of the event described by the verb or not. For example, tile omission of the goal argnment for a verb such as drive, push or carmj induces an atelic (process) interpretation as indicated by incompatibility with a terminative adverbial: (6) a. ,\]ohn drove his car to London in one hour</p>
                <p>b. John drove his ear (*in one hour) Within a (partial) deeomposltional approach to verb semantics (Tahny, 1985; Jackendolr, 1990; Sanfilippo, 1993; Sanlilil)po el al., 1992)), this contrast can be explained with reference to the rneaniug component path. In (6a), the goal argument (1o London) fixes a final bound for the path along which the driving event takes place. Assuming that, the compositional meaning of tile sentence involves establishing a homomorphism between tile event described by the verb and the path along which such an event takes place (l)owty, 1991; Sanlilippo, 1991), it follows that with an unbounded path (e.g. (6b)) only a process interpretation is possible, whereas with a bounded path (e.g. (6a)) a relic interpretation is more likely. P,y contr~t, the omission of the goal argument with verbs such ,an deliver, bring, dock and send does not inhibit amenability to a relic interpretation, e.g. (7) We can deliver the goods (to your door) in one hour</p>
                <p>Our aim, then, wtLs to capture regularities across distinct nses of the same verb sense by relating the subcategorization frames relative to these uses via regular syntactic and semantic changes. 'lb iLssess the feasibility of this approach, we attgmented the MtH)-to-LKB conversion code with facilities which make it possible to infer amenability to specific diathesis alternations from occurrence of multiple grammar codes and their ~ssociated semantic codes in the MR, Ds. To improve on the informational content of LDOCE grammar codes, we used an intermediate dictionary semiantonmtically derived from LDOCE (LI)OCEAnter) where the subcategorization information inferrablc from grammar codes and other orthographic conventions wi~s made more explicit (Boguraev &amp; Briscoe, 1989; Carroll &amp; Grovel 1989). Semantic inlbrmation about verb classes was obtained by mapping across LI)OCF, and II, LOCE so as to augment LI)OCE queries with thesaurus information, i.e. semantic codes (Santilippo &amp; Poznafiski, 1992).</p>
                <p>Syntactic and semantic intbrmation relative to verb senses was extracted through special functio,~s which operate on pointers to dictionary entries. The extracted info was used to generate FS representations of word senses. The collversion process was carried out in such a way that whenever multiple subcategorization frames were found in association with a verb sense, only those which could not be derived via diathesis alternation were expanded into LKII entries, l;br example, the LDOCEAnter entry for dock gives four subeategorization frames: (dock)</p>
                <p>(((Cat V) (Takes NP) (Type 1))</p>
                <p>((Cat 7) (Takes NP PP) (Type 2) (PFtlRH at))</p>
                <p>(((;at 7) (Takes NP NP) (Type 2 Transitive))</p>
                <p>((Cat 7) (Takes NP NP PP) (Type 3) (PF\[\]RH at)))</p>
                <p>In this case, the four uses of the verb can all be derived from the last one through application of the causative-inchoative and bounded-path alternations mentioned above; all that ,ceds doing is to mark what diatheses are possible in the LKB entry derived, e.g. o|)i-t ....... ig,t OIUI'II = do&lt;zk \] L 'l?he algorithm which guides this process checks whether information regardiug diathesis alternations can be inferred from dictionary entries iu the MRI) sources or must be manually supplied. In performlug this check, snbcategorization options relative to a given verb sense which can be inferred from a more informative subcategorization frame are ignored. This technique was successfidly employed in semiautomatic derivation of lexicons for 360 n~ow.'n~ent verbs yieldiug over 500 additional possible expansions by application of lexical rules. = \[ TI(.ANS-ALT = catm-lnch \] L OIJIrAI/I' ~ I)-IHLth J 4 Verbal Diatheses and Knowledge</p>
                <p>Representation To encode amenability to verbal diathescs, the feature D1ATHESES wiLs introduced ;Ls an extension of the morphological features ~Lssoeiated with verbs (see (8)). This feature takes as value the type altem,alions which is in turn (teiine(l ,'ks having a wtriety of I.~X AM 1) I,E l(im broke lhe glass vs. the glass broke 1(ira scares Sally vs. Sally scares easily John ale a sandwich vs. John ale John did nol notice the sign vs. John did nol notice Kim met Bill vs. Kim and Bill met Hill read the Guardian vs. The Guardian was read by Bill Kim relurned the book to Sue vs. Kim returned the book Kim came away vs. Kim came (particle alternation) Kim swam across lhc Fiver vs. Kim swam Kim walked away vs. Kim walked (particle alternation) John broughl a book to/for Sue vs. John brought Sue a book specialized types according to which diathesis alternations are admissible for each choice of verb type (e.g. intransitive, transit.ive, ditransitive), ~u shown in Figure 3 (see next page). The following table provides examples of the diatheses refi~rred to in l&quot;ig 3. cal rules whM,, on par with all other information structures in the LKP,, are hierarchically arranged, ~s shown in Fig 4 with reference t,o the bound and unl)ound path alternations for intransitive verbs. LexicM rules t~hl-ittt rann-.lt h!xleal-I'tde tl-l)ltt h-all I~-path-alt U_l)ath_ol)l_intrnns_al t b-l)ath-ol)l-|s~tratt a-all Figure 4: Lexical Rule llierarchy (fragment). enforcing diathesis alternations may involw~ a variety of syntactic, semantic and orthographic changes. For example, the u-path-ohl-intrans-alt rule shown in Fig 5 I)elow takes as input an I&quot;S of type obl-intranssign which represents a verb describing a non-stative eventuality (dyn-eve) whose subject participant (with semantics KI) is implied as moving along a directed path (th-move-dir) the endpoint of which is specitled by the oblique a,'gument (pp-sign), e.g. the use of swim in Kim swam acTvss the river. The output is an FS representing astrict intransitive verb (strietintrans-sign) which describes a process and whose subject participant is like that of the inlutt with the directed path speciIication removed (th-move instead of th-mow~-dir), e.g. swim in Kim swam). 5 Using the I,KB to Guide Dictionary</p>
                <p>Compilation There are at least two ways in which an LKB such as the one developed in ACQUILEX offers the means to I) IATI1ESIS caus-inch middle indef-obj defobj reel I) Itass b-lmth U-lmth to/tbre l)iathesis alternations are enforced by means of lexi\[alternations \] I)IVI'-AI,T = prt-or~obl-a|t \] PRT-ALT = prt-or-obl-alt l L TItANS-ALT PRT-ALT = prt-or-obl=alt \[ I'IUI'-AIA' = prt=or-obl-alt \] OBb-ALT = prt-or-obl-alt = trans-alt PRT-ALT = prt-or-obl-alt / OBL.ALT = prt-or-obl-alt J \[ PRT-ALT = l)rt-or-ol)l-alt l L TItANS-AI,T OBL-AIII&quot; = = l)rt-ttr-obl-alt traits-air j \] L PItT-ALT = prt-or-obl-alt TItANS. ALT : trana-alt J</p>
                <p>dltrana-dlatl .....</p>
                <p>&quot;PttANS-AL'r PPJt%ALT = prt-or-obl-alt OBL-ALT = = prt-or=obl-alt trana-alt L I)AT-MOVT = (|itt=nlovt \[t ........ bl-diatl ..... \] | PRT-ALT = prt-or-obl-alt l TRANS-AIIr = italia-air L ODL-AIIY = prt-or-obl-alt trana-alt _E caua-incl h middle) htdef-obj~ def-obj, recip, pass prt-or-obl-alt ~ b-path) u-path dat-niovt ~ to) for &quot; u-pat h-obl-lnt OUTPUT = &quot; ol)l-lnt rana-algn ORTII ~ \[\] CAT = INPUT = SEM =</p>
            </div>
            <note n="1-" place="below">lexical-rule \]</note>
            <note n="275" place="below"></note>
            <div1>
                <head xml:id="sec276"></head>
                <p>rana-alt • mtrlet-|ntrans-slgn ORTtl = \[fflorth</p>
                <p>&quot; atrlct-|ntrans-cat</p>
                <p>I~ESULT = Figure 3: Verbal Dhttheses Ilierarchy \] CAT = CT \[ nl~-a|gll</p>
                <p>A IVE = \[SEM = ~ \]</p>
                <p>st r|ct=intrans-sem</p>
                <p>IND = process</p>
                <p>FRED = and SEM = ARQ1 = \[\] r tll-formula \]</p>
                <p>ARG2 = \[\] \[ PItED = th-ntove</p>
                <p>L A~o~ = \[\] obl-|ntrans-cat</p>
                <p>\[ strlct=|ntrana-cat</p>
                <p>L | A OT IVE -- \[ Itl)'a|gll &quot;l ._ \[~EM=mJ ACTIVE = \[ \[ pp-slgn SEM = \[\] 1 \] lntrana-obl-sem IND = dylt-ev0 PRI,?,D = aitd AltGl ~ \[~\]Iv~r|)=fornlu|a i L Ar~a~ = ~ ~ q Figure 5: Tlle &quot;unbounded p~tth&quot; lexical rule for intransitives \] &quot;l CoRPus the ship NP LKll: docked safely ~ ADV \[ 1 ,,-p.th-obl-~,~~'\] OII'rPUT = ~t-illgran~_~igl,J 1 INPUT = ~ .j CORPUS T the strip (IOCKe(I t_t(, tal~Sgow Figure 6: TBA facilitate word classificatiol, in the compilation of new lexieal databases.</p>
                <p>First, the links between LKB types and dictionary entries established in the conversion stage can be used to run consistency cheeks on tile MRI) sources and to supply missing information or correct errors, This offers an efficient and cost-effective way of generating improved versions of the same (lietionary.</p>
                <p>Second, the types associated with specific word classes can he made to guide lexical aquisition from corpora when creating new dictionaries. It is now widely recognized that corpora are indispensable in the acquisition of lexicaI information relating to issues of usage such as the range and frequency of different patterns of syntactic realization. 'FILe availability of software tools for partial analysis of texts {e.g. morphological and semantic tagging, phrasal parsing etc.) has increased significantly the utility of corpora in lexical acquisition by providing ways to structure the information contained in them (see Briscoe {199l) and references therein), l'h~rther advances yet can be made by using LKB types to chLssify words in text eof pora. Suplmse , for example, we lit,ked the input and ontl)ut of lexical rules to semantically tagged subcategorization frames extracted from bracketed corpora| (Poznafiski &amp; Sautillppo, 1993). As indicated in Fig 6, this would allow us to assess which alternations might be of interest in establishing regular verb sense/usage shifts. Such an assessment wonhl provide an effective way to drive verb categorization from corpora in the domain of valency alternations. 6 Final Remarks A key element in our approach to \]exica\[ acquisition and representation f verbal diathesis concerns tim use of semantics constraints in formulating MI{I) queries and characterizing FS descriptions. This practice ensttres that the results achieved in this work for ,n(&gt; tiou verbs can be suitably extended to other semantic verb classes. For example, the c\[tuss of verbs which undergo &quot;extraposition&quot; --e.g. That Kim left early bothers Sue vs. It bothers 5'ue that Kim left early-- can be identified by using semantic constraints on MK1) queries which identify psychological verbs with stimuIlls subject such as bother, please, etc. (Sanfilippo &amp; Poznatiski, 1992). This approach provides an effective way of employing semiautomatic extraction of information from MIt,1)s for lexicon construction, and it facilitates word classification from text corpora when co,npiling new dictionary datab~es. References P, riseoe, T. (1991) I,exical Issues in Natural I,anguage</p>
                <p>Processing. In Klein, E. ,~ F. Veltman ((:(Is.). Nat-</p>
                <p>uml Language and 5'peecl h Springer-Verlag, 39-68. llriscoe, T., A. Copestake and V. de Paiva (1993)</p>
                <p>Defaull Inheritance within Unifiealion-Based Ap-</p>
                <p>p~vaches to the Lea:icon, CUP. P, oguraev, B. &amp; T. Briseoe (1989) Utilising the</p>
                <p>I,I)OCE (Jrammar Codes. In Boguraev, 1L &amp;</p>
                <p>Brlscoc, '\['. (eds.) Computational Leedcography for</p>
                <p>Natural Langvage Processing, Longman, London. Carroll, J. &amp; C. (Irover (1989) The Derivation of</p>
                <p>a Large Computational l,exicon for English from</p>
                <p>LI)OCE. In lloguraw, B. (~ llriscoe, T. (eds.) Com-</p>
                <p>pvlalional Le:eieography for Natural Language Pro-</p>
                <p>cessing. Longman, I,ondon.</p>
                <p>Col)estake, A. (1992) The ACQUILEX LI(B: Repre-</p>
                <p>sentation Issues in Semi-Antomatic Acquisition or</p>
                <p>Large I,exlcons. In l'roceedings of the 3rd Confer-</p>
                <p>ence on Applied Natural Langvage Processing.</p>
                <p>Copestake, A. and T. Briscoe (1994) Semi-productive</p>
                <p>|)olysemy and Sense Extensions. Ms. Computer</p>
                <p>Laboratory University of Cambridge, Xerox Park</p>
                <p>(Menlo l)ark) and Rank Xerox Research l,aboratory</p>
                <p>(Qrenohle).</p>
                <p>l)owty, D. (1991) Thematic Protc~Roles aml Arga-</p>
                <p>mcnt Selection. Language (}7, pp. 547-619.</p>
                <p>Jackendoff, ll.. (1990) ,5'emanlic Strvctures. MIT</p>
                <p>Press, Cambridge, Mass.</p>
                <p>McArthur, T, (1981) Longman Lexicon of Contempo-</p>
                <p>rary g'nglish. Lougman, l,oudon.</p>
                <p>Poznafiskl &amp;, Saufilipl)o (1993) l)eteeting I)ependen-</p>
                <p>eies between Semantic Verb Subclasses and Subc;~t-</p>
                <p>cgorization Frames in q'cxt Corpora. In IL Boguracv</p>
                <p>&amp;. J. Pustcjovsky (eds) Acquisition of Lez'ical Knowl-</p>
                <p>edge from Text, Proceedings of a ,gIG I, EX workshop,</p>
                <p>ACL-93, Ohio.</p>
                <p>Procter, I ). (1978) Longman Dictionary of Contem-</p>
                <p>porary English. l,ongnmn, London.</p>
                <p>San/ilippo, A. (1991) Thcnmtie aud Aspectua\] Infor-</p>
                <p>nmtion in Verb Se,nanl.ics. Belgian Journal of Lin-</p>
                <p>ffuislics, 6.</p>
                <p>Sanfilippo, A. (1993) LKII Eneodil,g of Lexical</p>
                <p>Knowledge. In llriseoe, '1'., A. Cot)e.stake and V.</p>
                <p>de l'aiva (eds.).</p>
                <p>Santiliplm, A &amp; V. Poznafiski (1992) The Acquisi-</p>
                <p>tion of Lexical Knowledge from Combined Machine-</p>
                <p>Readable l)ictionary Sources. In t'roceedings of the</p>
                <p>3rd Conference on Applied Natural Language lbv-</p>
                <p>cessing, 'lYento.</p>
                <p>Sanfilil)po, A., T. Briscoe, A. Copestake, M. Mar|l,</p>
                <p>M. Tauld and A. Alongc (1992) ~lYanslation I&quot;quiv-</p>
                <p>alence and Lexiealization in the ACQUILI';X I,l(P,.</p>
                <p>Proceedings of TM1-92, Montreal, Canada.</p>
                <p>Tahny, 1,. (1985) Lexiealizatio,~ Patterns: Semantic</p>
                <p>Structure in Lexical Form. In Shopen, T. (cd) Lan-</p>
                <p>guage &quot;\]~pology and Syntactic Description 3. (h'am-</p>
                <p>marital Categoriesv and the Lexicon, CUlL</p>
            </div1>
            <note n="1&quot;" place="below">binary-formula \[ l PP~ED = and \[ th-fi)rmula ARG2 ~ \[ ARGI = \[\] \] PILED = th-ltlOve-d|r L Alia2 = l~lol, j J</note>
            <note n="277" place="below"></note>
        </body>
        <back/>
    </text>
</TEI>
