<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>Evaluating Complement-Modifier Distinctions in a Semantically Annotated Corpus</title><author surname="McConville" givenname="Mark"><org  name="University of Edinburgh" country="United Kingdom" city="Edinburgh"/></author><author surname="Dzikovska" givenname="Myroslava O."><org  name="University of Edinburgh" country="United Kingdom" city="Edinburgh"/></author></firstpageheader><frontmatter><p><b>Evaluating Complement-Modifier Distinctions in a Semantically Annotated Corpus</b></p><p><b>Mark McConville, Myroslava O. Dzikovska</b></p><p>School of Informatics, University of Edinburgh 2 Buccleuch Place, Edinburgh, EH8 9LW, Scotland Mark.McConville@ed.ac.uk, M.Dzikovska@ed.ac.uk</p></frontmatter><abstract>We evaluate the extent to which the distinction between semantically core and non-core dependents as used in the FrameNet corpus corresponds to the traditional distinction between syntactic complements and modifiers of a verb, for the purposes of harvesting a wide-coverage verb lexicon from FrameNet for use in deep linguistic processing applications. We use the VerbNet verb database as our gold standard for making judgements about complement-hood, in conjunction with our own intuitions in cases where VerbNet is incomplete. We conclude that there is enough agreement between the two notions (0.85) to make practical the simple expedient of equating core PP dependents in FrameNet with PP complements in our lexicon. Doing so means that we lose around 13% of PP complements, whilst around 9% of the PP dependents left in the lexicon are not complements. </abstract></header><body><section number="1" title="Introduction"><p>The distinction between complements and modifiers is im­portant for parsing and language interpretation, since it im­pacts upon both syntactic and semantic decisions. On the syntactic level, the distinction relates to questions such as whether a particular sentence is syntactically correct if a given dependent is not realised. On the semantic level, the complement-modifier distinction raises important is­sues such as whether a semantic representation is 'com­plete', with nothing needing to be inferred from context, and whether a particular preposition denotes an indepen­dent predicate, as opposed to being a contentless argument marker. The answers to these questions may significantly impact the quality of parsing and semantic interpretation. As noted by Meyers et al. (1996), incorrectly classifying a complement as a modifier may cause a syntactic parser to miss a parse, whilst incorrestly classifying a modifier as a complement may cause a parser to add a spurious parse. This would be a relevant distinction to make in a (deep) syn­tactic parser. In addition, in an accurate representation of predicate argument structure, syntactic phrase heads pred­icate of their complements, but modifiers predicate of the syntactic phrase heads. This distinction is therefore rele­vant for both deep parsers that combine syntax and seman­tics, and shallow semantic parsers that attempt to induce a predicate structure based on semantic role labelling. Unfortunately, the precise boundary between complements and modifiers is notoriously difficult to define. Existing lex­ical semantic resources, for example PropBank (Palmer et al., 2005), FrameNet (Johnson and Fillmore, 2000), Verb-Net (Kipper et al., 2000) and OntoNotes (Hovy et al., 2006) , do include information related to the complement-modifier distinction, but each applies slightly different cri­teria, depending on whether the emphasis is syntactic or semantic. A number of recent projects have attempted to merge information from different resources (Kwon and Hovy, 2006; Crabbe et al., 2006) and use them in pars­ing (Shi and Mihalcea, 2005; McConville and Dzikovska, 2007) . For such applications it is important to be able to understand and evaluate to what extent the different ap­proaches to making complement-modifier distinctions are compatible, and which approach is the most appropriate for a given application.</p><p>In this paper, we investigate whether the semantic criteria for distinguishing complements and modifiers used by the creators of the FrameNet corpus (i.e. 'core' versus 'non-core' semantic roles) correspond to syntactic intuitions, in particular to the (primarily) syntactic criteria used in the VerbNet lexicon, which only lists syntactic arguments of the verbs, based on whether they can participate in a num­ber of syntactic alternations.</p><p>We show that while there is a reasonably good correla­tion (0.85 agreement) between the semantic 'coreness' of FrameNet verb dependents and the complements listed in VerbNet, these notions do not align perfectly, and discuss the implications for using FrameNet as a source of syntactic information for parsing. We argue that for deep parsers con­cerned with both syntactic and semantic representations, it may be beneficial to separate the semantic and syntactic as­pects of complement/modifier distinction into separate fea­tures.</p><p>This paper proceeds as follows: <b>section 2 </b>presents some necessary background, including an introduction to both FrameNet and VerbNet, as well as our lexical harvesting project; <b>section 3 </b>discusses the methodology used in our investigation, in particular the way in which we combined use of VerbNet with our own linguistic intuitions in making a decision about complement-hood; <b>section 4 </b>presents the results; and <b>section 5 </b>discusses the implications for deep parsing.</p></section><section number="2" title="Background 2.1   The FrameNet corpus"><p>FrameNet (Johnson and Fillmore, 2000) is a corpus of 140,000 English sentences (mainly drawn from the BNC), each annotated with both syntactic and semantic informa­tion. Underlying the corpus is an ontology of 800 'frames' (or semantic types), each of which is associated with a set of 'frame elements' (or semantic roles).<page local="2"/> Take for example the following sentence from the corpus:</p><p>(1) Overshadowed by Grigorovich, Kokonin nonetheless apparently eclipsed him in power in recent months.</p><p>In this example, the verb <i>eclipse </i>is associated with the Surpassing frame, which denotes situations where one entity is conceptualised as being superior to another in some way. This frame includes the following frame ele­ments, among others:</p><p>• Attribute - a property that invokes a scale (e.g. 'power', 'wealth')</p><p>• Item - the entity located closest to the end of the scale (i.e. the 'surpassor')</p><p>• Standard - the entity located farthest from the end of the scale (i.e. 'the surpassed')</p><p>• Time - the time when the Item is higher on the scale</p><p>Other verbs which are listed in the FrameNet ontology as 'evoking' the Surpassing frame include <i>surpass, better, outdo </i>and <i>outshine.</i></p><p>The FrameNet annotation process then runs as follows, as­suming the example sentence in (1):</p><p>1. identify a target word for the annotation, for example the main verb <i>eclipsed</i></p><p>2. identify the semantic frame which is evoked by the target word in the sentence - in this case the relevant frame is Surpassing</p><p>3. identify the sentential constituents which realise each frame element associated with the frame, i.e. <i>Overshadowed by Grigorovich, [Kokonin] </i>Item<i>nonetheless apparently eclipsed </i>[him]Standard [in <i>power]Attribute [in recent months]Time.</i></p><p>Finally, some basic syntactic information about the target word and the constituents realising the various frame ele­ments is also added:</p><p>• the part-of-speech of the target word (i.e. V, N, A or</p><doubt alpha="80.0" length="5" tooSmall="False" monospace="0.0">PREP)</doubt><p>• the syntactic category of each constituent realising a frame element (e.g. NP, PP, VPto, Sfin)</p><p>• the syntactic role, with respect to the target word, of each constituent realising a frame element (for exam­ple Ext (subject), Obj (object) or Dep (other depen­dent))</p><p>Thus, each sentence in the corpus can be seen to be anno­tated on at least three independent 'layers', as exemplified in Figure 1. The FrameNet corpus has proved to be a use­ful linguistic resource for a number of computational lin­guistics applications, for example semantic role labelling (Gildea and Jurafsky, 2002), information extraction (Sur-deanu et al., 2003), and question answering (Kaisser and Webber, 2007).</p><subsection number="2.2" title="Harvesting a verb lexicon from FrameNet"><p>McConville and Dzikovska (2007) present a procedure for harvesting a wide-coverage verb lexicon, for use with a deep semantic parser, from the FrameNet corpus. The technique used was to read off lexical entries from anno­tated sentences, and subsequently filter out spurious sub-categorisation frames. We took each sentence which had been annotated with respect to some target verb and con­verted it into a lexical entry whose subcategorisation frame contained the various annotated syntacto-semantic depen­dents. For example, the annotated sentence from Figure 1 was converted into the verb entry in Figure 2.<footnote anchor="1"/>This simple approach to deriving a verb lexicon gave rise to a number of spurious subcategorisation frames, involv­ing non-canonical verbal constructions and alternations (e.g. passives, imperatives, middles), which we then had to filter out. After the filtering process, the harvested lexicon had been reduced in size from 30,000 distinct verb/subcategorisation frame pairs to just 9,000. These were distributed across 2,600 verb senses, giving a ratio of 3.4 subcategorisation frames per verb sense. The process of filtering out spurious subcategorisation frames added value to the original FrameNet verb lexicon, making it more suit­able for hooking up to a deep parser where non-canonical constructions are generally handled in the rule component (e.g. using lexical rules).</p><p>One of the main issues we encountered in filtering the FrameNet verb lexicon involved distinguishing between those Dep dependents which are true complements of the verb and those which are generic modifiers, and filtering out the latter. To return to the example sentence in (1), the prepositional phrase <i>in power </i>is a complement of the target verb <i>eclipse, </i>since:</p><p>• although it is optional, it defines an argument specific to this class of verb, hence its existence cannot be pre­dicted from more general principles of grammar</p><p>• the preposition <i>in </i>is the only preposition which can be used to introduce the Attribute role of the verb</p><doubt alpha="100.0" length="7" tooSmall="False" monospace="0.0">eclipse</doubt><p>On the other hand, the prepositional phrase <i>in recent months </i>is a modifier of the target verb, since:</p><p>• it can be used with much the same meaning with al­most all classes of verb, hence its existence (and op-tionality) can be predicted from more general princi­ples of grammar</p><p>• the preposition <i>in </i>is <i>not </i>the only preposition which can be used to introduce the Time role — <i>Kokonin eclipsed Grigorovich {at the weekend, over three years, after a few days, on Saturday,... }</i></p><p>Thus, these two prepositional phrases need to be treated dif­ferently in parsing and construction of logical forms, as dis­cussed in the introduction. In particular, we need to make sure that lexical subcategorisation frames in our harvested<page local="3"/></p><footnote label="1">Note that the most recent version of the FrameNet corpus in­cludes these automatically-generated subcategorisation frames in lexical unit files.</footnote><p>Surpassing</p><doubt alpha="100.0" length="4" tooSmall="False" monospace="0.0">Item</doubt><p>Standard</p><p>Attribute</p><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">V</doubt><doubt alpha="100.0" length="4" tooSmall="False" monospace="0.0">ORTH</doubt><doubt alpha="100.0" length="6" tooSmall="False" monospace="0.0">SYNCAT</doubt><doubt alpha="100.0" length="7" tooSmall="False" monospace="0.0">SEMTYPE</doubt><doubt alpha="100.0" length="4" tooSmall="False" monospace="0.0">ARGS</doubt><p><i>{eclipse)</i></p><p>SYNROLE SYNCAT SEMROLE</p><doubt alpha="83.3" length="6" tooSmall="False" monospace="0.0">Ext NP</doubt><doubt alpha="83.3" length="6" tooSmall="False" monospace="0.0">Obj NP</doubt><doubt alpha="83.3" length="6" tooSmall="False" monospace="0.0">Dep PP</doubt><p>SYNROLE Dep SYNCAT PP SEMROLE Time</p><figure caption="Figure 2: A harvested verb entry"></figure><p>verb lexicon contain only those dependents which are sub­jects or complements, with modifiers filtered out. In order to do this simply and straightforwardly, we availed ourselves of one of the features built in to the FrameNet ontology — the 'coreness' feature on frame elements. In the ontology, the frame elements associated with each frame have been partitioned into two main groups, Core and non-Core, a distinction which (according to the anno­tation guidelines) is meant to cover the 'semantic spirit' of the distinction between complements and modifiers. Thus, for example, obligatory complements are always Core, as are:</p><p>• those which, when omitted, receive a definite interpre­tation, e.g. the Goal argument of the verb <i>arrive </i>(cf. <i>John arrived)</i></p><p>• those whose semantics cannot be predicted from their form, e.g. the Intermediary argument of the verb <i>relies </i>as in <i>John relies [on Mary] </i>— the preposition <i>on </i>does not encode the Intermediary role with any other class of verbs outwith the Reliance frame</p><p>Non-Core frame elements are themselves partitioned into two classes:</p><p><b>Peripheral </b>frame elements which do not introduce ad­ditional or distinct events from the main reported event, i.e. Time, Manner, Place, Degree, etc.</p><p><b>Extra-thematic </b>situate an event against a back­drop of another state-of-affairs, i.e. Frequency,</p><p>Containing_event, Beneficiary, etc.</p><p>To go back to the example sentence in Figure 1 in­volving the verb <i>eclipse </i>from the Surpassing frame, the FrameNet ontology classes the Item, Standard and Attribute roles as Core, and the Time role as Peripheral. This appears to correspond exactly with linguistic intuitions about which dependents are comple­ments and which are modifiers of the verb. For this reason, when it came to filtering out spurious subcategorisa­tion frames from the verb lexicon we had harvested from the FrameNet corpus, we decided on the simple expedient of deleting all and only those Dep dependents which evoke a non-Core frame element. This process resulted in the elimination of many spurious subcategorisation frames — the number of verb/subcategorisation frame pairs in the har­vested lexicon was cut by 45%, from 16,000 down to 9,000. However, it was clear from the start that the correlation be­tween the semantic 'coreness' and syntactic complement-hood is far from perfect. For instance, it was not difficult to find examples where syntactic complements evoked se­mantically non-Core frame elements. A number of con­stituents in the FrameNet corpus have been marked as di­rect objects, despite invoking non-Core frame elements, as in:</p><p>(2)   [John]Agent <i>ripped [the </i>top^ubregion <i>[from his packet of cigarettes</i><i>]Patient</i></p><p>In this instance, the verb <i>rip </i>has been assigned by the an-notators to the frame Damaging, where the Subregion frame element is marked as being Peripheral, based on examples like <i>John ripped his trousers [below the knee]. </i>In this particular case, the problem was probably caused by annotators not being careful enough when assigning verbs with different subcategorisation alternations to frames — it would have been better to have assigned the verb <i>rip </i>as used in (2) to the Removing frame, where the direct ob­ject invokes a Core frame element (i.e. Theme). Thus, the decision to retain all senses of the verb <i>rip </i>within the same frame has led to a situation where semantic and syntactic coreness have become dislocated.</p><p>The aim of the project reported here was to investigate the extent to which this kind of problem impacts upon the ef­fectiveness of the expedient we chose to distinguish com­plements from modifiers in the verb lexicon we harvested from the FrameNet corpus. In other words, we wanted to ascertain to what extent the 'coreness' feature on frame el­ements in the FrameNet ontology corresponds with linguistic intuitions as to which dependents are complements and which are modifiers.<page local="4"/></p><table caption="Figure 1: A FrameNet annotated sentence" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p><i>Kokonin</i></p></td><td class="cell"><p><i>apparently</i></p></td><td class="cell"><p><i>eclipsed</i></p></td><td class="cell"><p><i>him</i></p></td><td class="cell"><p><i>in power</i></p></td><td class="cell"><p><i>in recent months</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>target</p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>Surpassing</p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>frame element</p></td><td class="cell"><p>Item</p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>Standard</p></td><td class="cell"><p>Attribute</p></td><td class="cell"><p>Time</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>syntactic category</p></td><td class="cell"><p>NP</p></td><td class="cell"><p></p></td><td class="cell"><p>V</p></td><td class="cell"><p>NP</p></td><td class="cell"><p>PP</p></td><td class="cell"><p>PP</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>syntactic role</p></td><td class="cell"><p>Ext</p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>Obj</p></td><td class="cell"><p>Dep</p></td><td class="cell"><p>Dep</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></subsection><subsection number="2.3" title="The VerbNet verb lexicon"><p>Unfortunately, judgments on which verb dependents are complements and which are adjuncts are notoriously dif­ficult to make consistently. Although there are many cases which are clearly complements and many which are clearly modifiers, there are a large number of borderline cases where it is hard to make any kind of definite decision ei­ther way. As noted by the creators of the Penn TreeBank (Marcus et al., 1994): "After many attempts to find a re­liable test to distinguish between arguments and adjuncts, we have abandoned structurally marking this difference". The situation is muddied still further by the fact that certain of the suggested criteria found in the literature are plagued by issues of gradient grammaticality. For example, Mey­ers et al. (1996) states that pseudo-passivisation (e.g. <i>John is relied on by many people) </i>is a property of complement PPs but not modifier ones. However, Tseng (2006) points out that there is a continuum of acceptability for pseudo­passives, and many PPs which are clearly modifiers can be be pseudo-passivised, e.g. <i>David always takes that seat in the corner because he hates being sat next to. </i>In order to help us in deciding whether a given verb depen­dent is really a complement, we decided to make use of the VerbNet verb lexicon (Kipper et al., 2000) as a syntactic gold standard.</p><p>VerbNet is a lexicon of around 5,000 English verb senses, partitioned into 237 top-level classes. Each verb class spec­ifies, among other things, a set of associated subcategorisa­tion frames listing the arguments (i.e. subjects and com­plements) that are appropriate for all the verbs in the class. Take for example the VerbNet class exceed-90, which is the nearest equivalent to the Surpassing frame in FrameNet. This class encompasses verbs like <i>surpass, top </i>and <i>outstrip, </i>and specifies the following two subcategorisa-tion frames:<footnote anchor="2"/></p><doubt alpha="65.2" length="23" tooSmall="False" monospace="0.0">• NP:Theme1 V NP:Theme2</doubt><p>• NP:Theme1 V NP:Theme2 P NP:Attrib</p><p>In order to use VerbNet as our gold standard for making distinctions between complements and modifiers, we as­sumed that verb dependents which are listed in the relevant class are definitely complements. However, we were un­able to assume straightforwardly that any verb dependents which are <i>not </i>so listed are modifiers, since VerbNet is an incomplete resource, not being strictly corpus-based. We thus report two sets of results: those where we assume that VerbNet is a literal gold standard (i.e. pretend that it is com­plete), and those where we allow ourselves to make use of other criteria in deciding whether a dependent which is not listed in VerbNet is a complement or modifier. In addition to VerbNet, we considered ComLex (Grishman et al., 1994) and PropBank as possible sources of syntactic information. We decided that ComLex unsuitable because it is a purely syntactic resource and does not specify seman­tic roles for the complements listed in its subcategorisation frames, thus making it impossible to decide in many cases whether a given FrameNet dependent is listed ot not. In addition, the PropBank proved to be of little use, since, as pointed out in Palmer et al. (2005), 'We make no attempt to adhere to any linguistic distinction between arguments and adjuncts'.</p><footnote label="2">VerbNet subcategorisation frames are best thought of as flat­tened representations of LTAG elementary trees.</footnote></subsection></section><section number="3" title="Methodology"><p>We took the lexicon we had harvested from the FrameNet corpus<footnote anchor="3"/>, and extracted all and only those entries (incorpo­rating an orthographic base form, a semantic type, and a subcategorisation frame) which specify at least one PP de­pendent. Along with the information about annotated de­pendents, each entry was also associated with the corpus sentence which it had been harvested from. A total of 17,000 verb entries were extracted in this way. The next step was to select a random sample of these for manual checking. Unfortunately, this proved to be some­what more problematic than just picking a random sub­set. The FrameNet project's approach to annotation has proceeded on a 'frame-by-frame' basis rather than focus­ing on fully annotating running text, meaning that each frame in turn is fully annotated with respect to its lexical units, before moving on to the next identified frame. As a result, some frames contain many more lexical units than others, and thus are associated with many more annotated instances. Thus, of the 261 FrameNet frames implicated in our set of extracted verb entries, the most common 10% ac­count for 65% of the total number of entries. In particular, just the one frame (Self .motion) accounts for a quarter of all the extracted entries.</p><p>We attempted to counteract this bias by limiting each frame to a maximum of <i>two </i>entries, and moreover restricting each verb in a frame to a maximum of <i>one </i>entry. We were thus left with 593 verb entries in our sample, involving a total of 432 subcategorised Core PP dependents and 204 non-Core ones.</p><p>The annotation task involved going through each PP depen­dent in turn and deciding whether or not VerbNet classifies it as a complement of the target verb. The relevant decision tree runs as follows:</p><p>1. If the verb is listed in VerbNet with the appropriate sense:</p><p>• if the VerbNet class lists the relevant subcategorisation frame, including the PP dependent in question, then mark the dependent as a 'complement'</p><p>• if the VerbNet entry <i>does not </i>list the relevant subcate-gorisation frame, including the PP dependent in ques­tion:</p><p>- if you think that a more complete VerbNet entry for the relevant verb would list the relevant sub-categorisation frame, including the PP dependent in question, then mark the dependent as a 'com­plement'<page local="5"/></p><footnote label="3">More precisely, the lexicon after having removed all spuri­ous frames involving non-canonical constructions like passives and imperatives, but retaining all dependents, whether Core or non -Core.</footnote><p>- if not, then mark the dependent as a 'non-complement' 2. If the verb is <i>not </i>listed in VerbNet with the appropriate sense:</p><p>• if you think that a more complete VerbNet which <i>did </i>contain an entry for the relevant verb would list the appropriate subcategorisation frame, including the PP dependent in question, then mark the dependent as a 'complement'</p><p>• if not, then mark the dependent as a 'non-complement'</p><p>Thus, we first of all determined whether the relevant sense of the target verb was included in VerbNet. If so, and if a matching subcategorisation frame including the dependent as a complement was listed, then it was deemed to be a complement PP.</p><p>In cases where either the relevant sense does not exist in VerbNet, or where the relevant verb sense does exist but a matching subcategorisation frame is not listed, things be­come a little more complicated, due to the incomplete na­ture of VerbNet as a resource.</p><p>Take for example, the example sentence in Figure 2. Al­though the verb <i>eclipse </i>does not appear in VerbNet, the syn­onymous verb <i>surpass </i>does (cf. <i>John surpassed/eclipsed Mary in raw talent). </i>By close inspection of the subcat-egorisation frames listed in the VerbNet class in which <i>surpass </i>appears (i.e. exceed-90), we judge that the Attribute dependent of the verb <i>eclipse </i>in Figure 2 is a VerbNet complement — the class contains the subcategorisation frame NP:Theme1 V NP:Theme2 P NP:Attrib, where Theme1, Theme2 and Attrib match Item, Standard andAttribute respectively. Incases wheretherelevantverbsense<i>does</i>appearinVerb-Net, but there is no matching subcategorisation frame in­cluding thePPinquestionweappliedourownlinguisticin-tuitions (using standardly assumed criteria for distinguish­ing complements from modifiers) to judge whether a more complete version of VerbNet <i>should </i>list the PP dependent as a complement.</p></section><section number="4" title="Results"><p>One quarter of the verb entries in our sample were anno­tated by both authors, in order to test for inter-annotator agreement. The results were as follows:</p><p>• Is the appropriate sense of the target verb listed in VerbNet? Agreement: 0.95, kappa: 0.90</p><p>• Assuming both annotators agree that the target verb appears in VerbNet, is the PP dependent listed as a complement? Agreement: 0.97, kappa: 0.93</p><p>• Assuming that both annotators agree that the target verb appears in VerbNet but the PP dependent is not listed, is it a complement of the target verb? Agree­ment: 0.80, kappa: 0.60</p><p><i>• </i>Assuming that both annotators agree that the target verb is <i>not </i>listed in VerbNet, is the PP dependent a complement of the verb? Agreement: 0.94, kappa:</p><doubt alpha="0.0" length="4" tooSmall="False" monospace="0.0">0.87</doubt><p>The results of our investigation into the relation between coreness and compement-hood are presented in Table 1. The three 'experiments' listed are as follows:</p><p><i>• </i><b>Experiment 1 </b>only takes into account the 433 depen­dents whose verb senses were adjudged to be listed in VerbNet, ignoring annotator judgements about whether an unlisted PP dependent is a complement or not (i.e. it assumes that VerbNet is a complete resource for the verbs it lists)</p><p><i>• </i><b>Experiment 2 </b>also assumes the subset of 433 depen­dents whose verb senses were adjudged to be listed in VerbNet, but includes annotator judgements for un­listed PP dependents</p><p><i>• </i><b>Experiment 3 </b>includes results for all 634 PP de­pendents in the sample, including those whose target verbs do not appear in VerbNet</p><p>For each experiment, we present both the total assignments of dependents to classes (the columns represent FrameNet's Core versus non-Core distinction and the rows represent the judgements about syntactic complement-hood made by the annotators, in conjunction with VerbNet) as well as the interannotator agreement and Cohen's kappa scores. Note that agreement increases significantly when we take into account annotator judgement in cases where Verb­Net fails to list the relevant PP dependent. Note also that the amount of chance agreement depends on the relative proportion of Core and non-Core dependents annotated in the FrameNet corpus. It appears that the former have been annotated more completely, since over the corpus as a whole there are twice as many Core PPs listed than non-Core ones — we were able to find many PP modifiers which had been ignored by annotators. In conclusion, it appears that around 13% of syntactic PP complements will be lost if we simply delete all non-Core dependents from our harvested lexicon, and 9% of the de­pendents retained will <i>not </i>be syntactic complements. Finally, we manually examined all instances where a FrameNet Core PP dependent was judged <i>not </i>to be a com­plement of the relevant target verb. A significant propor­tion of these (around one third) appeared to involve some kind of bracketing mismatch between syntax and seman­tics, where the Core dependent annotated in FrameNet is <i>not </i>a syntactic dependent of the target verb. Take the fol­lowing example:</p><p>(3) She looked away quickly, and unfastened [the waistband]FASTENER [of her uniform skirt]CONTAINING_OBJECT.</p><p>Here the target verb is <i>unfasten. </i>The FrameNet annotation recognises both <i>the waistband </i>and <i>ofher uniform skirt </i>as distinct dependents of the verb. However, on the syntactic level it is the entire noun phrase <i>the waistband ofher uni­form skirt </i>which is a syntactic dependent of the verb (the direct object).</p><page local="6"/></section><section number="5" title="Discussion"><p>As can be seen from our analysis, while the distinction be­tween Core and non-Core semantic roles in the FrameNet ontology is highly correlated with the kind of syntactic cri­teria for complement-hood used in verb lexicons like Verb-Net, the two do not align perfectly. In other words, there are a significant number of cases where either a non-Core semantic dependent is realised by a verbal complement or a Core semantic dependent is realised by a verbal modi­fier. Although many of these mismatches can be put down to either annotation errors (generally failing to assign a par­ticular use of some verb to the optimal frame) or question­able annotation policy (e.g. annotating as semantic depen­dents phrases which are definitely not syntactic dependents of the target verb), there remain a large number of such mis­matches which are simply a result of the incompatibility of semantic and syntactic notions of 'coreness'. The fact that the complement-modifier distinction is so dif­ficult to pin down has important implications for parsing and semantic interpretation. As pointed out in the introduc­tion, a parser needs to have access to a complete, accurate list of what complements go with which verbs, if correct parses are not to be missed or spurious parses to be added. Consider the following contrasting examples:</p><doubt alpha="59.4" length="32" tooSmall="False" monospace="0.0">(4)   (a) John relied on the map</doubt><p>(b) John fell on the table (c) John slept on the table</p><p>The semantic representations corresponding to possible in­terpretations of those utterances are shown in Table 2. In (4a), 'rely' <i>requires </i>an on-PP complement (denoting the thing relied upon). In (4b), the on-PP is optional syntacti­cally, and can be either a complement or an adjunct seman-tically, depending on whether it denotes the trajectory of the fall, or the location where John fell. In (4c), the <i>on</i>-PP is syntactically and semantically optional, and it is definitely not a complement, since the location of the event is in no way unique to the <i>sleep </i>predicate. Unless the parser has access to this kind of information, then it will not be able to judge, for example, that the second sentence has two dis­tinct interpretations, whereas the other two only have one interpretation each.</p><p>In addition, note that there is an important contrast between the <i>on</i>-PP complements of <i>rely </i>and <i>fall. </i>In the former, the preposition <i>on </i>does not contribute any meaning to the sentence, other than to make clear what role its NP com­plement plays in the situation (i.e. it is basically a case-marker, cf. <i>John trusted the map). </i>The intrinsic meaning</p><p>Table 2: Possible semantic representations using FrameNet roles for utterances in example (4), and their relation to the complement/modifier distinction of the preposition <i>on, </i>involving some object being in con­tact with the top horizontal surface of some other object, plays no role in the predicate argument structure of this sentence, which can best be represented as something like <i>relyjjn(John,map).</i><i> </i>Furthermore, <i>on </i>is the only preposition which can be used here to encode the relevant semantic role (apart from its stylistic variant <i>upon). </i>In the second sentence, on the other hand, the preposition <i>does </i>contribute its intrinsic meaning to the predicate argu­ment structure — as the place where the trajectory of the fall ends. Thus, the predicate argument structure of this sentence is more like <i>fall(John,to(on(table))), </i>with the in­ference that, at the culmination of the event of falling, John is in fact 'on the table'. Also, a number of other preposi­tions can be used here to realise the relevant semantic role, e.g. <i>off, under, over, through, </i>etc.</p><p>The contrast between argument-marker and predicative uses of prepositions is important both for applications con­necting language to reasoning, and for applications using shallow semantic representations. For example, Kaisser and Webber (2007) describe the use of FrameNet in ques­tion answering where questions are paraphrased using verbs in the same frame. If we consider the paraphrases that apply in our cases, an appropriate paraphrase for (4a) would be <i>John trusted the map, </i>leaving out the preposi­tion <i>on, </i>while an appropriate paraphrase for (4b) would be <i>John dropped on the table, </i>where the preposition is re­tained, since it is crucial for the meaning. With this in mind, we decided to investigate how dif­ferent parsers handle the distinction between predicative and non-predicative uses of prepositions.<page local="7"/> We considered: (a) shallow semantic parsers that output FrameNet frames (Gildea and Jurafsky, 2002); (b) the LinGO English Re­source Grammar (Copestake and Flickinger, 2000), a deep HPSG grammar that produces semantic representations us­ing first-order logic,<footnote anchor="4"/>; and (c) the TRIPS parser Allen et al. (2007), a unification based parser for dialogue sys­tems. These systems represent several different ways of using logical forms: semantic parsers are used for question answering (Kaisser and Webber, 2007) and information re­trieval (Surdeanu et al., 2003); the LinGO ERG has been used for translating spoken dialogue (Kay et al., 1994); and the TRIPS parser produces semantic representations that are easy to map into representations used by domain spe­cific reasoners common in dialogue systems. It turns out that these three parsers have all taken different approaches to the distinction between predicative and non-predicative prepositions. Shallow semantic parsers output frame representations that identify frame element names, but not their specific meanings. So for (4a), they will iden­tify <i>on the map </i>as a Intermediary, and in 4(b) treat <i>on the table </i>as Goal, leaving open the question whether it is <i>the table </i>alone, or the whole PP, that fills the appropriate slot. The LinGO ERG parser represents some prepositions as argument markers, but not always consistently, e.g. <i>on </i>in (4a) will be represented as a case marker, but <i>for </i>in <i>John left for Boston </i>will be represented as a regular preposition, as will <i>on </i>in (4b). In contrast, the TRIPS dialogue parser makes strictly semantic decisions, and marks all PPs where the preposition is not clearly an argument marker as com­plements (thus ensuring that the the preposition predicates are included in the logical form), but at the expense of be­ing unable to rule out certain syntactically anomalous utter­ances, such as <i>*John put it.</i></p><table caption="Table l: Results" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Experiment l</p></td><td class="cell"><p>Experiment 2</p></td><td class="cell"><p>Experiment 3</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Core</p></td><td class="cell"><p>non-Core</p></td><td class="cell"><p>Core</p></td><td class="cell"><p>non-Core</p></td><td class="cell"><p>Core</p></td><td class="cell"><p>non-Core</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>complements</p></td><td class="cell"><p>l99</p></td><td class="cell"><p>37</p></td><td class="cell"><p>258</p></td><td class="cell"><p>49</p></td><td class="cell"><p>395</p></td><td class="cell"><p>59</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>non-complements</p></td><td class="cell"><p>82</p></td><td class="cell"><p>ll5</p></td><td class="cell"><p>23</p></td><td class="cell"><p>lG3</p></td><td class="cell"><p>37</p></td><td class="cell"><p>l45</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>agreement</p></td><td class="cell"><p>G.73</p></td><td class="cell"><p>G.83</p></td><td class="cell"><p>G.85</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>kappa</p></td><td class="cell"><p>G.65</p></td><td class="cell"><p>G.75</p></td><td class="cell"><p>G.65</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>John</b></p></td><td class="cell"><p><b>relied</b></p></td><td class="cell"><p><b>on the map</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>complement</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Protagonist</p></td><td class="cell"><p></p></td><td class="cell"><p>Intermediary</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>John</b></p></td><td class="cell"><p><b>fell</b></p></td><td class="cell"><p><b>on the table</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>complement</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Theme</p></td><td class="cell"><p></p></td><td class="cell"><p>Goal</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>modifier</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Theme</p></td><td class="cell"><p></p></td><td class="cell"><p>Place</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>John</b></p></td><td class="cell"><p><b>slept</b></p></td><td class="cell"><p><b>on the table</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p>modifier</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Sleeper</p></td><td class="cell"><p></p></td><td class="cell"><p>Place</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>These differences in approach are not surprising, given that the various syntactic and semantic criteria for identifying complements may not align, as we have shown in this pa­per, and therefore creating a lexicon that accurately iden­tifies such distinctions is a difficult task. We are currently working on methods to better represent this distinction in our lexicon.</p><p>We propose that for parsing lexicons, especially for deep parsers, a possible solution is to replace a single distinction with several finer-grained features, addressing the key is­sues raised in the introduction: Is some dependent syntacti­cally required to complete the utterance? Is some (optional) prepositional or adjectival phrase a possible dependent of a given verb? And does a particular preposition correspond to an independent predicate (regardless of whether the de­pendent can be classified as a complement)? The method adopted by VerbNet of defining syntactic frames based on alternations is a good approach to defining syntactic complements. It also takes the first step towards identifying those prepositions that may contribute meaning towards the predicate vs. those which do not, by defining classes of equivalent prepositions (such as location and direction) which may appear in the same PP. However, since VerbNet is incomplete, we are currently de­veloping a corpus-based approach to answer these ques­tions. We are particularly interested in the answer to the third question, whether a preposition in a given PP depen­dent is an argument marker, or contributes meaning to the logical form. This information is, to our knowledge, not coded in any existing resources. Adding it would provide essential information for building semantic representations, and therefore make such representations more usable in in­terpretation tasks.</p><footnote label="4">To be more precise, the LinGO ERG produces MRS repre­sentations that encode scope ambiguities. However, MRS repre­sentations resolve to fully instantiated logic formulas, which are our concern here</footnote></section><section number="6" title="Conclusion"><p>In this paper, we have attempted to evaluate the extent to which the distinction between semantically Core and non-Core dependents, as used in the FrameNet corpus, corre­sponds to the traditional distinction between syntactic com­plements and modifiers of a verb. We used the VerbNet verb database as our gold standard for making judgements about complement-hood, in conjunction with our own intuitions in cases where we considered VerbNet to be incomplete. We concluded that there is enough agreement between the two notions (0.85) to make practical the simple expedient of equating core PP dependents in FrameNet with PP com­plements in the wide-coverage verb lexicon we harvested from FrameNet. Doing so means that we lose around 13% of PP complements, whilst around 9% of the PP dependents left in the lexicon are not complements. We then discussed the implications of this result for deep parsing, suggesting that for parsers concerned with both syntactic and semantic representations, it may be beneficial to separate the seman­tic and syntactic aspects of complement/modifier distinc­tion into separate features</p></section><section title="Acknowledgements"><p>The work reported here was supported by grants N000140510043 and N000140510048 from the Office of Naval Research.</p></section><references><p>James Allen, Myroslava Dzikovska, Mehdi Manshadi, and Mary Swift. 2007. Deep linguistic processing for spo­ken dialogue systems. In <i>Proceedings of the ACL'07 Workshop on Deep Linguistic processing, </i>pages 49-56.</p><p>Ann Copestake and Dan Flickinger. 2000. An open source grammar development environment and broad-coverage English grammar using HPSG. In <i>Proceedings </i><i>ofthe</i><i> 2nd International Conference on Language Resources and Evaluation, </i>Athens, Greece.</p><p>Benoit Crabbe, Myroslava O. Dzikovska, William de Beau­mont, and Mary D. Swift. 2006. Increasing coverage of a domain independentdialogue lexicon with VerbNet. In <i>Proceedings of the Thurd International Workshop on Scalable Natural Language Understanding (ScaNaLU 2006), </i>New York City.</p><p>Daniel Gildea and Daniel Jurafsky. 2002. Automatic la­beling of semantic roles.  <i>Computational Linguistics,</i> 28(3):245-288.</p><page local="8"/><p>Ralph Grishman, Catherine Macleod, and Adam Meyers. 1994. COMLEX syntax: Building a computational lexi­con. In <i>Proceedings of COLING'94.</i></p><p>Eduard Hovy, Mitchell Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel. 2006. Ontonotes: The 90% solution. In <i>Proceedings of the Human Lan­guage Technology Conference of the NAACL, Compan­ion Volume: Short Papers, </i>pages 57-60, New York City, USA, June. Association for Computational Linguistics.</p><p>Christopher Johnson and Charles J Fillmore. 2000. The FrameNet tagset for frame-semantic and syntactic cod­ing of predicate-argument structure. In <i>Proceedings ANLP-NAACL 2000, </i>Seattle, WA.</p><p>Michael Kaisser and Bonnie Webber. 2007. Question an­swering based on semantic roles. In <i>Proceedings ofthe ACL 07 Workshop on Deep Linguistic processing.</i></p><p>Martin Kay, Jean Mark Gawron, and Peter Norvig. 1994. <i>Verbmobil: A Translation System for Face-To-Face Dia­log. </i>CSLI Press, Stanford, California.</p><p>Karin Kipper, Hoa Trang Dang, and Martha Palmer. 2000. Class-based construction of a verb lexicon. In <i>Pro­ceedings of the 7th Conference on Artificial Intelligence (AAAI-00) and of the 12th Conference on Innovative Ap­plications of Artificial Intelligence (IAAI-00), </i>pages 691­696, Menlo Park, CA, July 30- 3. AAAI Press.</p><p>Namhee Kwon and Eduard H. Hovy. 2006. Integrating se­mantic frames from multiple sources. In Alexander F. Gelbukh, editor, <i>CICLing, </i>volume 3878 of <i>Lecture Notes in Computer Science, </i>pages 1-12. Springer.</p><p>Mitchell Marcus, Grace Kim, Mary Ann Marcinkiewicz, Robert MacIntyre, Ann Bies, Mark Ferguson, Karen Katz, and Britta Schasberger. 1994. The penn treebank: Annotating predicate argument structure. In <i>Proceed­ings ofthe ARPA Human Language Technology Work­shop.</i></p><p>Mark McConville and Myroslava O. Dzikovska. 2007. Ex­tracting a verb lexicon for deep parsing from FrameNet. In <i>Proceedings ofthe ACL-07 Workshop on Deep Lin­guistic Processing, </i>Prague, Czech Republic, June.</p><p>Adam Meyers, Catherine Macleod, and Ralph Grishman. 1996. Standardization of the complement/adjunct dis­tinction. In <i>Proceedings of the Seventh Euralex Interna­tion Congress, Gothenburg, Sweden.</i></p><p>Martha Palmer, Paul Kingsbury, and Daniel Gildea. 2005. The proposition bank: An annotated corpus of semantic roles. <i>Computational Linguistics, </i>31(1):71-106.</p><p>Lei Shi and Rada Mihalcea. 2005. Putting pieces together: Combining framenet, verbnet and wordnet for robust se­mantic parsing. In <i>Proceedings ofthe Sixth International Conference on Intelligent Text Processing and Computa­tional Linguistics, </i>Mexico.</p><p>Mihai Surdeanu, Sanda M. Harabagiu, John Williams, and Paul Aarseth. 2003. Using predicate-argument struc­tures for information extraction. In <i>Proceedings of ACL'03, </i>pages 8-15.</p><p>Jesse Tseng. 2006. English prepositional passives in HPSG. In <i>Proceedings ofFormal Grammar 06, </i>pages 147-159.</p></references></body></article>