<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1" global="99"/><title>SemEval-2007 Task 19: Frame Semantic Structure Extraction</title><pubinfo>Proceedings of the 4th International Workshop on Semantic Evaluations (SemEval-2007),pages 99-104, Prague, June 2007. ©2007 Association for Computational Linguistics</pubinfo><author surname="Baker" givenname="Collin"><org  name="University of Jena" country="Germany" city="Jena"/></author><author surname="Ellsworth" givenname="Michael"><org  name="University of Jena" country="Germany" city="Jena"/></author><author surname="Erk" givenname="Katrin"><org  name="University of Jena" country="Germany" city="Jena"/></author></firstpageheader><frontmatter><p><b>SemEval'07 Task 19: Frame Semantic Structure Extraction</b></p><p><b>Collin Baker, Michael Ellsworth</b></p><p>International Computer Science Institute Berkeley, California {collinb,infinity} @icsi.berkeley.edu</p><p><b>Katrin Erk</b></p><p>Computer Science Dept. University of Texas Austin</p><p>katrin.erk@mail.utexas.edu</p></frontmatter><abstract>This task consists of recognizing words and phrases that evoke semantic <b>frames </b>as defined in the FrameNet project (http: //framenet.icsi.berkeley. edu), and their semantic dependents, which are usually, but not always, their syntactic dependents (including subjects). The train­ing data was FN annotated sentences. In testing, participants automatically annotated three previously unseen texts to match gold standard (human) annotation, including pre­dicting previously unseen frames and roles. Precision and recall were measured both for matching of labels of frames and FEs and for matching of semantic dependency trees based on the annotation. </abstract></header><body><section number="1" title="Introduction"><p>The task of labeling frame-evoking words with ap­propriate frames is similar to WSD, while the task of assigning frame elements is called <b>Semantic Role Labeling (SRL), </b>and has been the subject of several shared tasks at ACL and CoNLL. For example, in the sentence "Matilde said, 'I rarely eat rutabaga,"' <i>said </i>evokes the Statement frame, and <i>eat </i>evokes the Ingestion frame. The role of SPEAKER in the Statement frame is filled by <i>Matilda, </i>and the role of MESSAGE, by the whole quotation. In the Inges­tion frame, <i>I </i>is the Inges tor and <i>rutabaga </i>fills the INGES TIB LES role. Since the ingestion event is con­tained within the MESSAGE of the Statement event, we can represent the fact that the message conveyed was about ingestion, just by annotating the sentence with respect to these two frames.</p><p>After training on FN annotations, the participants' systems labeled three new texts automatically. The evaluation measured precision and recall for frames and frame elements, with partial credit for incorrect but closely related frames. Two types of evaluation were carried out: <b>Label matching evaluation, </b>in which the participant's labeled data was compared directly with the gold standard labeled data, and <b>Se­mantic dependency evaluation, </b>in which both the gold standard and the submitted data were first con­verted to semantic dependency graphs in XML for­mat, and then these graphs were compared.</p><p>There are three points that make this task harder and more interesting than earlier SRL tasks: (1) while previous tasks focused on role assignment, the current task also comprises the identification of the appropriate FrameNet frame, similar to WSD, (2) the task comprises not only the labeling of individ­ual predicates and their arguments, but also the inte­gration of all labels into an overall <b>semantic depen­dency graph, </b>a partial semantic representation of the overall sentence meaning based on frames and roles, and (3) the test data includes occurrences of frames that are not seen in the training data. For these cases, participant systems have to identify the closest known frame. This is a very realistic sce­nario, encouraging the development of robust sys­tems showing graceful degradation in the face of un­known events.</p><page local="2" global="100"/></section><section number="2" title="Frame semantics and FrameNet"><p>The basic concept of Frame Semantics is that many words are best understood as part of a group of terms that are related to a particular type of event and the participants and "props" involved in it (Fill-more, 1976; Fillmore, 1982). The classes of events are the semantic <b>frames </b>of the <b>lexical units (LU</b>s) that evoke them, and the roles associated with the event are referred to as <b>frame elements (FE</b>s). The same type of analysis applies not only to events but also to relations and states; the frame-evoking ex­pressions may be single words or multi-word ex­pressions, which may be of any syntactic category. Note that these FE names are quite frame-specific; generalizations over them are expressed via explicit FE-FE relations.</p><p>The Berkeley FrameNet project (hereafter FN) (Fillmore et al., 2003) is creating a computer- and human-readable lexical resource for English, based on the theory of frame semantics and supported by corpus evidence. The current release (1.3) of the FrameNet data, which has been freely available for instructional and research purposes since the fall of 2006, includes roughly 780 frames with roughly 10,000 word senses (lexical units). It also contains roughly 150,000 annotation sets, of which 139,000 are lexicographic examples, with each sentence an­notated for a single predicator. The remainder are from full-text annotation in which each sentence is annotated for all predicators; 1,700 sentences are an­notated in the full-text portion of the database, ac­counting for roughly 11,700 annotation sets, or 6.8 predicators (=annotation sets) per sentence. Nearly all of the frames are connected into a single graph by frame-to-frame relations, almost all of which have associated FE-to-FE relations (Fillmore et al., 2004a)</p><subsection number="2.1" title="Frame Semantics of texts"><p>The ultimate goal is to represent the lexical se­mantics of all the sentences in a text, based on the relations between predicators and their depen­dents, including both phrases and clauses, which may, in turn, include other predicators; although this has been a long-standing goal of FN (Fillmore and Baker, 2001), automatic means of doing this are only now becoming available.</p><p>Consider a sentence from one of the testing texts:</p><p>(1) This geography is important in understanding</p><p>Dublin.</p><p>In the frame semantic analysis of this sentence, there are two predicators which FN has analyzed: <i>important </i>and <i>understanding, </i>as well as one which we have not yet analyzed, <i>geography. </i>In addition, <i>Dublin </i>is recognized by the NER system as a loca­tion. In the gold standard annotation, we have the annotation shown in (2) for the Importance frame, evoked by the target <i>important, </i>and the annotation shown in (3) for the Grasp frame, evoked by <i>under­standing.</i></p><p>(2) [factor This geography] [cop is] IMPOR­TANT [undertaking in understanding Dublin].</p><p>tlNTERESTED_party <i>^NI]</i></p><p>(3) This geography is important in UNDER­STANDING Phenomenon Dublin]. [Cognizer</p><doubt alpha="75.0" length="4" tooSmall="False" monospace="0.0">CNI]</doubt><p>The definitions of the two frames begin like this:</p><p>Importance: A Factor affects the outcome of an Undertaking, which can be a goal-oriented activ­ity or the maintenance of a desirable state, the work in a Field, or something portrayed as affecting an Interested _party. ..</p><p>Grasp: A Cognizer possesses knowledge about the workings, significance, or meaning of an idea or object, which we call Phenomenon, and is able to make predictions about the behavior or occurrence of the Phenomenon. ..</p><p>Using these definitions and the labels, and the fact that the target and FEs of one frame are subsumed by an FE of the other, we can compose the mean­ings of the two frames to produce a detailed para­phrase of the meaning of the sentence: Something denoted by <i>this geography </i>is a factor which affects the outcome of the undertaking of understanding the location called "Dublin" by any interested party. We have not dealt with <i>geography </i>as a frame-evoking expression, although we would eventually like to. (The preposition <i>in </i>serves only as a marker of the frame element Undertaking.)</p><p>In (2), the interested .party is not a label on any part of the text; rather, it is marked INI, for "in­definite null instantiation", meaning that it is con­ceptually required as part of the frame definition, absent from the sentence, and not recoverable from the context as being a particular individual-meaning that <i>this geography </i>is important for anyone in gen­eral's understanding of Dublin.<page local="3" global="101"/> In (3), the COG­NIZER is "constructionally null instantiated", as the gerund <i>understanding </i>licenses omission of its sub­ject. The marking of null instantiations is important in handling text coherence and was part of the gold standard, but as far as we know, none of the partici­pants attempted it, and it was ignored in the evalua­tion.</p><p>Note that we have collapsed the two null instan­tiated FEs, the INTERESTED_PARTY of the impor­tance frame and the COGNIZER in the Grasp frame, since they are not constrained to be distinct.</p></subsection><subsection number="2.2" title="Semantic dependency graphs"><p>Since the role fillers are dependents (broadly speak­ing) of the predicators, the full FrameNet annotation of a sentence is roughly equivalent to a dependency parse, in which some of the arcs are labeled with role names; and a dependency graph can be derived algo-rithmically from FrameNet annotation; an early ver­sion of this was proposed by (Fillmore et al., 2004b) Fig. 1 shows the semantic dependency graph de­rived from sentence (1); this graphical representa­tion was derived from a semantic dependency XML file (see Sec. 5). It shows that the top frame in this sentence is evoked by the word <i>important, </i>although the syntactic head is the copula <i>is </i>(here given the more general label "Support"). The labels on the arcs are either the names of frame elements or indi­cations of which of the daughter nodes are seman­tic heads, which is important in some versions of the evaluation. The labels on nodes are either frame names (also colored gray), syntactic phrases types (e.g. NP), or the names of certain other syntactic "connectors", in this case, Marker and Support.</p></subsection></section><section number="3" title="Definition of the task 3.1   Training data"><p>The major part of the training data for the task con­sisted of the current data release from FrameNet (Release 1.3), described in Sec.2 This was supple­mented by additional training data made available through SemEval to participants in this task. In ad­dition to updated versions of some of the full-text an­notation from Release 1.3, three files from the ANC were included: from Slate.com, "Stephanopoulos</p><figure caption="Figure 1: Sample Semantic Dependency Graph"></figure><p>Crimes" and "Entrepreneur as Madonna", and from the Berlitz travel guides, "History of Jerusalem".</p><subsection number="3.2" title="Testing data"><p>The testing data was made up of three texts, none of which had been seen before; the gold standard consisted of manual annotations (by the FrameNet team) of these texts for all frame evoking expres­sions and the fillers of the associated frame ele­ments. All annotation of the testing data was care­fully reviewed by the FN staff to insure its cor­rectness. Since most of the texts annotated in the FN database are from the NTI website (www.nti. org), we decided to take two of the three test­ing texts from there also. One, "China Overview", was very similar to other annotated texts such as "Taiwan Introduction", "Russia Overview", etc. available in Release 1.3. The other NTI text, "Work Advances", while in the same domain, was shorter and closer to newspaper style than the rest of the NTI texts.   Finally, the "Introduction to<page local="4" global="102"/></p><p>Dublin", taken from the American National Cor­pus (ANC,www.americannationalcorpus. org) Berlitz travel guides, is of quite a different genre, although the "History of Jerusalem" text in the training data was somewhat similar. Table 1 gives some statistics on the three testing files. To give a flavor of the texts, here are two sentences; frame evoking words are in boldface:</p><p>From "Work Advances": "The <b>Iranians </b>are <b>now willing </b>to <b>accept </b>the <b>installation </b>of cameras only <b>outside </b>the <b>cascade halls, </b>which will not enable the IAEA to <b>monitor </b>the <b>entire uranium enrichment process," </b>the <b>diplomat said.</b></p><p>From "Introduction to Dublin": And <b>in </b>this <b>city, where literature </b>and <b>theater </b>have <b>historically dominated </b>the scene, visual <b>arts </b>are <b>finally </b>com­ing into their own with the <b>new Museum </b>of Modern <b>Art </b>and the <b>many galleries </b>that display the work of <b>modern Irish artists.</b></p></subsection></section><section number="4" title="Participants"><p>A number of groups downloaded the training or test­ing data, but in the end, only three groups submitted results: the UTD-SRL group and the LTH group, who submitted full results, and the CLR group who submitted results for frames only. It should also be noted that the LTH group had the testing data for longer than the 10 days allowed by the rules of the exercise, which means that the results of the two teams are not exactly comparable. Also, the results from the CLR group were initially formatted slightly differently from the gold standard with regard to character spacing; a later reformatting allowed their results to be scored with the other groups'.</p><p>The LTH system used only SVM classifiers, while the UTD-SRL system used a combination of SVM and ME classifiers, determined experimentally. The CLR system did not use classifiers, but hand-written symbolic rules. Please consult the separate system papers for details about the features used.</p></section><section number="5" title="Evaluation"><p>The labels-only matching was similar to previous shared tasks, but the dependency structure evalua­tion deserves further explanation: The XML seman­tic dependency structure was produced by a program called fttosem, implemented in Perl, which goes sentence by sentence through a FrameNet full-text XML file, taking LU, FE, and other labels and using them to structure a syntactically unparsed piece of a sentence into a syntactic-semantic tree. Two basic principles allow us to produce this tree: (1) LUs are the sole syntactic head of a phrase whose semantics is expressed by their frame and (2) each label span is interpreted as the boundaries of a syntactic phrase, so that when a larger label span subsumes a smaller one, the larger span can be interpreted as a the higher node in a hierarchical tree. There are a fair num­ber of complications, largely involving identifying mismatches between syntactic and semantic headed-ness. Some of these (support verbs, copulas, mod­ifiers, transparent nouns, relative clauses) are anno­tated in the data with their own labels, while oth­ers (syntactic markers, e.g. prepositions, and auxil­iary verbs) must be identified using simple syntactic heuristics and part-of-speech tags.</p><p>For this evaluation, a non-frame node counts as matching provided that it includes the head of the gold standard, whether or not non-head children of that node are included. For frame nodes, the partici­pants got full credit if the frame of the node matched the gold standard.</p><subsection number="5.1" title="Partial credit for related frames"><p>One of the problems inherent in testing against un­seen data is that it will inevitably contain lexical units that have not previously been annotated in FrameNet, so that systems which do not generalize well cannot get them right. In principle, the deci­sion as to what frame to add a new LU to should be helped by the same criteria that are used to assign polysemous lemmas to existing frames. However, in practice this assignment is difficult, precisely be­cause, unlike WSD, there is no assumption that all the senses of each lemma are defined in advance; if the system can't be sure that a new use of a lemma is in one of the frames listed for that lemma, then it must consider all the 800+ frames as possibili­ties.<page local="5" global="103"/> This amounts to the automatic induction of fine-grained semantic similarity from corpus data, a notoriously difficult problem (Stevenson and Joanis, 2003; Schulte im Walde, 2003).</p><table caption="Table 1: Summary of Testing Data" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>Sents</p></td><td class="cell"><p>NEs</p></td><td class="cell"><p>Fran Tokens</p></td><td class="cell"><p>ties Types</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Work</p></td><td class="cell"><p>14</p></td><td class="cell"><p>31</p></td><td class="cell"><p>174</p></td><td class="cell"><p>77</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>China</p></td><td class="cell"><p>39</p></td><td class="cell"><p>90</p></td><td class="cell"><p>405</p></td><td class="cell"><p>125</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dublin</p></td><td class="cell"><p>67</p></td><td class="cell"><p>86</p></td><td class="cell"><p>480</p></td><td class="cell"><p>165</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Totals</p></td><td class="cell"><p>120</p></td><td class="cell"><p>207</p></td><td class="cell"><p>1059</p></td><td class="cell"><p>272</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>For LUs which clearly do not fit into any exist­ing frames, the problem is still more difficult. In the course of creating the gold standard annotation of the three testing texts, the FN team created almost 40 new frames. We cannot ask that participants hit upon the new frame name, but the new frames are not cre­ated in a vacuum; as mentioned above, they are al­most always added to the existing structure of frame-to-frame relations; this allows us to give credit for assignment to frames which are not the precise one in the gold standard, but are close in terms of frame-to-frame relations. Whenever participants' proposed frames were wrong but connected to the right frame by frame relations, partial credit was given, decreas­ing by 20% for each link in the frame-frame relation graph between the proposed frame and the gold stan­dard. For FEs, each frame element had to match the gold standard frame element and contain at least the same head word in order to gain full credit; again, partial credit was given for frame elements related via FE-to-FE relations.</p></subsection></section><section number="6" title="Results"><p>The strictness of the requirement of exact bound­ary matching (which depends on an accurate syntac­tic parse) is compounded by the cascading effect of semantic classification errors, as seen by comparing the F-scores in Table 3 with those in Table 2. The difficulty of the task is reflected in the F-scores of around 35% for the most difficult text in the most difficult condition, but participants still managed to reach F-scores as high as 75% for the more limited task of Frame Identification (Table 2), which more closely matches traditional Senseval tasks, despite the lack of a full sense inventory. The difficulty posed by having such an unconstrained task led to understandably low recall scores in all participants (between 25 and 50%). The systems submitted by the teams differed in their sensitivity to differences in the texts: UTD-SRL's system varied by around 10% across texts, while LTH's varied by 15%.</p><table caption="Table 3: Results for combined Frame and FE recog­nition"></table><p>There are some rather encouraging results also. The participants rather consistently performed bet­ter with our more complex, but also more useful and realistic scoring, including partial credit and grad­ing on semantic dependency rather than exact span match (compare the top and bottom halves of Table 3). The participants all performed relatively well on the frame-recognition task, with precision scores av­eraging 63% and topping 85%.</p></section><section number="7" title="Discussion"><p>The testing data for this task turned out to be espe­cially challenging with regard to new frames, since, in an effort to annotate especially thoroughly, almost 40 new frames were created in the process of an­notating these three specific passages.<page local="6" global="104"/> One result of this was that the test passages had more unseen frames than a random unseen passage, which prob­ably lowered the recall on frames. It appears that this was not entirely compensated by giving partial credit for related frames.</p><table caption="Table 3: Results for combined Frame and FE recognition" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Text</p></td><td class="cell"><p>Group</p></td><td class="cell"><p>Recall</p></td><td class="cell"><p>Prec.</p></td><td class="cell"><p>Fl</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Label matching only</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dublin</p></td><td class="cell"><p>UTD-SRL</p></td><td class="cell"><p>0.27699</p></td><td class="cell"><p>0.55663</p></td><td class="cell"><p>0.36991</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>China</p></td><td class="cell"><p>UTD-SRL</p></td><td class="cell"><p>0.31639</p></td><td class="cell"><p>0.51715</p></td><td class="cell"><p>0.39260</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Work</p></td><td class="cell"><p>UTD-SRL</p></td><td class="cell"><p>0.31098</p></td><td class="cell"><p>0.62408</p></td><td class="cell"><p>0.41511</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dublin</p></td><td class="cell"><p>LTH</p></td><td class="cell"><p>0.36536</p></td><td class="cell"><p>0.55065</p></td><td class="cell"><p>0.43926</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>China</p></td><td class="cell"><p>LTH</p></td><td class="cell"><p>0.39370</p></td><td class="cell"><p>0.54958</p></td><td class="cell"><p>0.45876</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Work</p></td><td class="cell"><p>LTH</p></td><td class="cell"><p>0.41521</p></td><td class="cell"><p>0.61069</p></td><td class="cell"><p>0.49433</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Semantic dependency matching</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dublin</p></td><td class="cell"><p>UTD-SRL</p></td><td class="cell"><p>0.26238</p></td><td class="cell"><p>0.53432</p></td><td class="cell"><p>0.35194</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>China</p></td><td class="cell"><p>UTD-SRL</p></td><td class="cell"><p>0.31489</p></td><td class="cell"><p>0.53145</p></td><td class="cell"><p>0.39546</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Work</p></td><td class="cell"><p>UTD-SRL</p></td><td class="cell"><p>0.30641</p></td><td class="cell"><p>0.61842</p></td><td class="cell"><p>0.40978</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dublin</p></td><td class="cell"><p>LTH</p></td><td class="cell"><p>0.36345</p></td><td class="cell"><p>0.54857</p></td><td class="cell"><p>0.43722</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>China</p></td><td class="cell"><p>LTH</p></td><td class="cell"><p>0.40995</p></td><td class="cell"><p>0.57410</p></td><td class="cell"><p>0.47833</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Work</p></td><td class="cell"><p>LTH</p></td><td class="cell"><p>0.45970</p></td><td class="cell"><p>0.67352</p></td><td class="cell"><p>0.54644</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><table caption="Table 2: Frame Recognition only" class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Text</p></td><td class="cell"><p>Group</p></td><td class="cell"><p>Recall</p></td><td class="cell"><p>Prec.</p></td><td class="cell"><p>Fl</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dublin</p></td><td class="cell"><p>UTD-SRL</p></td><td class="cell"><p>0.4188</p></td><td class="cell"><p>0.7716</p></td><td class="cell"><p>0.5430</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>China</p></td><td class="cell"><p>UTD-SRL</p></td><td class="cell"><p>0.5498</p></td><td class="cell"><p>0.8009</p></td><td class="cell"><p>0.6520</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Work</p></td><td class="cell"><p>UTD-SRL</p></td><td class="cell"><p>0.5251</p></td><td class="cell"><p>0.8382</p></td><td class="cell"><p>0.6457</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dublin</p></td><td class="cell"><p>LTH</p></td><td class="cell"><p>0.5184</p></td><td class="cell"><p>0.7156</p></td><td class="cell"><p>0.6012</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>China</p></td><td class="cell"><p>LTH</p></td><td class="cell"><p>0.6261</p></td><td class="cell"><p>0.7731</p></td><td class="cell"><p>0.6918</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Work</p></td><td class="cell"><p>LTH</p></td><td class="cell"><p>0.6606</p></td><td class="cell"><p>0.8642</p></td><td class="cell"><p>0.7488</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Dublin</p></td><td class="cell"><p>CLR</p></td><td class="cell"><p>0.3984</p></td><td class="cell"><p>0.6469</p></td><td class="cell"><p>0.4931</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>China</p></td><td class="cell"><p>CLR</p></td><td class="cell"><p>0.4621</p></td><td class="cell"><p>0.6302</p></td><td class="cell"><p>0.5332</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Work</p></td><td class="cell"><p>CLR</p></td><td class="cell"><p>0.5054</p></td><td class="cell"><p>0.7452</p></td><td class="cell"><p>0.6023</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p>This task is a more advanced and realistic version of the Automatic Semantic Role Labeling task of Senseval-3 (Litkowski, 2004). Unlike that task, the testing data was previously unseen, participants had to determine the correct frames as a first step, and participants also had to determine FE boundaries, which were given in the Senseval-3.</p><p>A crucial difference from similar approaches, such as SRL with PropBank roles (Pradhan et al., 2004) is that by identifying relations as part of a frame, you have identified a gestalt of relations that enables far more inference, and sentences from the same passage that use other words from the same frame will be easier to link together. Thus, the FN SRL results are translatable fairly directly into formal representations which can be used for rea­soning, question answering, etc. (Scheffczyk et al., 2006; Frank and Semecky, 2004; Sinha and Narayanan, 2005).</p><p>Despite the problems with recall, the participants have expressed a determination to work to improve these results, and the FN staff are eager to collabo­rate in this effort. A project is now underway at ICSI to speed up frame and LU definition, and another to speed up the training of SRL systems is just begin­ning, so the prospects for improvement seem good.</p><p>This material is based in part upon work sup­ported by the National Science Foundation under Grant No. IIS-0535297.</p></section><references><p>Charles J. Fillmore and Collin F. Baker. 2001. Frame semantics for text understanding. In <i>Proceedings of WordNet and Other Lexical Resources Workshop, </i>Pittsburgh, June. NAACL.</p><p>Charles J. Fillmore, Christopher R. Johnson, and Miriam R.L Petruck. 2003. Background to FrameNet. <i>International Journal of Lexicography, </i>16.3:235-250.</p><p>2004a. FrameNet as a "Net". In <i>Proceedings of LREC, </i>volume 4, pages 1091-1094, Lisbon. ELRA.</p><p>Charles J. Fillmore, Josef Ruppenhofer, and Collin F. Baker. 2004b. FrameNet and representing the link between semantic and syntactic relations. In Chu-ren Huang and Winfried Lenders, editors, <i>Frontiers in Linguistics, </i>volume I of <i>Language and Linguisitcs Monograph Series B, </i>pages 19-59. Inst. ofLinguistics, Acadmia Sinica, Taipei.</p><p>Charles J. Fillmore. 1976. Frame semantics and the na­ture of language. <i>Annals ofthe New York Academy of Sciences, </i>280:20-32.</p><p>Charles J. Fillmore. 1982. Frame semantics. In <i>Lin­guistics in the Morning Calm, </i>pages 111-137. Han­shin Publishing Co., Seoul, South Korea.</p><p>Anette Frank and Jiri Semecky. 2004. Corpus-based induction of an LFG syntax-semantics interface for frame semantic processing. In <i>Proceedings of the 5th International Workshop on Linguistically Interpreted Corpora (LINC 2004), </i>Geneva, Switzerland.</p><p>Ken Litkowski. 2004. Senseval-3 task: Automatic label­ing of semantic roles. In Rada Mihalcea and Phil Ed­monds, editors, <i>Senseval-3: Third International Work­shop on the Evaluation ofSystems for the Semantic Analysis ofText, </i>pages 9-12, Barcelona, Spain, July. Association for Computational Linguistics.</p><p>Sameer S. Pradhan, Wayne H. Ward, Kadri Hacioglu, James H. Martin, and Dan Jurafsky. 2004. Shallow semantic parsing using support vector machines. In Daniel Marcu Susan Dumais and Salim Roukos, ed­itors, <i>HLT-NAACL 2004: Main Proceedings, </i>pages 233-240, Boston, Massachusetts, USA, May 2 - May 7. Association for Computational Linguistics.</p><p>Jan Scheffczyk, Collin F. Baker, and Srini Narayanan. 2006. Ontology-based reasoning about lexical re­sources. In Alessandro Oltramari, editor, <i>Proceedings of ONTOLEX 2006, </i>pages 1-8, Genoa. LREC.</p><p>Sabine Schulte im Walde. 2003. Experiments on the choice of features for learning verb classes. In <i>Pro­ceedings of the 10th Conference of the EACL (EACL-</i> <i>03).</i><i></i></p><p>Steve Sinha and Srini Narayanan. 2005. Model based answer selection. In <i>Proceedings ofthe Workshop on Textual Inference, 18th National Conference on Artifi­cial Intelligence, </i>PA, Pittsburgh. AAAI.</p><p>Suzanne Stevenson and Eric Joanis. 2003. Semi-supervised verb class discovery using noisy features. In <i>Proceedings of the 7th Conference on Natural Lan­guage Learning (CoNLL-03), </i>pages 71-78.</p><p>Charles J. Fillmore, Collin F. Baker, and Hiroaki Sato.</p></references></body></article>