<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>Establishing the Upper Bound and Inter-judge Agreement of a Verb Classification Task</title><author surname="Merlo" givenname="Paola"><org  name="University of Geneva" country="Switzerland" city="Geneva"/></author><author surname="Stevenson" givenname="Suzanne"><org  name="Rutgers University" country="USA" city="New Brunswick"/></author></firstpageheader><frontmatter><p><b>Establishing the Upper Bound and Inter-judge Agreement of a Verb Classification Task</b></p><p><b>Paola Merlo* and Suzanne Stevenson</b></p><p>* University of Geneva Department of Linguistics 2 rue de Candolle 1211 Geneve 4 Switzerland merlo@lettres.unige.ch</p><p>t Department of Computer Science, and Center for Cognitive Science Rutgers University 110 Frelinghuysen Road Piscataway, NJ 08854-8019 USA suzanne@cs.rutgers.edu</p></frontmatter><abstract>Detailed knowledge about verbs is critical in many NLP and IR tasks, yet manual determination of such knowledge for large numbers of verbs is difficult, time-consuming and resource intensive. Recent responses to this problem have attempted to classify verbs automatically, as a first step to automatically build lexical resources. In order to estimate the upper bound of a verb classification task, which appears to be difficult and subject to variability among experts, we investigated the performance of human experts in controlled classification experiments. We report here the results of two experiments—using a forced-choice task and a non-forced choice task—which measure human expert accuracy (compared to a gold standard) in classifying verbs into three pre-defined classes, as well as inter-expert agreement. To preview, we find that the highest expert accuracy is 86.5% agreement with the gold standard, and that inter-expert agreement is not very high (K between .53 and .66). The two experiments show comparable results. </abstract></header><body><section number="1." title="Introduction"><p>Automatic lexical acquisition is a fundamental problem in natural language processing (Boguraev and Pustejovsky, 1996). Detailed knowledge about verbs in particular is crit­ical in many NLP and IR tasks, yet manual determination of such knowledge for large numbers of verbs is a dif­ficult, time-consuming and resource intensive task (Dang et al., 1998; Fellbaum, 1998; Levin, 1993; Miller et al., 1990). Furthermore, while syntactic properties of verbs such as subcategorization frames may be gleaned from a machine readable dictionary (Dorr, 1997), or from exam­ples of usage in a corpus (Brent, 1993; Briscoe and Car­roll, 1997; Lapata, 1999; Manning, 1993; McCarthy and Korhonen, 1998), the extraction of semantic properties of verbs poses a more challenging problem. On the assump­tion that syntactic properties of verb usage, and their fre­quency distributions, reflect underlying semantic proper­ties of the verb, recent research has developed statistical corpus-based methods for learning selectional restrictions (Resnik, 1996; Riloff and Schmelzenbach, 1998), verbal aspect (Klavans and Chodorow, 1992; Siegel, 1999), and lexical semantic classification (Aone and McKee, 1996; La­pata and Brew, 1999; Schulte im Walde, 1998; Stevenson et al., 1999; Stevenson and Merlo, 2000).</p><p>Classification in particular has recently attracted atten­tion, as it imposes a hierarchical organization that enables an NLP system to efficiently maintain and exploit gener­alizations over lexical items (Palmer, to appear). One type of classification of verbs that has inspired recent work is the one in (Levin, 1993), which explores the hypothesis that the semantics of a verb determines the possible syntactic ex­pressions of its arguments. Based on this assumption, Levin has developed a very influential framework, where verbs are classified into semantic classes that are revealed empiri­cally by diathesis alternations—alternations in the syntactic expressions of arguments.</p><p>This kind of approach is very fruitful for NLP for two reasons. First, it organises the verbal lexicon around gener­alizations about argument structure—i.e., the thematic roles assigned to the arguments of a verb. These generalisa­tions play a major role in many language engineering tasks. For example, knowledge about argument structure is cru­cial for dealing with dependency relations in parsing and generation (e.g., (Srinivas and Joshi, 1999; Stede, 1998)): for handling thematic divergences in machine translation (Dorr, 1997); for document profiling in information re­trieval (Klavans and Kan, 1998); and for template filling in information extraction (Riloff and Schmelzenbach, 1998). Secondly, it supports corpus-based approaches. Corpus-based approaches to lexical semantic classification have drawn on Levin's hypothesis (that the semantics of a verb determines the possible syntactic expressions of its argu-ments)by applying the converse reasoning—that is, assum­ing that similar subcategorization alternations correspond to similar underlying semantics. This method is especially promising for automatic lexical acquisition, since it is the syntactic properties of arguments that are more easily ex-tractable from a corpus (Dorr and Jones, 1996; Lapata and 2000).<page local="2"/></p><doubt alpha="62.5" length="56" tooSmall="False" monospace="0.0">Brew, 1999; Schulte im Walde, 1998; Stevenson and Merlo,</doubt><p>These approaches have thus far led to some promis­ing preliminary successes. For example, in using fre­quencies of syntactic features to automatically classifying verbs which share subcategorization frames but differ in argument structure, (Stevenson and Merlo, 2000) achieve 69.5% accuracy in a task whose baseline performance is 34%, and (Lapata and Brew, 1999) achieve 83.9% accuracy compared to a 61.3% baseline, both attaining reductions in error rate of over 50%. However, while accuracies of 70­84% seem like respectable initial results, the problem arises that we simply do not know what is a reasonable expecta­tion for performance in this kind of task.</p><p>Our observation is that (human) classification based on alternations and argument structures of verbs engenders a lively theoretical debate on class membership of verbs, which requires complex linguistic information and exper­tise. This leads us to believe that the task is intrinsically difficult, and also likely to show differences in classifica­tion between experts. There is reason then to hypothesize that the expert-defined upper bound of classification algo­rithms will be lower than 100% accuracy, and show vari­ability. Our goal then is to determine a realistic approxima­tion to this upper bound, in order to enable more informed evaluation of verb classification algorithms.</p><p>We report here the results of two experiments—using a forced-choice task and a non-forced choice task—which measure expert accuracy (compared to a gold standard) in classifying verbs into three pre-defined classes, as well as inter-expert agreement. The forced-choice version corre­sponds most closely to the typical computational classifi­cation task, while the non-forced choice version provides a measure on what is assumed to be a more natural task for a human expert. To preview our results, we find that the highest expert accuracy is 86.5% agreement with the gold standard, and that expert agreement is not very high (a kappa value between .53 and .66). Moreover, the two experiments show comparable results.</p><p>This study provides several results of general interest:</p><p>1. It provides the experts' accuracy on a controlled clas­sification task—providing the upper bound of perfor­mance for an automatic verb classification algorithm on a representative classification task.</p><p>2. It provides the degree of agreement between experts in both experiments—indicating how stable the expert performance is.</p><p>3. It provides results that are comparable across the two experiments—indicating that the controlled experi­ment is representative of a natural classification task.</p></section><section number="2." title="The Classification Task"><p>In choosing a representative type of classification task for our experiments on expert performance, we focused on the verb classification task that we have been addressing through automatic corpus-based learning methods (Steven­son and Merlo, 1999; Stevenson and Merlo, 2000). In our computational experiments, we have been using frequen­cies of syntactic distributions from a very large corpus to trainan automatic classifier to discriminate three lexical se­mantic classes from (Levin, 1993). This choice of task for our experiments with human experts clearly helps us in the practical problem of evaluating our own automatic classi­fier. However, the task is also representative of much of the recent work on verb classification, which is largely inspired by Levin's work, as discussed above.</p><p>Thus, we adopted a definition of verb classes based on their argument structure, and we used Levin's classification as the benchmark. Our experimental focus is also notewor­thy in that the differences in the classes lies in their <i>ar­gument structure, </i>not in their possible subcategorizations. That is, the three verb classes represented in the stimuli each support both the transitive and intransitive subcatego-rization frame. This allows us to manipulate only the ar­gument structure variable, while holding subcategorization constant across the classes.</p><p>The stimuli were thus chosen from the following three classes, as exemplified in these sentences:</p><p>Unergative:   The horse raced past the barn.</p><p>The jockey raced the horse past the barn.</p><p><u>Unaccusative</u>: The butter melted in the pan.</p><p>The cook melted the butter in the pan.</p><p>Object-drop: The boy washed the hall. The boy washed.</p><p>Each class is distinguished by the content of the thematic roles assigned by the verb. For object-drop verbs, the subject is an Agent and the optional object is a Theme, yielding the thematic assignments (Agent, Theme) and (Agent) for the transitive and intransitive alternants respec­tively. Unergatives and unaccusatives differ from object-drop verbs in participating in the causative alternation, and also differ from each other in their core thematic argument. In an intransitive unergative, the subject is an Agent, and in an intransitive unaccusative, the subject is a Theme. In the causative transitive form of each, this core semantic argu­ment is expressed as the direct object, with the addition of a Causal Agent (the causer of the action) as subject in both cases.</p><p>The fact that the classes from which we draw the experi­mental stimuli are distinguished by argument structure and not subcategorization is important for two reasons. First, the task represents a key problem in verb classification, as alluded to above, since the ability to automatically learn thematic distinctions among verbs is a necessary compo­nent of automatic lexicon acquisition. Thus, it is important to determine a realistic upper bound on performance for this kind of task. Second, this property entails a fine-grained discrimination task that provides a challenging testbed for automatic classification, and (we hypothesize) for human expert classification as well. The verbs cannot be distin­guished simply by the subcategorization frames they allow,<page local="3"/></p><doubt alpha="37.5" length="8" tooSmall="False" monospace="0.0">E1 E2 E3</doubt><p>Table 1: Percent Agreement (%Agr) and Pair-wise Agree­ment ( , Calculated by the Kappa Statistic) of Three Ex­perts (E1, E2, E3), Compared to Each Other and to a Gold Standard (Levin).</p><p>enabling us to explore the level of performance that can be expected when discriminating verbs based on argument structure alone.</p></section><section number="3." title="Experiments"><p>The format, the task and the stimuli of both experi­ments where chosen to produce a controlled experimen­tal situation. We decided to measure the upper bound for a very narrowly characterised classification task first, and to measure a slightly more relaxed version of it in com­parison. Moreover, we needed to be able to compare re­sponses. Therefore, we performed a closed-form ques­tionnaire study, where the number and types of the target classes are defined in advance, for which we prepared a forced-choice and a non-forced-choice variant. The forced-choice study provides data for a maximally restricted ex­perimental situation, which corresponds most closely to the automatic verb classification task. However, we are also in­terested in slightly more natural results—provided by the non-forced-choice task—where the experts can assign the verbs to an "Others" category.</p><subsection number="3.1." title="Forced-Choice Experiment"><p>We asked three experts in lexical semantics (all native speakers of English) to complete a forced-choice electronic questionnaire study. Materials consisted of individually randomized lists of 59 verbs selected using Levin's index (Levin, 1993). The verbs were to be classified into three tar­get classes—unergative, unaccusative, and object-drop— which were described in the instructions. The definitions of the classes were as follows. Unergative: A verb that as­signs an agent theta role to the subject in the intransitive. If it is able to occur transitively, it can have a causative mean­ing. Unaccusative: A verb that assigns a patient/theme theta role to the subject in the intransitive. When it occurs tran­sitively, it has a causative meaning. Object-Drop: A verb that assigns an agent role to the subject and patient/theme role to the object, which is optional. When it occurs transi­tively, it does not have a causative meaning. There were 20 unergative verbs, 19 unaccusative verbs, and 20 object-drop verbs. The list of verbs is given in the appendix; complete materials and instructions are available from the authors.</p><p>We calculated the pairwise agreement of the experts' classifications, with the gold standard (Levin) and with each other. The results are indicated in Table 3.1. The per­centage of verbs to which they all gave the same classifica­tion (60%) is smaller than any of the pairwise agreements, indicating that the experts do not all agree on the same sub­set of verbs. Assessing the percentage of verbs on which the experts agree gives us an intuitive measure. However, this measure does not take into account how much the ex­perts agree <i>over </i>the expected agreement by chance. This value is provided by the Kappa statistic, , which we cal­culated following (Klauer, 1987, pages 55-57), to measure the experts' degree of agreement over chance, with the gold standard and with each other (using the distribution to determine significance; p&lt;0.001 for all reported results). These results are also shown inTable 3.1.</p><p>Expected chance agreement varies with the number and the relative proportions of categories used by the experts. This means that two given pairs of experts might reach the same percent agreement on a given task, but not have the same expected chance agreement, if they assigned verbs to classes in different proportions. The Kappa statistic ranges from 0, for no agreement above chance, to 1, for perfect agreement. The interpretation of the scale of agreement de­pends on the domain. (Carletta, 1996) cites the conven­tion from the domain of content analysis indicating that</p><p>indicates marginal agreement, while is an indication of good agreement. However, we can ob­serve that only one of our agreement figures almost reaches what would be considered "good" under this interpretation. Given the very high level of expertise of our human experts, we suspect then that this is too stringent a scale for our task, which is qualitatively quite different from content analysis.</p><p>Evaluating the experts' performance summarized in Ta­ble 3.1., we can remark two main things, which confirm our expectations. First, the percent agreement with the gold standard is not close to 100% even by trained experts. The best accuracy of a human expert (in comparison with the gold standard) is 86.5%, and the average accuracy of the three experts is 80%. Second, with respect to comparison of the experts among themselves, the rate of agreement is never very high, and the variability in agreement is consid­erable, ranging from .53 to .66. We can conclude then that the task we have set is very difficult. Based on these results, we note that 86.5% is a more realistic upper bound (than 100%) for evaluating machine performance in this three-way classification task.</p><p>Examining the pattern of average agreement between the experts and Levin, we find agreement on 17.7 (of 20) unergatives, 16.7 (of 19) unaccusatives, and 11.3 (of 20) object-drops. This clearly indicates that the object-drop class is the most difficult for the human experts to define. This class is the most heterogeneous in our verb list, con­sisting of verbs from several subclasses of the "unspecified object alternation" class in (Levin, 1993). We conclude that the verb classification task is likely easier for very homoge­neous classes, and more difficult for more broadly defined classes, even when the exemplars share the critical syntac­tic behaviours.</p><p>The results from this first experiment are useful as they provide an upper bound for our task in a very controlled ex­perimental situation, and they confirm the original expecta­tions. Nonetheless, one possible shortcoming of the above experiment is that the forced-choice task, while maximally comparable to our computational experiments, may not be a natural one for human experts.<page local="4"/> To explore this issue, we performed a non-forced choice version of the experiment.</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p>%Agr</p></td><td class="cell"><p><i>K</i></p></td><td class="cell"><p>%Agr</p></td><td class="cell"><p><i>K</i></p></td><td class="cell"><p>%Agr <i>K</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>E2</p></td><td class="cell"><p>75%</p></td><td class="cell"><p>.59</p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>E3</p></td><td class="cell"><p>70%</p></td><td class="cell"><p>.53</p></td><td class="cell"><p>77%</p></td><td class="cell"><p>.66</p></td><td class="cell"><p></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Levin</p></td><td class="cell"><p>71%</p></td><td class="cell"><p>.56</p></td><td class="cell"><p>86.5%</p></td><td class="cell"><p>.80</p></td><td class="cell"><p>83% .74</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></subsection><subsection number="3.2." title="Non-Forced Choice Experiment"><p>We asked two additional experts in lexical semantics to complete the non-forced-choice electronic questionnaire study. (One expert was a native speaker of English and one bi-lingual; neither participated in the first experiment.) In addition to the three verb classes of interest, an answer of "Others" was allowed, for verbs that the experts could not classify within the three target classes. Materials con­sisted of individually randomized lists of 119 target and filler verbs taken from the electronic version of Levin's in­dex, available through Chicago University Press. The tar­gets were again the same 59 verbs used in the forced-choice experiment. To avoid unwanted priming of target items, the 60 fillers were carefully selected from the set of verbs that do not share any class in Levin's index with any of the senses of the 59 target verbs.</p><p>In this task, if we take only the target items into ac­count, the experts agreed 74.6% of the time (#=0.64) with each other, and 86% (#= 0.80) and 69% (#= 0.57) with the gold standard. (If we take all the verbs into consid­eration, they agreed in 67% of the cases (#=0.56) with each other, and 68% (#=0.55) and 60.5% (#= 0.46) with the gold standard, respectively.) These results show that the forced-choice and non-forced-choice task are compara­ble in accuracy of classification and inter-judge agreement on the target classes, giving us confidence that the forced-choice results provide a reasonably stable upper bound for computational experiments.</p></subsection></section><section number="4." title="General Discussion"><p>Going back to the original issues raised in the introduc­tion, we can then draw the following conclusions. First, accuracy by experts in the 3-way forced choice task is be­tween 71% and 86.5%, confirming the initial intuition that the task is difficult. Since it is not likely for the classifi­cation task to be solved at more than expert accuracy by automatic methods, we can conclude that 86.5% is a more realistic upper limit on the performance that we can expect in this task from an automatic classifier.</p><p>Second, the degree of agreement among experts in both tasks is relatively low. Carletta (1996) suggests that a kappa value of 0.8 or greater is an indication of good agreement. In neither of our tasks do highly trained experts reach that degree of agreement. The À' values we obtain (between .53 and .66) are in a way rather striking because in fact they are quite low.</p><p>One might think that if classification is so difficult, then it might not be useful. We remark however that an indi­vidual expert consistently over- or under-estimates specific classes (compared to the gold standard). The fact that the experts make consistent mistakes suggests that they inter­preted the defining properties given in the instructions, such as agent or theme, differently from each other. If this is the case, the disagreement can probably be partly solved by dis­cussion and negotiation, as is done for instance in tagging and bracketting a corpus. If establishing a new classifica­tion, we can therefore reccommend to follow corpus annotation practice, and adopt a two stage classification process, where experts can discuss results after the first stage, and reclassify verbs after discussion.</p><p>Furthermore, it might be objected that the classes we have chosen make the task too hard. We should recall that the way we have chosen the verbs kept their subcate-gorization alternations constant between the classes, while changing only their argument structure. This set-up likely constitutes a more difficult discrimination task than when both argument structures and subcategorization frames can vary. From existing classification work, we know that many verbs can be classified by only looking at their subcate-gorization frames (Lapata and Brew, 1999). However, in some cases, subcategorization information might not be sufficiently informative for the purposes of classification. Our verbs are a case in point, since the three classes share subcategorization frames but differ in their pattern of the­matic assignments. Thus, it seemed preferable to us to keep the subcategorization frame constant rather than varying both argument structures and subcategorization frames at the same time, as it provided a controlled experimental sit­uation.</p><p>Finally, we observe that the measure of agreement across the two tasks is similar for the target items: in the first task, #=.53, .59, and .66 among experts, and À'=.56, .74, and .80 with the gold standard; in the second task, A'=.64 among experts and À'=.57 and .80 with the gold standard. The average accuracy in the two experiments is also very close: 80% in the first, and 77% in the second. We conclude that, although not ideal, a forced-choice task can be representative of a natural task, and is informative about the difficulty of more open-ended classifications.</p></section><section number="5." title="Conclusions"><p>In conclusion, we have demonstrated an expert-based upper bound of 86.5%, far below the default maximum ac­curacy of 100%, in a lexical semantic classification task ap­plied to verbs that share subcategorization frames but dif­fer in argument structure (i.e., in the thematic roles they assign). Our results indicate that in a verb classification task of this kind, which depends on complex linguistic in­formation and relationships, it is important to experimen­tally determine a realistic upper bound of accuracy for the task from human expert performance, in order to enable in­formed evaluation of automatic methods.</p></section><references><p>Aone, Chinatsu and Douglas McKee, 1996. Acquiring predicate-argument mapping information in multilingual texts. In Branimir Boguraev and James Pustejovsky (eds.), <i>Corpus Processing for Lexical Acquisition. </i>MIT Press, pages 191-202. Boguraev, Branimir and James Pustejovsky, 1996. Issues in text-based lexicon acquisition. In Branimir Boguraev and James Pustejovsky (eds.), <i>Corpus Processing for Lexical Acquisition. </i>MIT Press, pages 3-20. Brent, Michael, 1993. From grammar to lexicon: Unsuper-vised learning oflexical syntax. <i>ComputationalLinguis-</i> <i>tics, </i>19(2):243-262.</p><p>Briscoe, Ted and John Carroll, 1997. Automatic extraction ofsubcategorizationfrom corpora. In<i>Proceedings ofthe Fifth Applied Natural Language Processing Conference. </i>Carletta, Jean, 1996. Assessing agreement on classification tasks: the Kappa statistics. <i>Computational Linguistics,</i> 22(2):249-254.</p><p>Dang, Hoa Trang, Karin Kipper, Martha Palmer, and Joseph Rosenzweig, 1998. Investigating regular sense extensions based on intersective Levin classes. In <i>Proc. of the 36th Annual Meeting of the ACL and the 17th International Conference on Computational Linguistics (COLING-ACL '98). </i>Montreal, Canada: Universite de Montreal.</p><p>Dorr, Bonnie, 1997. Large-scale dictionary construction for foreign language tutoring and interlingual machine translation. <i>Machine Translation, </i>12:1-55. Dorr, Bonnie and Doug Jones, 1996. Role of word sense disambiguation in lexical acquisition: Predicting seman­tics from syntactic cues. In <i>Proc. of the 16th Interna­tional Conference on Computational Linguistics. </i>Copen­hagen, Denmark. Fellbaum, Christiane, 1998. <i>Wordnet: an Electronic Lexi­cal Database. </i>MIT Press. Klauer, Klaus J., 1987. <i>Kriteriumsorientierte Tests</i><i>.Got-</i></p><p>tingen, Germany: Verlag fur Psychologie. Klavans, Judith and Min-Yen Kan, 1998. Role of verbs in document analysis. In <i>Proc. ofthe 36th Annual Meet-ingofthe ACL and the 17th InternationalConference on Computational Linguistics (COLING-ACL '98). </i>Mon­treal, Canada: Universite de Montreal. Klavans, Judith L. and Martin Chodorow, 1992. Degrees of stativity: The lexical representation of verb aspect. In <i>Proceedings of the Fourteenth International Conference on Computational Linguistics. </i>Lapata, Maria, 1999. Acquiring lexical generalizations from corpora: A case study for diathesis alternations. In <i>Proceedings ofthe 37th Annual Meeting ofthe Associ­ation for Computational Linguistics (ACL'99). </i>College</p><p>Park, MD.</p><p>Lapata, Maria and Chris Brew, 1999. Using subcategoriza-tion to resolve verb class ambiguity. In <i>Proceedings of Joint SIGDAT Conference on Empirical Methods in Nat­ural Language. </i>College Park, MD.</p><p>Levin, Beth, 1993. <i>English Verb Classes and Alternations. </i>Chicago, IL: University ofChicago Press.</p><p>Manning, Christopher D., 1993. Automatic acquisition of a large subcategorization dictionary from corpora. In <i>Proceedings ofthe 31st Annual Meeting ofthe Associ­ation for Computational Linguistics. </i>Ohio State Univer­sity, Association for Computational Linguistics.</p><p>McCarthy, Diana and Anna Korhonen, 1998. Detecting verbal participation in diathesis alternations. In <i>Proc. ofthe 36th Annual Meeting ofthe ACL and the 17th International Conference on Computational Linguistics (COLING-ACL '98). </i>Montreal, Canada: Universite de Montreal.</p><p>Miller, George, R. Beckwith, C. Fellbaum, D. Gross, and K. Miller, 1990. Five papers on Wordnet. Technical re­port, Cognitive Science Lab, PrincetonUniversity.</p><p>Palmer, Martha, to appear. Consistent criteria for sense dis­tinctions. <i>Computing for the Humanities.</i></p><p>Resnik, Philip, 1996. Selectional constraints: an information-theoretic model and its computational real­ization. <i>Cognition, </i>61(1-2):127-160.</p><p>Riloff, Ellen and Mark Schmelzenbach, 1998. An empir­ical approach to conceptual case frame acquisition. In <i>Proceedings of the Sixth Workshop on Very Large Cor­pora.</i></p><p>Schulte im Walde, Sabine, 1998. Automatic semantic classification of verbs according to their alternation be­haviour. Technical Report AIMS Report 4(3), Institutfur Maschinelle Sprachverarbeitung, Universitat Stuttgart.</p><p>Siegel, Eric, 1999. Corpus-based linguistic indicators for aspectual classification. In <i>Proceedings ofACL'99. </i>Col­lege Park, MD: University ofMaryland.</p><p>Srinivas, Bangalore and Aravind K. Joshi, 1999. Supertag-ging: An approach to almost parsing. <i>Computational Linguistics, </i>25(2):237-265.</p><p>Stede, Manfred, 1998. A generative perspective on verb alternations. <i>ComputationalLinguistics, </i>24(3):401-430.</p><p>Stevenson, Suzanne and Paola Merlo, 1999. Verb classi­fication using distributions of grammatical features. In <i>Proc. ofthe 9th Conference ofthe European Chapter for Computational Linguistics (EACL'99). </i>Bergen,Norway.</p><p>Stevenson, Suzanne and Paola Merlo, 2000. Automatic lexical acquisition based on statistical distributions. In <i>Proceedings </i><i>ofthe</i><i> 18th International Conference on Computational Linguistics (COLING 2000). </i>Saar-bruecken, Germany.</p><p>Stevenson, Suzanne, Paola Merlo, Natalia Kariaeva, and Kamin Whitehouse, 1999. Supervised learning of lexi­cal semantic verb classes using frequency distributions. In <i>Proceedings of SigLex99: Standardizing Lexical Re­sources (SigLex'99)</i>. College Park, Maryland.</p></references><appendix title="Appendix"><p>The unergatives are manner of motion verbs: <i>floated, galloped, glided, hiked, hopped, hurried, jogged, jumped, leaped, marched, paraded, raced, rushed, scooted, scur­ried, skipped, tiptoed, trotted, vaulted, wandered.</i></p><p>The unaccusatives are verbs of change of state: <i>boiled, changed, cleared, collapsed, cooled, cracked, dissolved, divided, exploded, flooded, folded, fractured, hardened, melted, opened, simmered, solidified, stabilized, widened.</i></p><p>The object-drop verbs are unspecified object alternation verbs: <i>borrowed, called, carved, cleaned, danced, inher­ited, kicked, knitted, organised, packed, painted, played,</i> <i>reaped, rented, sketched, studied, swallowed, typed, washed, yelled.</i><page local="5"/><i></i></p></appendix></body></article>