<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>The Automated Acquisition of Topic Signatures for Text Summarization</title><author surname="Lin" givenname="Chin-Yew"><org  name="Duke University" country="USA" city="Durham"/></author><author surname="Hovy" givenname="Eduard"><org  name="University of Southern California" country="USA" city="Marina del Rey"/></author></firstpageheader><frontmatter><p>The Automated Acquisition of Topic Signatures for Text</p><p>Summarization</p><p>Chin-Yew Lin and Eduard Hovy</p><p>Information Sciences Institute University of Southern California Marina del Rev, CA 90292, USA {cyl,hovy}@isi.edu</p></frontmatter><abstract>In order to produce a good summary, one has to identify the most relevant portions of a given text. We describe in this paper a method for au­tomatically training topic signatures-sets of related words, with associated weights, organized around head topics-and illustrate with signatures we cre­ated with 6,194 TREC collection texts over 4 se­lected topics. We describe the possible integration of topic signatures with ontologies and its evaluaton on an automated text summarization system. </abstract></header><body><section number="1" title="Introduction"><p>This paper describes the automated creation of what we call <i>topic signatures, </i>constructs that can play a central role in automated text summarization and information retrieval. Topic signatures can be used to identify the presence of a complex concept—a concept that consists of several related components in fixed relationships. <i>Restaurant-visit, </i>for example, involves at least the concepts <i>menu, eat, pay, </i>and possibly <i>waiter, </i>and <i>Dragon Boat Festival </i>(in Tai­wan) involves the concepts <i>calamus </i>(a talisman to ward off evil), <i>moxa </i>(something with the power of preventing pestilence and strengthening health), pic­tures of <i>Chung Kuei </i>(a nemesis of evil spirits), <i>eggs </i>standing on end, etc. Only when the concepts co-occur is one licensed to infer the complex concept; <i>eat </i>or <i>moxa </i>alone, for example, are not sufficient. At this time, we do not consider the interrelationships among the concepts.</p><p>Since many texts may describe all the compo­nents of a complex concept without ever explic­itly mentioning the underlying complex concept—a topic—itself, systems that have to identify topic (s), for summarization or information retrieval, require a method of inferring complex concepts from their component words in the text.</p></section><section number="2" title="Related Work"><p>In late 1970's, DeJong (DeJong, 1982) developed a system called FRUMP (Fast Reading Understand­ing and Memory Program) to skim newspaper sto­ries and extract the main details.   FRUMP uses a data structure called <i>sketchy script </i>to organize its world knowledge. Each sketchy script is what FRUMP knows about what can occur in particu­lar situations such as demonstrations, earthquakes, labor strikes, and so on. FRUMP selects a partic­ular sketchy script based on clues to styled events in news articles. In other words, FRUMP selects an empty template<footnote anchor="1"/> whose slots will be filled on the fly-as FRUMP <i>reads </i>a news article. A summary is gen­erated based on what has been captured or filled in the template.</p><p>The recent success of information extraction re­search has encouraged the FRUMP approach. The SUMMONS (SUMMarizing Online NewS articles) system (McKeown and Radev, 1999) takes tem­plate outputs of information extraction systems de­veloped for MUC conference and generating sum­maries of multiple news articles. FRUMP and SUM­MONS both rely on prior knowledge of their do­mains. However, to acquire such prior knowledge is labor-intensive and time-consuming. For exam­ple, the University of Massachusetts CIRCUS sys­tem used in the MUC-3 (SAIC, 1998) terrorism do­main required about 1500 person-hours to define ex­traction patterns<footnote anchor="2"/> (Riloff, 1996). In order to make them practical, we need to reduce the knowledge en­gineering bottleneck and improve the portability of FRUMP or SUMMONS-like systems.</p><p>Since the world contains thousands, or perhaps millions, of complex concepts, it is important to be able to learn sketchy scripts or extraction patterns automatically from corpora—no existing knowledge base contains nearly enough information. (Riloff and Lorenzen, 1999) present a system AutoSlog-TS that generates extraction patterns and learns lexical con­straints automatically from preclassified text to al­leviate the knowledge engineering bottleneck men­tioned above. Although Riloff applied AutoSlog-TS to text categorization and information extraction, the concept of <i>relevancy signatures </i>introduced <b>by </b>her is very similar to the <i>topic signatures </i>we pro­posed in this paper.<page local="2"/> Relevancy signatures and topic signatures are both trained on preclassified docu­ments of specific topics and used to identify the presence of the learned topics in previously unseen documents. The main differences to our approach are: relevancy signatures require a parser. They are sentence-based and applied to text categorization. On the contrary, topic signatures only rely on cor­pus statistics, are document-based<footnote anchor="3"/> and used in text summarization.</p><footnote label="1">We viewed sketchy scripts and templates as equivalent constructs in the sense that they specify high level entities and relationships for specific topics.</footnote><footnote label="2">An extraction pattern is essentially a case frame contains its trigger word, enabling conditions, variable slots, and slot constraints. CIRCUS uses a database of extraction patterns to parse texts (Riloff, 1996).</footnote><p>In the next section, we describe the automated text summarization system SUMMARIST that we used in the experiments to provide the context of discussion. We then define topic signatures and de­tail the procedures for automatically constructing topic signatures. In Section 5, we give an overview of the corpus used in the evaluation. In Section 6 we present the experimental results and the possibility of enriching topic signatures using an existing ontol­ogy. Finally, we end this paper with a conclusion.</p></section><section number="3" title="SUMMARIST"><p>SUMMARIST (Hovy and Lin, 1999) is a system designed to generate summaries of multilingual in­put texts. At this time, SUMMARIST can process English, Arabic, Bahasa Indonesia, Japanese, Ko­rean, and Spanish texts. It combines robust natural language processing methods (morphological trans­formation and part-of-speech tagging), symbolic world knowledge, and information retrieval tech­niques (term distribution and frequency) to achieve high robustness and better concept-level generaliza­tion.</p><p>The core of SUMMARIST is based on the follow­ing 'equation':</p><p>summarization = topic identification + topic interpretation + generation.</p><p>These three stages are:</p><p><b>Topic Identification: </b>Identify the most important (central) topics of the texts. SUMMARIST uses positional importance, topic signature, and term frequency. Importance based on discourse structure will be added later. This is the most developed stage in SUMMARIST.</p><p><b>Topic Interpretation: </b>To fuse concepts such as waiter, menu, and food into one generalized concept restaurant, we need more than the sim­ple word aggregation used in traditional infor­mation retrieval. We have investigated concept <u>I    </u>ABCMEWS.com<b><u> :  Delay in Handing Flight 990 Probe to FBI</u></b></p><footnote label="3">We would like to use only the relevant parts of documents to generate topic signatures in the future. Text segmentation algorithms such as TextTiling (Hearst, 1997) can be used to find subtopic segments in text.</footnote><doubt alpha="83.6" length="67" tooSmall="True" monospace="0.0">NT SB Chairman James Hall Egyptian official! want to review résulta</doubt><doubt alpha="79.7" length="74" tooSmall="True" monospace="0.0">of the investigation into the crash of EgyptAir Flight 990 before the case</doubt><doubt alpha="74.1" length="27" tooSmall="True" monospace="0.0">is turned over to the FBI._</doubt><doubt alpha="78.1" length="493" tooSmall="True" monospace="0.0">Nov. 16 - U.S. investigators appear to be leaning more than ever toward the possibility that one of the co-pilots of EgyptAir Flight 990 may havedeliberately crashed the plane last month, killing all 21? people on board.However, U.S. officials say the National Transportation Safety Board will delay transferring the investigation of the Oct. 31 crash to the FBI - the agency that would lead a criminal probe - for at least a few days, to allow Egyptian experts to review evidence in the case.</doubt><doubt alpha="82.8" length="320" tooSmall="True" monospace="0.0">Suspicions of foul play were raised after investigators listening to a tape from the cockpit voice recorder isolated a religious prayer or statement made by the co-pilot just before the plane's autopilot was turned off and the plane began its initial plunge into the Atlantic Ocean off Mas­sachusetts' Nantucket Island._</doubt><doubt alpha="78.9" length="171" tooSmall="True" monospace="0.0">Over the past week, after much effort, the NT SB and the Navy succeeded in locating the plane's two "black boxes," the cockpit voice recorder and the flight data recorder.</doubt><doubt alpha="78.6" length="238" tooSmall="True" monospace="0.0">The tape indicates that shortly after the plane leveled off at its cruising altitude of 33,000 feet, the chief pilot of the aircraft left the plane's cockpit, leaving one of the two co-pilots alone there as the aircraft began its descent.</doubt><p>Figure <b>1: </b>A Nov. 16 1999 ABC News page summary-generated by SUMMARIST.</p><p>counting and topic signatures to tackle the fu­sion problem.</p><p><b>Summary Generation: </b>SUMMARIST can pro­duce keyword and extract type summaries.</p><p>Figure 1 shows an ABC News page summary about EgyptAir Flight 990 by SUMMARIST. SUM­MARIST employs several different heuristics in the topic identification stage to score terms and sen­tences. The score of a sentence is simply the sum of all the scores of content-bearing terms in the sen­tence. These heuristics are implemented in separate modules using inputs from preprocessing modules such as tokenizer, part-of-speech tagger, morpholog­ical analyzer, term frequency and <i>tfidf </i>weights cal­culator, sentence length calculator, and sentence lo­cation identifier. We only activate the position mod­ule, the <i>tfidf </i>module, and the topic signature module for comparison. We discuss the effectiveness of these modules in Section 6.</p></section><section number="4" title="Topic Signatures"><p>Before addressing the problem of world knowledge acquisition head-on, we decided to investigate what type of knowledge would be useful for summariza­tion. After all, one can spend a lifetime acquir­ing knowledge in just a small domain. But what is the minimum amount of knowledge we need to enable effective topic identification as illustrated by the <i>restaurant-visit </i>example? Our idea is simple. We would collect a set of terms<footnote anchor="4"/> that were typi­cally highly correlated with a target concept from a preclassified corpus such as TREC collections, and then, during summarization, group the occurrence of the related terms by the target concept. For exam­ple, we would replace joint instances of <i>table, menu, waiter, order, eat, pay, tip, </i>and so on, by the single phrase <i>restaurant-visit, </i>in producing an indicative summary.<page local="3"/> We thus defined a topic signature as a family of related terms, as follows:</p><footnote label="4">Terms can be stemmed words, bigrams, or trigrams.</footnote><doubt alpha="59.3" length="27" tooSmall="False" monospace="0.0">TS   =   {topic, signature}</doubt><doubt alpha="30.0" length="30" tooSmall="False" monospace="0.0">=   {topic, &lt; .,(tn, wn)&gt;} (1)</doubt><p>where <i>topic </i>is the target concept and <i>signature </i>is a vector of related terms. Each <i>ti </i>is an term highly-correlated to <i>topic </i>with association weight w, . The number of related terms <i>n </i>can be set empirically-according to a cutoff associated weight. We describe how to acquire related terms and their associated weights in the next section.</p><subsection number="4.1" title="Signature Term Extraction and Weight Estimation"><p>On the assumption that semantically related terms tend to co-occur, one can construct topic signa­tures from preclassified text using the x<footnote anchor="2"/> test, mu­tual information, or other standard statistic tests and information-theoretic measures. Instead of <i>x<footnote anchor="2"/>-, </i>we use <i>likelihood ratio </i>(Dunning, 1993) A, since À is more appropriate for sparse data than x<footnote anchor="2"/> test and the quantity <i>—2logX </i>is asymptotically x<footnote anchor="2"/> dis­tributed<footnote anchor="5"/>. Therefore, we can determine the confi­dence level for a specific <i>—2logX </i>value by looking up X<footnote anchor="2"/> distribution table and use the value to select an appropriate cutoff associated weight.</p><p>We have documents preclassified into a set <i>TZ </i>of relevant texts and a set <i>TZ </i>of nonrelevant texts for a given topic. Assuming the following two hypotheses:</p><p><b>Hypothesis 1 </b><i>(Hi): P</i><i>(1Z</i><i>\ti) = p = P(TZ\ii), </i>i.e. the relevancy of a document is independent of <i>ti.</i></p><p><b>Hypothesis 2 </b><i>(H2): P</i><i>(1Z</i><i>\ti) = px # p2 = P(TZ\ii), </i>i.e. the presence of <i>ti </i>indicates strong relevancy assuming <i>p\</i><i> </i><i>^&gt;</i><i> </i><i>p2.</i></p><p>and the following 2-by-2 contingency table:</p><p>where On is the frequency of term <i>ti </i>occurring in the relevant set, <i>Oi2 </i>is the frequency of term <i>ti </i>oc­curring in the nonrelevant set, <i>0</i><i>2i </i>is the frequency of term <i>ti ^ ti </i>occurring in the relevant set, 022 is the frequency of term <i>ti ^ ti </i>occurring in the non-relevant set.</p><p>Assuming a binomial distribution:</p><doubt alpha="39.3" length="28" tooSmall="False" monospace="0.0">b(k;n,x)= (^jxk(l^x)(n-k)(2)</doubt><footnote label="5">This assumes that the ratio is between the maximum like­lihood estimate over a subpart of the parameter space and the maximum likelihood estimate over the entire parameter space. See (Manning and Schütze, 1999) pages 172 to 175 for details.</footnote><p>then the likelihood for <i>Hi </i>is:</p><doubt alpha="38.6" length="57" tooSmall="False" monospace="0.0">L(#i) = 6(0ii;0ii + 0i2,p)fc(02i;02i + 022,p)and forH2is:</doubt><doubt alpha="30.4" length="46" tooSmall="False" monospace="0.0">L(H2) =fe(0ii;0ii + 012,Pl)fe(02i;021 +022,P2)</doubt><p>The <i>—2logX </i>value is then computed as follows:</p><doubt alpha="40.0" length="5" tooSmall="True" monospace="0.0">L(H1)</doubt><doubt alpha="50.0" length="8" tooSmall="True" monospace="0.0">= -2lofl</doubt><doubt alpha="40.0" length="5" tooSmall="True" monospace="0.0">L(H2)</doubt><doubt alpha="33.3" length="36" tooSmall="False" monospace="0.0">b(On; On+Oi2, p)b(021 ;O21 + O22' p)</doubt><doubt alpha="21.4" length="14" tooSmall="True" monospace="0.0">=       —2iog-</doubt><doubt alpha="27.4" length="179" tooSmall="False" monospace="0.0">b(On: On +O12. Pl)b(02i; °21 + 022'P2) =      -2((On+ O2l)loflp +(O12+ O22)lofl (1 - P) - (3)(OnlogPI +Oi2log(1 — p1) + 021*ogP2 + 022        (1 — P2))) =      2JV X - «(R|T)) (4)</doubt><doubt alpha="27.3" length="22" tooSmall="True" monospace="0.0">=      2JV XX{H;T) (5)</doubt><p>where <i>N </i>= On + 0i2 + 02i + 022 is the total num­ber of term occurrence in the corpus, <i>/H(Jl) </i>is the entropy of terms over relevant and nonrelevant sets of documents, <i>%(TZ\T) </i>is the entropy of a given term over relevant and nonrelevant sets of documents, and <i>1(TZ; </i><i>T)</i><i> </i>is the mutual information between docu­ment relevancy and a given term. Equation 5 indi­cates that mutual information<footnote anchor="6"/> is an equivalent mea­sure to likelihood ratio when we assume a binomial distribution and a 2-by-2 contingency table. To create topic signature for a given topic, we:</p><p>1. classify documents as relevant or nonrelevant according to the given topic</p><p>2. compute the <i>—2logX </i>value using Equation 3 for each term in the document collection</p></subsection></section><section number="3." title="rank terms according to their —2logX value"><p>4. select a confidence level from the x<footnote anchor="2"/> distribution table; determine the cutoff associated weight and the number of terms to be included in the signatures</p></section><section number="5" title="The Corpus"><p>The training data derives from the Question and Answering summary evaluation data provided <b>by </b>TIPSTER-SUMMAC <b>(Ma</b>ni et al., 1998) that is a subset of the TREC collections. The TREC data is a collection of texts, classified into various topics, used for formal evaluations of information retrieval sys­tems in a series of annual comparisons. This data set contains essential text fragments (phrases, clauses, and sentences) which must be included in summaries to answer some TREC topics. These fragments are each judged by a human judge. As described in Sec­tion 3, SUMMARIST employs several independent modules to assign a score to each sentence, and then combines the scores to decide which sentences to ex­tract from the input text. One can gauge the efficacy<page local="4"/></p><footnote label="6">The mutual information is defined according to chapter 2 of (Cover and Thomas, 1991) and is not the pairwise mutual information used in (Church and Hanks, 1990).</footnote><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p></p></td><td class="cell"><p><i>TZ</i></p></td><td class="cell"><p><i>TZ</i></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>ti</i></p></td><td class="cell"><p><i>On</i></p></td><td class="cell"><p>012</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><i>ti</i></p></td><td class="cell"><p>021</p></td><td class="cell"><p>022</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><p><b><u>TREC Topic Description</u>_</b> <b><u>Test Questions</u></b></p><doubt alpha="52.9" length="17" tooSmall="True" monospace="0.0">&lt;num) Number: 151</doubt><doubt alpha="80.3" length="66" tooSmall="True" monospace="0.0">(title) Topic: Coping with overcrowded prisons (desc) Description:</doubt><doubt alpha="82.6" length="207" tooSmall="True" monospace="0.0">The document will provide information on jail and prison overcrowding and how inmates are forced to cope with those conditions; or it will reveal plans to relieve the overcrowded condition, (narr) Narrative:</doubt><doubt alpha="82.4" length="324" tooSmall="True" monospace="0.0">A relevant document will describe scenes of overcrowding that have become all too common in jails and prisons around the country. The document will identify how inmates are forced to cope with those over­crowded conditions, and/or what the Correctional System is doing, or planning to do, to alleviate the crowded condition.</doubt><doubt alpha="42.9" length="7" tooSmall="True" monospace="0.0">&lt;/top)_</doubt><doubt alpha="52.3" length="172" tooSmall="False" monospace="0.0">(DOCNO)AP891027-0063(/DOCNO)(FILEID) AP-NR-10-27-89 0615EDT(/FILEID) (lST_LINE)r a PM-Chainedlnmates 10-27 0335 ( /1 ST.LINE) (2ND_LINE)PM-Chained In mates ,0344(/2ND_LINE)</doubt><doubt alpha="81.8" length="66" tooSmall="True" monospace="0.0">(HEAD)Inmates Chained to Walls in Baltimore Police Stations(/HEAD)</doubt><doubt alpha="72.1" length="43" tooSmall="True" monospace="0.0">(DATELINE)BALTIMORE (AP) (/DATELINE) (TEXT)</doubt><doubt alpha="78.0" length="168" tooSmall="True" monospace="0.0">(Q3)Prisoners are kept chained to the walls of local police lockups for as long as three days at a time because of overcrowding in regular jail cells, police said.(/Q3)</doubt><doubt alpha="74.7" length="91" tooSmall="True" monospace="0.0">Overcrowding at the ( Ql ) Bait imore County Detention Center(/Ql) has forced police to ...</doubt><doubt alpha="50.0" length="8" tooSmall="True" monospace="0.0">(/TEXT)_</doubt><p>Table 1: TREC topic description for topic 151, test questions expected to be answered by relevant doc­uments, and a sample document with answer keys.</p><p>of each module by comparing, for different amounts of extraction, how many 'good' sentences the module selects by itself. We rate a sentence as good simply if it also occurs in the ideal human-made extract, and measure it using combined recall and precision (F-score). We used four topics<footnote anchor="7"/> of total 6,194 doc­uments from the TREC collection. 138 of them are relevant documents with TIPSTER-SUMMAC pro­vided answer keys for the question and answering evaluation. Model extracts are created automati­cally from sentences containing answer keys. Table 1 shows TREC topic description for topic 151, test questions expected to be answered by relevant doc­uments<footnote anchor="8"/>, and a sample relevant document with an­swer keys markup.</p><footnote label="7">These four topics are: topic 151: Overcrowded Prisons, 1211 texts, 85 relevant;</footnote><doubt alpha="66.7" length="57" tooSmall="False" monospace="0.0">topic 257:Cigarette Consumption,1727 texts, 126 relevant;</doubt><doubt alpha="65.4" length="52" tooSmall="False" monospace="0.0">topic 258:Computer Security,1701 texts, 49 relevant;</doubt><p>topic 271: <i>Solar Power, </i>1555 texts, 59 relevant.</p><footnote label="8">A relevant document only needs to answer at least one of the five questions.</footnote></section><section number="6" title="Experimental Results"><p>In order to assess the utility of topic signatures in text summarization, we follow the procedure de­scribed at the end of Section 4.1 to create topic signature for each selected TREC topic. Docu­ments are separated into relevant and nonrelevant sets according to their TREC relevancy judgments for each topic. We then run each document through a part-of-speech tagger and convert each word into its root form based on the WordNet lexical database. We also collect individual root word (unigram) fre­quency, two consecutive non-stopword<footnote anchor="9"/> (bigram) fre­quency, and three consecutive non-stopwords (tri-gram) frequency to facilitate the computation of the <i>—2logX </i>value for each term. We expect high rank­ing bigram and trigram signature terms to be very-informative. We set the cutoff associated weight at 10.83 with confidence level <i>a </i>= 0.001 by looking up a x<footnote anchor="2"/> statistical table.</p><p>Table 2 shows the top 10 unigram, bigram, and tri­gram topic signature terms for each topic <footnote anchor="10"/>. Several conclusions can be drawn directly. Terms with high <i>—2logX </i>are indeed good indicators for their corre­sponding topics. The <i>—2logX </i>values decrease as the number of words in a term increases. This is rea­sonable, since longer terms usually occur less often than their constituents. However, bigram terms are more informative than unigram terms as we can ob­serve: <i>jail/prison overcrowding </i>of topic 151, <i>tobacco industry </i>of topic 257, <i>computer security </i>of topic 258, and <i>solar energy/power </i>of topic 271. These auto­matically generated signature terms closely resemble or equal the given short TREC topic descriptions. Although trigram terms shown in the table, such as <i>federal court order, philip morris rjr, jet propul­sion laboratory, </i>and <i>mobile telephone system </i>are also meaningful, they do not demonstrate the closer term relationship among other terms in their respective topics that is seen in the bigram cases. We expect that more training data can improve the situation.</p><p>We notice that the <i>—2logX </i>values for topic 258 are higher than those of the other three topics. As indicated by (Mani et al., 1998) the majority of rel­evant documents for topic 258 have the query topic as their main theme; while the others mostly have the query topics as their subsidiary themes. This implies that it is too liberal to assume all the terms in relevant documents of the other three topics are relevant. We plan to apply text segmentation algo­rithms such as TextTiling (Hearst, 1997) to segment documents into subtopic units. We will then per­form the topic signature creation procedure only on the relevant units to prevent inclusion of noise terms.</p><footnote label="9">We use the stopword list supplied with the SMART re­trieval system.</footnote><footnote label="10">The —2log\ values are not comparable across ngram cat­egories, since each ngram category has its own sample space.</footnote><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Ql</b></p></td><td class="cell"><p><b>What   are   name   and/or   location   of  the   correction facilities where the reported overcrowding exists?</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Q2</b></p></td><td class="cell"><p><b>What negative experiences have there been at the overcrowded facilities (whether or not they are thought to have been caused by the overcrowding)?</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Q3</b></p></td><td class="cell"><p><b>What measures have been taken/plan ned/recom mended (etc.) to accommodate more inmates at penal facilities, e.g., doubling</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Q4</b></p></td><td class="cell"><p><b>What measures have been taken/plan ned/recom mended (etc.)</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Q5</b></p></td><td class="cell"><p><b>What measures have been taken/plan ned/recom mended (etc.) to reduce the number of existing inmates at  an overcrowded facility, e.g., granting early release, transfering to uncrowded facilities?</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Sample Answer Keys</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="5"/><p>The topic signature module scans each sentence, assigning to each word that occurs in a topic signa­ture the weight of that keyword in the topic signa­ture. Each sentence then receives a topic signature score equal to the total of all signature word scores it contains, normalized by the highest sentence score. This score indicates the relevance of the sentence to the signature topic.</p><p>SUMMARIST produced extracts of the same texts separately for each module, for a series of ex­tracts ranging from 0% to 100% of the original text.</p><p>Although many relevant documents are available for each topic, only some of them have answer key</p><subsection number="6.1" title="Comparing Summary Extraction"><p><b>Effectiveness Using Topic Signatures, </b><i>TFIDF, </i><b>and Baseline Algorithms</b></p><p>In order to evaluate the effectiveness of topic signa­tures used in summary extraction, we compare the summary sentences extracted by the topic signature module, baseline module, and <i>tfidf </i>modules with hu­man annotated model summaries. We measure the performance using a combined measure of recall <i>(R) </i>and precision <i>(P),</i><i> </i><i>F.</i><i> </i>F-score is defined by:</p><doubt alpha="33.3" length="3" tooSmall="False" monospace="0.0">F =</doubt><doubt alpha="50.0" length="14" tooSmall="False" monospace="0.0">(l+ß2)PRß2P+R'</doubt><doubt alpha="100.0" length="5" tooSmall="False" monospace="0.0">where</doubt><doubt alpha="50.0" length="2" tooSmall="False" monospace="0.0">R=</doubt><doubt alpha="50.0" length="2" tooSmall="False" monospace="0.0">N„</doubt><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">P</doubt><doubt alpha="100.0" length="3" tooSmall="False" monospace="0.0">Nme</doubt><doubt alpha="100.0" length="2" tooSmall="False" monospace="0.0">Nm</doubt><doubt alpha="100.0" length="1" tooSmall="False" monospace="0.0">ß</doubt><doubt alpha="100.0" length="5" tooSmall="False" monospace="0.0">NmeNe</doubt><doubt alpha="0.0" length="7" tooSmall="False" monospace="0.0">(6) (7)</doubt><p><b><i>#</i></b><b><i> </i></b><b><i>of</i></b><b><i> </i></b><i>sentences </i><b><i>extrated </i></b><i>that also appear in </i><b><i>the </i></b><i>model summary</i> <b><i>#</i></b><b><i> </i></b><b><i>of</i></b><b><i> </i></b><i>sentences in </i><b><i>the </i></b><i>model summary</i> <b><i>#</i></b><b><i> </i></b><b><i>of</i></b><b><i> </i></b><i>sentences extracted by </i><b><i>the </i></b><i>system relative importance </i><i>of</i><i> R and P</i></p><p>We assume equal importance of recall and preci­sion and set <i>ß </i>to 1 in our experiments. The baseline (position) module scores each sentence by its posi­tion in the text. The first sentence gets the high­est score, the last sentence the lowest. The baseline method is expected to be effective for news genre. The <i>tfidf </i>module assigns a score to a term <i>ti </i>accord­ing to the product of its frequency within a doc­ument <i>j</i><i> </i><i>(tfij)</i><i> </i>and its inverse document frequency  <i>(idfj</i><i> = log-^-).</i><i> N </i>is the total number of documents in the corpus and <i>dfj</i><i> </i>is the number of documents containing term <i>ti.</i></p><table caption="Table 2: Top 10 signature terms of unigram, bigram, and trigram for four TREC topics." class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Topic 10 Signature Terms of Topic 151 -         Overcrowded Prisons</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Unigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"><p><b>Bigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"><p><b>Trigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>jail</b></p></td><td class="cell"><p><b>461.044</b></p></td><td class="cell"><p><b>county jail</b></p></td><td class="cell"><p><b>160.273</b></p></td><td class="cell"><p><b>federal court order</b></p></td><td class="cell"><p><b>45.960</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>county</b></p></td><td class="cell"><p><b>408.821</b></p></td><td class="cell"><p><b>early release</b></p></td><td class="cell"><p><b>85.361</b></p></td><td class="cell"><p><b>comply consent decree</b></p></td><td class="cell"><p><b>35.121</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>overcrowding</b></p></td><td class="cell"><p><b>342.349</b></p></td><td class="cell"><p><b>state prison</b></p></td><td class="cell"><p><b>74.372</b></p></td><td class="cell"><p><b>dekalb county sheriff</b></p></td><td class="cell"><p><b>35.121</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>inmate</b></p></td><td class="cell"><p><b>234.765</b></p></td><td class="cell"><p><b>state prisoner</b></p></td><td class="cell"><p><b>67.666</b></p></td><td class="cell"><p><b>gov. jo frank</b></p></td><td class="cell"><p><b>35.121</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>sheriff</b></p></td><td class="cell"><p><b>154.440</b></p></td><td class="cell"><p><b>day fine</b></p></td><td class="cell"><p><b>61.465</b></p></td><td class="cell"><p><b>joe frank harris</b></p></td><td class="cell"><p><b>35.121</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>state</b></p></td><td class="cell"><p><b>151.940</b></p></td><td class="cell"><p><b>jail overcrowding</b></p></td><td class="cell"><p><b>61.329</b></p></td><td class="cell"><p><b>prisoner county jail</b></p></td><td class="cell"><p><b>35.121</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>prisoner</b></p></td><td class="cell"><p><b>148.178</b></p></td><td class="cell"><p><b>court order</b></p></td><td class="cell"><p><b>60.090</b></p></td><td class="cell"><p><b>state prison county</b></p></td><td class="cell"><p><b>28.043</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>prison</b></p></td><td class="cell"><p><b>145.306</b></p></td><td class="cell"><p><b>local jail</b></p></td><td class="cell"><p><b>56.440</b></p></td><td class="cell"><p><b>'t put prison</b></p></td><td class="cell"><p><b>26.341</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>city</b></p></td><td class="cell"><p><b>133.477</b></p></td><td class="cell"><p><b>prison overcrowding</b></p></td><td class="cell"><p><b>55.373</b></p></td><td class="cell"><p><b>county jail state</b></p></td><td class="cell"><p><b>26.341</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>overcrowded</b></p></td><td class="cell"><p><b>128.008</b></p></td><td class="cell"><p><b>central facility</b></p></td><td class="cell"><p><b>52.909</b></p></td><td class="cell"><p><b>hold local jail</b></p></td><td class="cell"><p><b>26.341</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Topic 10 Signature Terms of Topic 257 -        Cigarette Consumption</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Unigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"><p><b>Bigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"><p><b>Trigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>cigarette</b></p></td><td class="cell"><p><b>476.038</b></p></td><td class="cell"><p><b>tobacco industry</b></p></td><td class="cell"><p><b>80.768</b></p></td><td class="cell"><p><b>philip morris rjr</b></p></td><td class="cell"><p><b>28.061</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tobacco</b></p></td><td class="cell"><p><b>313.017</b></p></td><td class="cell"><p><b>bn cigarette</b></p></td><td class="cell"><p><b>67.429</b></p></td><td class="cell"><p><b>rothmans benson hedge</b></p></td><td class="cell"><p><b>26.969</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>smoking</b></p></td><td class="cell"><p><b>284.198</b></p></td><td class="cell"><p><b>philip morris</b></p></td><td class="cell"><p><b>54.073</b></p></td><td class="cell"><p><b>lung cancer death</b></p></td><td class="cell"><p><b>22.214</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>smoke</b></p></td><td class="cell"><p><b>159.134</b></p></td><td class="cell"><p><b>cigarette year</b></p></td><td class="cell"><p><b>48.045</b></p></td><td class="cell"><p><b>qtr firm chg</b></p></td><td class="cell"><p><b>21.418</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>rothmans</b></p></td><td class="cell"><p><b>156.675</b></p></td><td class="cell"><p><b>rothmans international</b></p></td><td class="cell"><p><b>44.434</b></p></td><td class="cell"><p><b>qtr qtr firm</b></p></td><td class="cell"><p><b>21.418</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>osha</b></p></td><td class="cell"><p><b>148.372</b></p></td><td class="cell"><p><b>tobacco smoke</b></p></td><td class="cell"><p><b>44.269</b></p></td><td class="cell"><p><b>bn bn bn</b></p></td><td class="cell"><p><b>20.226</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>seita</b></p></td><td class="cell"><p><b>126.421</b></p></td><td class="cell"><p><b>sir patrick</b></p></td><td class="cell"><p><b>40.455</b></p></td><td class="cell"><p><b>consumption bn cigarette</b></p></td><td class="cell"><p><b>20.226</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>ban</b></p></td><td class="cell"><p><b>113.849</b></p></td><td class="cell"><p><b>cigarette company</b></p></td><td class="cell"><p><b>39.399</b></p></td><td class="cell"><p><b>great american smokeout</b></p></td><td class="cell"><p><b>20.226</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>smoker</b></p></td><td class="cell"><p><b>104.110</b></p></td><td class="cell"><p><b>cent market</b></p></td><td class="cell"><p><b>36.223</b></p></td><td class="cell"><p><b>lung cancer heart</b></p></td><td class="cell"><p><b>20.226</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>bat</b></p></td><td class="cell"><p><b>79.903</b></p></td><td class="cell"><p><b>tax increase</b></p></td><td class="cell"><p><b>36.223</b></p></td><td class="cell"><p><b>malaysian Singapore company</b></p></td><td class="cell"><p><b>20.226</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Topic 10 Signature Terms of Topic 258 -         Computer Security</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Unigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"><p><b>Bigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"><p><b>Trigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>computer</b></p></td><td class="cell"><p><b>1159.351</b></p></td><td class="cell"><p><b>computer security</b></p></td><td class="cell"><p><b>213.331</b></p></td><td class="cell"><p><b>jet propulsion laboratory</b></p></td><td class="cell"><p><b>98.854</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>virus</b></p></td><td class="cell"><p><b>927.674</b></p></td><td class="cell"><p><b>graduate student</b></p></td><td class="cell"><p><b>178.588</b></p></td><td class="cell"><p><b>robert t. mo</b></p></td><td class="cell"><p><b>98.854</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>hacker</b></p></td><td class="cell"><p><b>887.377</b></p></td><td class="cell"><p><b>computer system</b></p></td><td class="cell"><p><b>146.328</b></p></td><td class="cell"><p><b>Cornell university graduate</b></p></td><td class="cell"><p><b>79.081</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>morris</b></p></td><td class="cell"><p><b>666.392</b></p></td><td class="cell"><p><b>research center</b></p></td><td class="cell"><p><b>132.413</b></p></td><td class="cell"><p><b>lawrence berkeley laboratory</b></p></td><td class="cell"><p><b>79.081</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Cornell</b></p></td><td class="cell"><p><b>385.684</b></p></td><td class="cell"><p><b>computer virus</b></p></td><td class="cell"><p><b>126.033</b></p></td><td class="cell"><p><b>nasa jet propulsion</b></p></td><td class="cell"><p><b>79.081</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>university</b></p></td><td class="cell"><p><b>305.958</b></p></td><td class="cell"><p><b>Cornell university</b></p></td><td class="cell"><p><b>108.741</b></p></td><td class="cell"><p><b>university graduate student</b></p></td><td class="cell"><p><b>79.081</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>system</b></p></td><td class="cell"><p><b>290.347</b></p></td><td class="cell"><p><b>nuclear weapon</b></p></td><td class="cell"><p><b>107.283</b></p></td><td class="cell"><p><b>lawrence livermore national</b></p></td><td class="cell"><p><b>69.195</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>laboratory</b></p></td><td class="cell"><p><b>287.521</b></p></td><td class="cell"><p><b>military computer</b></p></td><td class="cell"><p><b>106.522</b></p></td><td class="cell"><p><b>livermore national laboratory</b></p></td><td class="cell"><p><b>69.195</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>lab</b></p></td><td class="cell"><p><b>225.516</b></p></td><td class="cell"><p><b>virus program</b></p></td><td class="cell"><p><b>106.522</b></p></td><td class="cell"><p><b>computer security expert</b></p></td><td class="cell"><p><b>66.196</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>mcclary</b></p></td><td class="cell"><p><b>128.515</b></p></td><td class="cell"><p><b>west german</b></p></td><td class="cell"><p><b>82.210</b></p></td><td class="cell"><p><b>security center bethesda</b></p></td><td class="cell"><p><b>49.423</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Topic 10 Signature Terms of Topic 271 -         Solar Power</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Unigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"><p><b>Bigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"><p><b>Trigram</b></p></td><td class="cell"><p><b>-2Io</b><b>flA</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>solar</b></p></td><td class="cell"><p><b>484.315</b></p></td><td class="cell"><p><b>solar energy</b></p></td><td class="cell"><p><b>268.521</b></p></td><td class="cell"><p><b>division multiple access</b></p></td><td class="cell"><p><b>31.347</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>mazda</b></p></td><td class="cell"><p><b>308.015</b></p></td><td class="cell"><p><b>solar power</b></p></td><td class="cell"><p><b>94.210</b></p></td><td class="cell"><p><b>mobile telephone service</b></p></td><td class="cell"><p><b>31.347</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>leo</b></p></td><td class="cell"><p><b>276.932</b></p></td><td class="cell"><p><b>christian aid</b></p></td><td class="cell"><p><b>86.211</b></p></td><td class="cell"><p><b>british technology group</b></p></td><td class="cell"><p><b>23.510</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>iridium</b></p></td><td class="cell"><p><b>258.705</b></p></td><td class="cell"><p><b>leo system</b></p></td><td class="cell"><p><b>70.535</b></p></td><td class="cell"><p><b>earth height mile</b></p></td><td class="cell"><p><b>23.510</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>pavilion</b></p></td><td class="cell"><p><b>203.811</b></p></td><td class="cell"><p><b>mobile telephone</b></p></td><td class="cell"><p><b>70.535</b></p></td><td class="cell"><p><b>financial backing iridium</b></p></td><td class="cell"><p><b>23.510</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>pound</b></p></td><td class="cell"><p><b>128.121</b></p></td><td class="cell"><p><b>iridium project</b></p></td><td class="cell"><p><b>62.697</b></p></td><td class="cell"><p><b>global mobile satellite</b></p></td><td class="cell"><p><b>23.510</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>tower</b></p></td><td class="cell"><p><b>126.353</b></p></td><td class="cell"><p><b>real goods</b></p></td><td class="cell"><p><b>61.901</b></p></td><td class="cell"><p><b>handheld mobile telephone</b></p></td><td class="cell"><p><b>23.510</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>lookout</b></p></td><td class="cell"><p><b>125.406</b></p></td><td class="cell"><p><b>science park</b></p></td><td class="cell"><p><b>54.859</b></p></td><td class="cell"><p><b>mobile satellite system</b></p></td><td class="cell"><p><b>23.510</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>Inmarsat</b></p></td><td class="cell"><p><b>109.728</b></p></td><td class="cell"><p><b>solar concentrator</b></p></td><td class="cell"><p><b>54.859</b></p></td><td class="cell"><p><b>motorola iridium project</b></p></td><td class="cell"><p><b>23.510</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>boydston</b></p></td><td class="cell"><p><b>78.373</b></p></td><td class="cell"><p><b>bp solar</b></p></td><td class="cell"><p><b>31.347</b></p></td><td class="cell"><p><b>active solar system</b></p></td><td class="cell"><p><b>15.673</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table><page local="6"/><p>markups. The number of documents with answer keys are listed in the row labeled: "# of Relevant Docs Used in Training". To ensure we utilize all the available data and conduct a sound evaluation, we perform a three-fold cross validation. We re­serve one-third of documents as test set, use the rest as training set, and repeat three times with non-overlapping test set. Furthermore, we use only uni­gram topic signatures for evaluation.</p><p>The result is shown in Figure 2 and Table 3. We find that the topic signature method outperforms the other two methods and the <b><i>tfidf </i></b>method performs poorly. Among 40 possible test points for four topics with 10% summary length increment (0% means se­lect at least one sentence) as shown in Table 3, the topic signature method beats the baseline method 34 times. This result is really encouraging and in­dicates that the topic signature method is a worthy-addition to a variety of text summarization methods.</p></subsection><subsection number="6.2" title="Enriching Topic Signatures Using Existing Ontologies"><p>We have shown in the previous sections that topic signatures can be used to approximate topic iden­tification at the lexical level. Although the au­tomatically acquired signature terms for a specific topic seem to be bound by unknown relationships as shown in Table 2, it is hard to image how we can enrich the inherent flat structure of topic signatures as defined in Equation 1 to a construct as complex as a MUC template or script.</p><p>As discussed in (Agirre et al., 2000), we propose using an existing ontology such as SENSUS (Knight and Luk, 1994) to identify signature term relations. The external hierarchical framework can be used to generalize topic signatures and suggest richer rep­resentations for topic signatures. Automated entity-recognizers can be used to classify unknown enti­ties into their appropriate SENSUS concept nodes. We are also investigating other approaches to auto­matically learn signature term relations. The idea mentioned in this paper is just a starting point.</p></subsection></section><section number="7" title="Conclusion"><p>In this paper we presented a procedure to automati­cally acquire topic signatures and valuated the effec­tiveness of applying topic signatures to extract topic relevant sentences against two other methods. The topic signature method outperforms the baseline and the <i>tfidf </i>methods for all test topics. Topic signatures can not only recognize related terms (topic identifi­cation), but group related terms together under one target concept (topic interpretation). Topic identi­fication and interpretation are two essential steps in a typical automated text summarization system as we present in Section 3.</p><p>Topic signatures can also been viewed as an in­verse process of query expansion. Query expansion intends to alleviate the word mismatch problem in information retrieval, since documents are normally-written in different vocabulary. How to automati­cally identify highly correlated terms and use them to improve information retrieval performance has been a main research issue since late <b>1960</b>'s. Re­cent advances in the query expansion (Xu and Croft, 1996) can also shed some light on the creation of topic signatures. Although we focus the use of topic signatures to aid text summarization in this paper, we plan to explore the possibility of applying topic signatures to perform query expansion in the future.</p><p>The results reported are encouraging enough to allow us to continue with topic signatures as the ve­hicle for a first approximation to world knowledge. We are now busy creating a large number of signa­tures to overcome the world knowledge acquisition problem and use them in topic interpretation.</p></section><section number="8" title="Acknowledgements"><p>We thank the anonymous reviewers for very use­ful suggestions. This work is supported in part <b>by </b>DARPA contract N66001-97-9538.</p><table class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>151-tfidf</b></p></td><td class="cell"><p><b>-39.93</b></p></td><td class="cell"><p><b>-35.82</b></p></td><td class="cell"><p><b>-15.26</b></p></td><td class="cell"><p><b>-7.22</b></p></td><td class="cell"><p><b>-6.35</b></p></td><td class="cell"><p><b>-3.50</b></p></td><td class="cell"><p><b>2.55</b></p></td><td class="cell"><p><b>3.76</b></p></td><td class="cell"><p><b>3.83</b></p></td><td class="cell"><p><b>-0.75</b></p></td><td class="cell"><p><b>0.00</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>151_topic_aig</b></p></td><td class="cell"><p><b>-2.76</b></p></td><td class="cell"><p><b>-2.19</b></p></td><td class="cell"><p><b>+6.02</b></p></td><td class="cell"><p><b>+ 4.58</b></p></td><td class="cell"><p><b>+ 7.48</b></p></td><td class="cell"><p><b>+ 13.77</b></p></td><td class="cell"><p><b>+ 15.63</b></p></td><td class="cell"><p><b>+ 14.17</b></p></td><td class="cell"><p><b>+8.66</b></p></td><td class="cell"><p><b>+3.59</b></p></td><td class="cell"><p><b>0.00</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>257_baaeHne    |        0.098   |        0.155   |        0.191   |        0.184   |        0.199   |        0.193   |        0.189   |        0.181   |        0.181   |     0.185   | 0.190</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>257_tfidf</b></p></td><td class="cell"><p><b>-55.11</b></p></td><td class="cell"><p><b>-38.56</b></p></td><td class="cell"><p><b>-20.59</b></p></td><td class="cell"><p><b>+ 6.54</b></p></td><td class="cell"><p><b>+6.06</b></p></td><td class="cell"><p><b>+ 16.44</b></p></td><td class="cell"><p><b>+ 18.34</b></p></td><td class="cell"><p><b>+ 21.68</b></p></td><td class="cell"><p><b>+ 14.49</b></p></td><td class="cell"><p><b>+ 7.09</b></p></td><td class="cell"><p><b>0.00</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>257_topic_aig</b></p></td><td class="cell"><p><b>+ 45.53</b></p></td><td class="cell"><p><b>+64.06</b></p></td><td class="cell"><p><b>+31.86</b></p></td><td class="cell"><p><b>+34.91</b></p></td><td class="cell"><p><b>+ 16.07</b></p></td><td class="cell"><p><b>+ 20.40</b></p></td><td class="cell"><p><b>+ 20.60</b></p></td><td class="cell"><p><b>+ 18.01</b></p></td><td class="cell"><p><b>+ 12.48</b></p></td><td class="cell"><p><b>+ 4.24</b></p></td><td class="cell"><p><b>0.00</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>258-baaeHne    |        0.141   |        0.270   |        0.428   |        0.463   |        0.462   |        0.471   |        0.470   |        0.502   |        0.512   |     0.528   | 0.527</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>258-tfidf</b></p></td><td class="cell"><p><b>-18.84</b></p></td><td class="cell"><p><b>-27.88</b></p></td><td class="cell"><p><b>-16.57</b></p></td><td class="cell"><p><b>-5.21</b></p></td><td class="cell"><p><b>+6.61</b></p></td><td class="cell"><p><b>+ 16.13</b></p></td><td class="cell"><p><b>+ 16.90</b></p></td><td class="cell"><p><b>+9.56</b></p></td><td class="cell"><p><b>+4.74</b></p></td><td class="cell"><p><b>+ 1.92</b></p></td><td class="cell"><p><b>0.00</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>258-topic_aig</b></p></td><td class="cell"><p><b>+ 1.55</b></p></td><td class="cell"><p><b>-21.82</b></p></td><td class="cell"><p><b>-17.19</b></p></td><td class="cell"><p><b>-6.56</b></p></td><td class="cell"><p><b>+8.96</b></p></td><td class="cell"><p><b>+ 14.16</b></p></td><td class="cell"><p><b>+ 20.40</b></p></td><td class="cell"><p><b>+ 11.44</b></p></td><td class="cell"><p><b>+ 7.74</b></p></td><td class="cell"><p><b>+3.43</b></p></td><td class="cell"><p><b>0.00</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>271_baaeHne    |        0.167   |        0.316   |        0.368   |        0.338   |        0.335   |        0.351   |        0.331   |        0.310   |        0.321   |     0.333   | 0.332</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>271_tfidf</b></p></td><td class="cell"><p><b>-56.75</b></p></td><td class="cell"><p><b>-51.35</b></p></td><td class="cell"><p><b>-35.25</b></p></td><td class="cell"><p><b>-13.20</b></p></td><td class="cell"><p><b>-5.99</b></p></td><td class="cell"><p><b>-4.83</b></p></td><td class="cell"><p><b>+5.96</b></p></td><td class="cell"><p><b>+ 10.97</b></p></td><td class="cell"><p><b>+5.76</b></p></td><td class="cell"><p><b>-1.22</b></p></td><td class="cell"><p><b>0.00</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p><b>271_topic_aig</b></p></td><td class="cell"><p><b>+ 17.97</b></p></td><td class="cell"><p><b>-4.40</b></p></td><td class="cell"><p><b>+0.09</b></p></td><td class="cell"><p><b>+ 13.67</b></p></td><td class="cell"><p><b>+ 14.10</b></p></td><td class="cell"><p><b>+ 4.63</b></p></td><td class="cell"><p><b>+ 10.26</b></p></td><td class="cell"><p><b>+ 15.70</b></p></td><td class="cell"><p><b>+9.65</b></p></td><td class="cell"><p><b>+ 2.29</b></p></td><td class="cell"><p><b>0.00</b></p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></section><references><p>Eneko Agirre, Olatz Ansa, Eduard Hovy, and David Martinez. 2000. Enriching very large ontologies using the www. In <i>Proceedings of the Workshop on Ontology Construction of the European Con­ference of AI (ECAI).</i></p><p>Kenneth Church and Patrick Hanks. 1990. Word as­sociation norms, mutual information and lexicog­raphy. In <i>Proceedings of the 28th Annual Meeting of the Association for Computational Linguistics (ACL-90), </i>pages 76-83.</p><p>Thomas Cover and Joy A. Thomas. 1991. <i>Elements of Information Theory. </i>John Wiley &amp; Sons.</p><p>Gerald DeJong. 1982. An overview of the FRUMP system. In Wendy G. Lehnert and Martin H. Ringle, editors, <i>Strategies for natural language processing, </i>pages 149-76. Lawrence Erlbaum As­sociates.</p><p>Ted Dunning. 1993. Accurate methods for the statistics of surprise and coincidence. <i>Computa­tional Linguistics, </i>19:61-74.</p><p>Marti Hearst. 1997. TextTiling: Segmenting text into multi-paragraph subtopic passages. <i>Compu­tational Linguistics, </i>23:33-64.</p><p>Eduard Hovy and Chin-Yew Lin. 1999. Automated text summarization in SUMMARIST. In Inder-jeet Mani and Mark T. Maybury, editors, <i>Ad­vances in Automatic Text Summarization, </i>chap­ter 8, pages 81-94. MIT Press.</p><p>Kevin Knight and Steve K. Luk. 1994. Building a large knowledge base for machine translation. In <i>Proceedings of the Eleventh National Conference on Artificial Intelligence (AAAI-94).</i></p><page local="7"/><p>Figure 2: F-measure vs. summary length for all four topics. Topic signature clearly outperform <i>tfidf </i>and baseline except for the case of topic 258 where performance for the three methods are roughly equal.</p><doubt alpha="9.7" length="72" tooSmall="False" monospace="0.0">II0%I10%I20%I30%I40%I50%   |60%   |70%   |80%   |       90%   |   100% |</doubt><doubt alpha="4.6" length="196" tooSmall="False" monospace="0.0">I151-baaeline    |        0.308   |        0.349   |        0.400   |        0.434   |        0.423   |        0.402   |        0.385   |        0.373   |        0.376   |     0.370   |    0.355 |</doubt><p>Table 3: F-measure performance difference compared to baseline method in percentage. Columns indicate at different summary lengths related to full length documents. Values in the baseline rows are F-measure scores. Values in the <i>tfidf </i>and topic signature rows are performance increase or decrease divided by their corresponding baseline scores and shown in percentage.</p><p>Inderjeet Mani, David House, Gary Klein, Lynette Hirschman, Leo Obrst, Thérèse Firmin, Michael Chrzanowski, and Beth Sundheim. 1998. The TIPSTER SUMMAC text summariza­tion evaluation final report. Technical Report MTR98W0000138, The MITRE Corporation.</p><p>Christopher Manning and Hinrich Schütze. 1999. <i>Fondations of Statistical Natural Language Pro­cessing. </i>MIT Press.</p><p>Kathleen McKeown and Dragomir R. Radev. 1999. Generating summaries of multiple news articles. In Inderjeet Mani and Mark T. Maybury, edi­tors, <i>Advances in Automatic Text Summarization, </i>chapter 24, pages 381-389. MIT Press.</p><p>Ellen Riloff and Jeffrey Lorenzen. 1999. Extractionbased text categorization: Generating domain-specific role relationships automatically. In Tomek Strzalkowski, editor, <i>Natural Language In­formation Retrieval. </i>Kluwer Academic Publishers. Ellen Riloff. 1996. An empirical study of automated dictionary construction for information extraction in three domains. <i>Artificial Intelligence Journal, </i>85, August.</p><p>SAIC. 1998. Introduction to information extraction, http: //www.mue.saic.com.</p><p>Jinxi Xu and W. Bruce Croft. 1996. Query ex­pansion using local and global document analysis. In <i>Proceedings of the 17th Annual International ACM SIGIR Conference on Research and Devel­opment in Information Retrieval, </i>pages 4-11.</p></references></body></article>