<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>Computers in the Yugoslav Serbo-Croat/English Contrastive Analysis Project</title><author surname="Bujas" givenname="Zeljko"><org  name="Zagreb University" city="Zagreb Yugoslavia"/></author></firstpageheader><frontmatter><p><u>Computers in the Yugoslav Serbo-Croat/English Contrastive</u></p><p><u>Analysis Project</u></p><p>2eljko Bujas, Ph.D. Assistant Professor Department of English Zagreb University, Zagreb, Yugoslavia</p></frontmatter><abstract></abstract></header><body><section title=""><p><u>0.1.</u>       As far as the present writer is aware, the Yugo­slav Serbo-Croat/English Contrastive Analysis Project' is the first contrastive analysis effort to use a large cor­pus of parallel texts. The corpus is made up of the Brown Corpus (reduced by <i>50%) </i>with its Serbo-Croat translation, and a smaller Control Corpus (Serbo-Croat originals and English translation). A total, thus, of twice 500,000 words plus twice 150,000 words, or a grand total of some 1,300,000 words of running text.</p><p><u>0.2.</u>       The Project, let us make it clear, is not exclu­sively based on this corpus. Compilation and confron­tation of grammatical statements by various authors, plus plain old intuition, figure prominently in the methodol­ogy. The insistence on a large corpus, however, is due to the conviction, prevailing among the Project workers, that only an extensive investigation of correspondences (original-language elements and their translations) can adequately reveal the less predictable patterns which tend to have a considerable contrastive analysis potential.</p><p><u>0.21.</u>     The most productive method of obtaining correspon­dences from our corpus is to concordance separately its Serbo-Croat and English parts,  then to merge the resulting KWIC concordances into a contrastive KWIC concordance (with English keywords and alternating English and Serbo-Croat lines). For the more promising patterns, the merging procedure will be used twice, with both English and Serbo-Croat keywords.</p><p><u>0.22.</u>     In view of the size of the corpus, and the exten­sive concordancing required as a major procedure in the Project, the need for computer processing is obvious. It requires no undue strain on imagination to realize the soul-numbing effect of sheer physical handling of this mass of text if written out on slips.</p><p>Even in its most efficient and flexible form of a manual concordance (a sentence-slip file with keywords underlined monolingually), without which no manual pairing<page local="2"/></p><p>Df correspondences is possible, the manual handling of this 1,300,000-word corpus calls for a staggering amount 3f time and effort to prepare. According to our careful sstimate, a total of 7,100 man-hours is required to make such a concordance (without the 1,900 hours of transla­tion from English to Serbo-Croat, and vice versa).</p><p><u>3.23.</u> The slip file thus obtained would, however, secure snly a one-way approach: either from English or Serbo-2roat. A slip-file allowing a two-way approach would re­quire an additional effort of at least 4,500 man-hours.</p><p><u>3.24.</u> Finally, even these two manual concordances would still leave unfilled the need for reverse concordancing, so important for morphosyntactic research. To meet this leed,  two additional (though less ample) slip files would :iave to be established.</p><p><u>UP.</u>       In view of all this,  the Yugoslav Serbo-Croat/ Snglish Contrastive Analysis Project has from the outset Linked the planning of its work to the services of a local computer,  the City of Zagreb IBM 360/30 machine*</p><p><u>L</u><u>.1.       Stafie 1</u> of computer processing. The tape with the full text of the Brown Corpus (purchased from Brown Uni­versity, Providence, R.I., L'.S.h.), which had been pre­pared on an 1BK 7090 machine, had first to be converted from the density of 800 BP.l to 1,600 BPI, required by the Zagreb computer.</p><p><u>L</u><u>.11.</u>     After this, a printout of the entire text was ob­tained on the Zagreb machine. The printing took about sight hours, with a. special program<footnote anchor="3"/>restructuring the original format of the Brown Corpus text. This program Left out the location-marker column on the right-hand mar­gin of printout*, and added a sequence of sentence numbers (from 00001 to 52533) on the left.</p><p><u>L.12.</u>     The full text of the Brown Corpus was now reduced -y 50/', retaining, however, as closely as possible, the same proportions of the 15 genres (styles) contained in the Corpus. , <u>L.</u><u>13.</u>     Printouts of the samples retained in this reduced /ersion were then sent out to reliable translators, se­Lected to be representative of the three major regional /ariants of Serbo-Croat (western, central and eastern). Their instructions were to translate at normal speed, and is carefully as when they do any other paid translation »ork. The only limitation imposed upon them was to observe the sentence limit in the original (English or, in the used for the preparation on the IBM 360/30 of a full for­ward KWIC concordance of the Serbo-Croat Corpus.<page local="3"/></p><p><u>Sta^e 8.</u> Using the same tape, we now plan to pro­duce a reverse KWIC concordance of the Serbo-Croat text, ihis concordance will be selective in the same sense that the English reverse concordance was (cf. Stage 4).</p><p><u>1.9.     Sta^e 9.</u> With the normal and reverse KWIC concord­ances of both the English and Serbo-Croat corpora now ob­tained*, we can move on to the final stage(s) of central importance to the Froject, i.e. the merging of these mono­lingual concordances to get contrastive concordances (cf. ~.Z_.). We have planned four such concordances, and have "'tempted to illustrate them here by short simulated sam­ples, ns at the time of writing this no concordances of the Brown Corpus text (.either original or translation) were available,  the text used for these samples is the ierbo-Croat original and its translation into English of the novel <u>Povratak Filipa Latinovicza</u> (The Return of Ihilip Latinovicz) by the contemporary Croat writer Miro­slav Krleza.</p><p><u>1.91.     Forward contrastive concordance</u> <footnote anchor="9"/>(English to Serbo-Croat)</p><p>)JJ3 NO THE DOOR LOCKED, AND HIMSELF SHUT ÜUI IN I HE STREET. AND EVER SINCE THEN ;.m ASZAJ ZAKLJUCYANA VRATA I DSUU NA ULICI.  IE ÜIADA ZZIV1  NA ULIC1 VECZ MNOG 1312 i T3NGJL OF THE COKE FLAME FLICKERED OUT FROM UNjER <i>1 </i>HE PAINTED IRON STOVE/ 3312 SU JtZlCVAC &lt;0KS0V0S PLAMENA PU3 STALKOH NASLIKANE GVOZDENE PATENT-PfcCZ 1/2 3J-.D RIANLS STATJE ANC THEY NcrfER GOT HIM OUT. ANt THc WATER AHOVE HIM MAS STAIN ::&lt;0 JRIJ AN A I  DA OA V1SZE NIKADA NISU 1ZVUKLI , NEGO SE JE  SAMO VODA ZAKRVARILA ))•,•, o WHEN HIS 3HN MOTHER HAÜ lURNtD HIM OUT INTO THÊ STREET IN MORAL INDIGNAI I Dj«.&lt;i    J'JTRA,  KADA GA JE R3DZENA MAJKA  1ZBAC1LA NA ULICU S MORALNIM ZGRAZZ ANJEM,</p><p>Jibb iE DISTANCE. EVERYTHING HAS SHELLING OUT IN THE SILENT INSTRUMENTATION OF T JjSb J DAL J INAMAt  SVE JE RASLO KAO TIHA INS TRUMENi AC 1 JA MODRÜG JUT ARNJEG BUDZENJ 3UJ WITH ITS RULLS OF DRlAO, <b>— </b>ALL GAVE OUT THE ACRID AND PUNGENTLY ACRID S MEL 3HJ  tMLJAMA.  KAO KOPR ENA/1/  1Z S VEGA  STRUJI  OS/TAR I  OSJETLJIVO VLAZZAN VONJ DJ jl96 ARLIER, THE STUFFING HAD BEEN COMING OUT, A MASS OF BANDS, CURLY FEATHERS A <i>:</i><i>Hb  </i>tT  G30UA,  PROVIRIVALA UTROBA.  1SPUNJENA GURTAMA»  PERASTIM KOLUTI MA I CYUPE</p><p>A HAX CANDLE HAS BURNING OUT ON A MARBLE SiuARt OF THE CHJRCH F G3G3R1 Je VALA JE NA MRAMORNOJ CYEIVOR1NI  CRKVcNOG PODA JEDNA VOS 3271 RY HJMAN EYE. LIKE AN ANIMAL PEEPING OUT OF A CAGE/2/ HUMAN GESTURES ARE LI 32/1  JDSKOM OKU  IMA  IUGE.  KAKVOM Ü01ADZAJE PROMATRAJU 2Z1VOT1NJE  1Z KAVEZA/2/ KR</p><doubt alpha="60.0" length="5" tooSmall="False" monospace="0.0">It1 J</doubt><p>FIRE BLAZES OUT OF   1 HE  Ir6n THROATS  AND THERE  IS A</p><p>SUKLJA J3ANJ  IZ  ZZELJiiZNIH 2Z0R1JELA I  MIRISZE BARUT/1/ JE<page local="4"/></p><p>Control Corpus, in Serbo-Croat). They were not to split the English sentence into two or more Serbo-Croat senten­ces, nor were they allowed to combine two or more English sentences into one Serbo-Croat sentence.</p><p>The reason for this was the need to secure a me­chanical pairing of the English (or Serbo-Croat) keyword! marked by its sentence number, with the same-numbered, parallel, Serbo-Croat (or English) sentence in the two-language concordancing planned for the later Project stages.</p><p><u>1.2.</u> <u>Stage 2.</u> A new magnetic tape will be prepared of the reduced Brown Corpus text, and with the sentence se­quence numbers interpolated. This version will be used for all subsequent concordancing.</p><p><u>1.3.</u> <u>Stage 3.</u> Using this magnetic tape, the IBM 360/30 will now prepare a full forward KWIC concordance of the reduced Brown Corpus textf <u>1.</u><u>4.</u> <u>Stage 4.</u> Now (while the reduced Brown Corpus is still being translated) we shall use the same tape to ob­tain a reverse KWIC concordance of the same text. Since all "function words" - such as of, <u>had</u>. <u>most</u>, <u>those</u>, <u>did</u>, etc. - were already isolated in the previous stage (in the forward concordance)*this will further reduce the mass of text to be concordanced by one-halff <u>1.</u><u>5.</u> <u>Stage 5.</u> The Serbo-Croat translation of the reduced Brown Corpus, by now in an advanced stage, will be copied out on a Flexowriter in batches (as translators send in their typescripts), resulting in a paper tape.</p><p><u>1.51 •</u>    The same procedure can, at thir. stage, be applied to the 300,000 words of the Control Corpus. No time for translation has to be set apart here, since only already published English translations of Serbo-Croat originals are to be used.</p><p><u>1.6.</u> <u>Stage 6.</u> Although the 3erbo-Croa!t paper tapes ob­tained in the preceding stage are immediately computer-processable, we shall convert them to a magnetic tape, be­cause this medium secures an incomparably speedier proces­sing on the computer.</p><p><u>1.61.</u>   We hope that stages 2 to 6 will not take more than twenty weeks (if enough personnel can be hired simulta­neously) .</p><p><u>1.7.</u> <u>Stage 7.</u> The Serbo-Croat magnetic tape will now be<page local="5"/></p><p><u>1.92.</u> <u>Forward contrastive concordance </u>(Serbo-Croat to English)</p><p>2205 8UKVOM, GDJE SU SE 8ILI SKLONILI ONE BURNE NÜCZI , POSLIJE ROKUVUG PRUSZTENJ 2205 ID OAK-TREE WHERE THEY HAD FOUND  SHELTER THAT STORMY NIGHT  ON THEIR WAY BAC 2144 AJJ LIJECYNICI U SVOJIM TAJAN STVENIM BURNUSI MA /SZTO 12GLEDAJU KAO ST AROMOD 214* MOVED PHYSICIANS  IN   THEIR MYSTERIOUS &amp;URNOUSES  LIKE OLO-FASHIONED NIGHTSHIR 0216 1SERA, NAKOSTRIJESZENA LAVLJA GRIVAi BURSKE SATERI JE PRt"0 LADYSMITHOM, MARS 3216 RL-DIVERS,   THE L10NCS BRISTLING MANE.   THE BOER BATTERIES  AT LADYSMITH, THE 2144 MIRISZU JE. IMA LI U NJOJ KARAHELA, BUSZE MU PO ZUBIMA. MJERE MU TLAK KRV I 2144 S  AND SMELLING  IT  TO  FIND OUT WHETHER  THERE  WAS ANY SUGAR IN IT,  DRILLING H 1546 SMATO TALASANJE OUZOVA  I  LI SNJAC YA 1  BUTINA,  DEBELIH MASnIH ZZENSKIH NOGU, 1546    HAIRY BUTTOCKS AND CALVES AND  THIGHS,  FAT WOMENCS  LEGS,  ANKLES,  JOINTS, SK 0210 AKAVA STEGNA KONJS&lt;A&gt; KRVAVE RANJENE BUTINE, UZNE MI RE ne CRNE REPINE, RASKRV 3210  LANKS,   BLOOD-STAINED AND  WUUNÜED,   THEIR  LONG  BLACK  LASHING TAILS,   THEIR  r:L L 0473 EKANE POJASE MESA OKO KUKOVA I IZNA3 8UTUVA U LJtLlNl POItZA, A ÜVAJ TU <i>Lit </i>0473 SOFT  ROLLS OF FLESH ROUND  I HE HIPS <i>M3 </i>ABOVE   THÉ  IH1GHS,  WHILE THIS FELLIM 1314 AVODL AKAVOJ OBLINI KDNJSKIH STtGMA I RUTÜVA, TO JE JbDINI VtLIKl DOZZ WL J AJ 1314  1NING HAIRY FLANKS AND HINDQUARTERS, HAD BEEN THE ONLY GREAT  EXPERIENCE OF 0248 LAVE, ZZALOSNE PTICYJE OCYI, KRAVLJE BUTOVE, KONJSKA STEGNA, A SINOCZ JÜSZ 0248 SC LEGS,  WRETCHED BIRDSC WINGS, COWSi. BUTTOCKS, HORSESt HAUNCHES, WHILE ÜNL</p><doubt alpha="64.3" length="42" tooSmall="False" monospace="0.0">0984 CtPROSIM VAS,  JAGO, BUTE SPAMETNl/4/</doubt><doubt alpha="50.0" length="42" tooSmall="False" monospace="0.0">39B4 (.(.PLEASE,   YAGA » BE SENS Iß LE/4/</doubt><p><u>1.93.</u> <u>Reverse contrastive concordance </u>(English to Serbo-Croat)</p><p>0002 WAS  ALL STILL  FAMILIAR  TO HIM/l/  THt ROTTING, SUMY RUOFS,  THE ROUND BALL U 0002 NAO JE JOSZ UVIJEK SVL KAKO 0ÜLAZI/1/ I   IRULI  SLINÄVI  KROVOVI   I  JABJKA FRAT 0003 NTY-THRtE YEARS HAD PASSED SINCE THE MORNING WHEN HE HAD SLUNK UP TO THAT 0 0003 DES ET   I  TRI  G3DINE SU PROSZLE OD ONOG  JUTRA,  KADA SE OUVUKAO POD OVA VRATA 0003 EET, AND EVER SINCE THEN HE HAO BEEN LIVING IN THE STRECI, AND NUTHING HAD 0003 LICI,   TE UTADA ZZIVI NA ULIC1  VLCZ H.mOGO GODINA,  A NI S2TA SL NUE PR0M1JCNI 0003 £ HAD BEEN LIVING IN  THE  STREET,  ANB NUTHING HAD REALLY CHANGED.</p><p>0003 VI  NA UL ICI VECZ  MNOGO GODINA, A NISZTA SE NI JE  PRuMIJENILU UGLAVNUM.</p><p>0004 DLY LOCKED DOOR AND, JUST AS ON THAT MURNIN3, HE COULD FEEL THE COLD, IRON 0004 M ZAKLJUCYANiM VRATIMA,  1  KAJ I  ONÜG  JUTRA I MAO JE OSJECZAJ HLADNOG, GVUZDE 0034 AS HE PUSHED IT, HOW THE LEAVES WEHE 0UIVER|N3 IN THE UPPER BRANCHES OF THE 0004 NJEGOVUM RUKOM  I  ZNAO JE,  KAKO SE  LISZCZt  MlC YE U KRUSZNJAMA KESTENOVA  I CY 0034 RED, IN NEED OF SLEEP, HE CUULD FEEL SOMETHING CRAWLING INSIDE HIS COLLAR -0034 RAN,  NE ISPAVAN, OSJECZAJUCZI  KAKO MU NESZTO PLAZI   OKU UKOif RAT NI KA/1/  PO SVO 0005 RO,   LAST  DRUNKEN NIGHT,  AND  THE  GREY MORNING.</p><doubt alpha="65.2" length="161" tooSmall="False" monospace="0.0">0004 ASIFIN A DREAM -- ASONTHAT OTHfcR MORNING — /!/ HE WAS ALL DIRTY, TIRtD 0004  ILO MU JE /ONOG JUT^A/ KAO DA  SANJA/1/ BIO JE  SAV CYAUZAV, 'JMORAN, NEISPAVA</doubt><doubt alpha="62.0" length="92" tooSmall="False" monospace="0.0">0004  ED OF  SLEEP,   HE  COULD  FEEL   SOMETHING  CRAWLING  INSIDE  HIS  COLLAR —  A 8ED-EJ</doubt><doubt alpha="64.3" length="84" tooSmall="False" monospace="0.0">0004 N,  0SJECZAJUCZ1 KAK3 HU NESZTO PLAZI  UKO 0K0VRATN1KA/1/  PU SVOJ PRIl.ICI STJ</doubt><p>0005 PIJANL,  POSLJEDNJE,   TRECZE N0CZ1   I  ONUS  SIV06 JUTRA — DDK ZZIVI.</p><page local="6"/><doubt alpha="0.0" length="1" tooSmall="False" monospace="0.0">6</doubt><p><u>.94.     Reverse contrastive concordance </u>(Serbo-Croat to English)</p><p>0082 RVOREDA,  MEDUZINA GLAVA 00 SA DRE NAD TE SZKIH, OKOVANIH HRASTOVIM VRAI IHA I 0002 LASTER HEAD OF MEDUSA SURMOUNTING THE HEAVY, IRON-BOUND OAK DOOR WITH ITS C 0002 NEDJZINA GLAVA 00 SADRE NAD TESZKIM, OKOVANIH HRASTOVIM VRATIMA I ( »L*ONA KV 0002    OF MEDUSA SURMOUNTING THE HEAVY,  IRON-BOUND OAK DOUR WITH ITS  COLO LATCH.</p><p>0002 GLAVA 00 SADRE NAD TESZKIM. OKOVANIH HRASTOVIM VRATIMA I  HLADNA KVAKA. 0002 USA SURMOUNTING THE HEAVY,   IRON-BOUND OAK DOOR WITH ITS COLO LATCH.</p><p>03n4 ZASTAO JE PREO  STRAN1M ZAKLJUCYANIM VRATIMA,  I KAO I</p><p>JJoîJ*, HE STOPPED IN FRONT OF  THE UNFRIENDLY LOCKED DOOR AND,  JUST  AS ON T</p><p>Q004 ZASTAO JE PRED STRANIM ZAKLJUCYANIM VRATIMA,  I  KAO I ONOG JUT 0304    HE STOPPED  IN FRONT OF  THE UNFRIENDLY LOCKED OOOR AND,  JUST AS ON THAT MOR</p><p>0006    GDJE SE JE KAO MALI  DECYKO IGRAO  SA  SV0J1M BIJELIM JANJCEM, STAJALO JE GRA ERE AS A BOY HE HAD PLAYED WITH HIS WHITE  LAMB,   THERE WAS A BUILDING-SITE W 0006 E JE KAO MALI OECYKO IGRAO SA SVOJIM BIJELIM JANJCEM, STAJALO JE GRADIL1SZT <i>lilt </i>S  A BOY HE HAD PLAYED WITH HIS WHITE  LAMB,  THERE WAS A BUILDING-SITE WALLED 0006 JE GRADILISZTE OBZIDANO KAO CYOVJEK VIS0K1M ZIDOM 1 NA TOM VISÜKÜM ZIDU Bl <i>Hit </i>-SITE WALLED IN LIKE A MAN BEHIND A HUH WALL, AND ON THIS HIGH WALL THERE 0000 DUGO JE  STAJAO POD VI HUM ZMNSKIM SIÉZNICIMA,  A PRSTI SU</p><p><i>III, </i>RE FOR A LONG TIME UNDER  THE  SLIM CORSETS, AND HIS FINGERS WERE ALL DIRTY W</p><p>0009 OUGO JE STAJAO POD VI TKIM ZZENSKIH STEZNICIMA,  A PRST1 SU MU BIL</p><p><i>Uli </i>A LONG TIME UNDER THE SLIM CORSETS, AND HIS FINSERS WERE ALL DIRTY WITH DUS</p><p><u>.10.</u>     The reason why these four concordances have been resented under one processing stage (9) is that, first, e are not sure whether we can afford the computer for ach of them, and, second, we do not, at this point, know ow selective each of them is going to be. A considerable eduction of the text to be concordanced can be achieved n reverse concordancing if we restrict ourselves only to orde, ending in a characteristic morpheme with clearly oreseeable contrastive analysis potential (such as -ed. lv, -<u>est</u>, -<u>ing</u>, -<u>ness</u>, -<u>less</u>, etc. in English, and -ao, <u>vsi</u>, -en, -<u>scu</u>, -<u>ost</u>, -<u>ste</u>, etc. in Serbo-Croat).</p><p><u>.11.</u>     It may be pointed out here that, irrespective of ow restrictive the selection of keywords for concordanc-ng œay have to be, no concessions should be made in the rinciple of bilingual approach. Only if,,in our investi-ation of the contrastive potential of individual ele-ents, we strictly observe the approach from both the nglish and the Serbo-Croat texts, can we be certain that e shall have covered all possible contrastive description atterns based on correspondences in both corpora.</p><p><u>.0.</u>       Once contrastive concordancing has been completed we shall still be facing some practical technical problems.<page local="7"/></p><p><u>2.1.</u> Project analysts, for instance, will often have to be provided with slips instead of computer printout sheets. Only if the material being analyzed is in the form of slips will they be able to classify and reclassify the key elements swiftly and flexibly (by putting together, break­ing up and re-establishing batches of slips).</p><p><u>2.11.</u>     Cutting up the concordance printouts to get the slips is not very practical in view of the varying size of contrasted pairs of elements with their context (cf. n. 9, second half). The way around this, clearly, is to have the pairs printed out at regular intervals with sufficient blank space in between. This, however, would probably triple the amount of printout paper required. Also, this is complicated further by the need for a number of copies for each pair (slip), because of simultaneous demands that may often be made upon the same slip by several Project analysts, approaching the same element from various des­criptive levels. These copies could be secured by using special, multiple-carbon printout paper, but this might prove quite expensive.</p><p><u>2.2.</u> In view of all this, the Yugoslav Serbo-Croat/ English Contrastive Analysis Project has envisaged the use of a Flexowriter here as an alternative method. This ma­chine has already provided us with the paper tape of the Serbo-Croat translation of the reduced Brown Corpus, plus the tapes of Serbo-Croat originals and English transla­tions of the Control Corpus (cf. Stage 5). The missing paper tape of the English text of the Brown Corpus can be obtained on a magtape-to-papertape converter. Once both paper tapes are ready, running them through the Flexo­writer provides us with up to 13 (some claim 20) carbons of each contrasted pair. An additional advantage of using the Flexowriter for slip duplication is in the less awk­ward shape of slips. Paper tapes reproduce the text in 60-character-wide lines of the original translators' type­script, as opposed to the 110 to 120-character streamers of normal computer printout (unless the concordance print­out was programmed for a narrower format, requiring con­siderably more paper).</p><p><u>2.3.</u> The resulting slip files of sentence-numbered</p><p>English and Serbo-Croat texts, coupled with the Project's basic (monolingual - forward and reverse) concordances, can now be used as a replacement for contrastive concor­dances. It would work approximately like this: upon receiving an analyst's request for examples of all corresponden­ces in the corpus of an element under analysis,  the Project headquarters in Zagreb would look the element up in one of the basic concordances, record sentence numbers of all the occurrences, extract slips bearing these numbers from the Flexowricer-produced slip file, and forward them to the analyst for further research.<page local="8"/></p><p>Footnotes</p><p>1. Launched in 1968, at the Institute of Linguistics, Faculty of Arts and Letters, Zagreb University. Direc­tor: Professor Rudolf Filipovic, Ph.D.  (Fostal address: Jugoslavenski projekt za kontrastivnu analizu srpsko-hrvatskog i engleskog jezika, Institut za lingvistiku, Filozofski fakultet, Djure Salaja 3, Zagreb, Yugosla­via). Project analysts, numbering 20, are on English department staffs from all parts of Yugoslavia.</p><p>2. oize of storage: 32k. Other equipment: three 2311 discs, two 2415/4 tape drives, one 2540 card reader, one 2671 -pper-tape reader, one 1403/2 printer.</p><p>..ritten by Dipl. ing. Kilutin Gihlar, Chief Programmer of the Zagreb system.</p><p>4. Cf. <u>i-.anual of Infor</u>m<u>ation</u> (for the Brown Corpus), Brown Iniversity, 1964, p. 7.</p><p>5. «e hope to use forward and reverse concordancing prog­rams developed by a US project for an IBK 360/30, or a similar machine.</p><p>6. In a total reverse concordance they would only appear in a different place: of under F, <u>had</u> under D, etc.</p><p>7. Putting the top 100 words from the Brown Corpus Rank List on the exclusion list (compared to a total of some ISO 'function words", in the present author's estimate), would reduce the text by 47.4 per cent, while including only one morphologically marked word (YEARS) and two lexical words (liEW, TIKE). Expanding the exclusion list to cover the top 200 words would probably not be econom­ical (though only two additional morphologically marked words would be included: UKITED and STATES), because the computer would be slowed down, whereas the textual mass would be reduced by only 6 more per cent (to 53.<page local="9"/>6 per cent.</p><p>8. Which may take between 40 and 60 computer hours, as op­posed to an estimated 2,350 hours of manual processing (for only the English forward concordance at that).</p><p>9; In addition to being simulations, all these concordance samples are in an idealized format, with the correspon­dences spatially parallel to the keyword. In practice, however, it is impossible to achieve this ideal textual parallelism, because there are no other formal signals to govern it, except the sentence sequence number which can only mark the sentence as a whole.</p><p>For this reason,  the actual computer concordances will, when ready, have the correspondence to the keyword printed out with the whole sentence in which it occurs, under the single line with the keyword. This will, natu­rally, increase the size of the concordance, but not more than about 50 per cent in our estimate. This is be­cause only an approximate 40 per cent of all sentences in the original text of the Brown Corpus are in excess of 20 words (which can be accommodated by the average printout line). A mere 6 per cent of these sentences are longer than 40 words, requiring, consequently, r..ore than two printout lines.</p></section></body></article>