<?xml version="1.0"?><!DOCTYPE article SYSTEM "/project/take/software/searchbench_offline_processing/paperxml_generator/aclextractor/src/python/../resource/dtd/paperxml.dtd"><article><header><firstpageheader><page local="1"/><title>The PIT Corpus of German Multi-Party Dialogues</title><author surname="Strauß" givenname="Petra-Maria"><org  name="University of Ulm" country="Germany" city="Ulm"/></author><author surname="Hoffmann" givenname="Holger"><org  name="University of Ulm" country="Germany" city="Ulm"/></author><author surname="Minker" givenname="Wolfgang"><org  name="University of Ulm" country="Germany" city="Ulm"/></author><author surname="Neumann" givenname="Heiko"><org  name="University of Ulm" country="Germany" city="Ulm"/></author><author surname="Palm" givenname="Günther"><org  name="University of Ulm" country="Germany" city="Ulm"/></author><author surname="Scherer" givenname="Stefan"><org  name="University of Ulm" country="Germany" city="Ulm"/></author><author surname="Traue" givenname="Harald"><org  name="University of Ulm" country="Germany" city="Ulm"/></author><author surname="Weidenbacher" givenname="Ulrich"><org  name="University of Ulm" country="Germany" city="Ulm"/></author></firstpageheader><frontmatter><p><b>The PIT Corpus Of German Multi-Party Dialogues</b></p><p><b>Petra-Maria Strauß*, Holger Hoffmann*, Wolfgang Minker* Heiko Neumann*, Günther Palm*, Stefan Scherer* Harald C. Traue*, and Ulrich Weidenbacher*</b></p><p>University of Ulm *Inst. of Information Technology * Inst. of Neural Information Processing *Inst. of Medical Psychology Ulm/Donau, Germany {firstname.lastname} @uni-ulm.de</p></frontmatter><abstract>The PIT corpus is a German multi-media corpus of multi-party dialogues recorded in a Wizard-of-Oz environment at the University of Ulm. The scenario involves two human dialogue partners interacting with a multi-modal dialogue system in the domain of restaurant selection. In this paper we present the characteristics of the data which was recorded in three sessions resulting in a total of 75 dialogues and about 14 hours of audio and video data. The corpus is available at http://www.uni-ulm.de/in/pit. </abstract></header><body><section number="1." title="Introduction"><p>In this paper we present the PIT corpus of German multi­party dialogues recorded in the context of the <i>Competence Centre Perception and Interactive Technologies<footnote anchor="1"/> </i>(PIT) at the University of Ulm. PIT joins various institutes of the University to perform interdisciplinary research in the field of advanced human-computer interaction. Our objective is to develop components and technologies for intelligent and user-friendly human-computer interaction in multi-user en­vironments.</p><p>Future dialogue systems will be endowed with more human-like capabilities: They should be adaptive in terms of the users' needs and preferences. They should be flexible in terms of the number of users that interact with the sys­tem. They should be endowed with perceptive skills from different sensory channels (vision, hearing, haptic, etc.). And they should be aware of the users' emotional state and possess enhanced conversational skills, just to name a few of the desired character traits of future dialogue sys­tems. Integration of perception, emotion processing, and multimodal dialogue skills in interactive systems will not only improve the human-computer communication but also human-human communication over networked systems. At present, research on multi-party interaction is very pop­ular especially in the context of the meeting scenario. Var­ious corpora have been published, e.g. the ICSI (Janin et al., 2003) and the AMI (Carletta et al., 2005) corpus. The meeting scenario requires intelligent computer systems to enhance and assist the human communication during meet­ings, however, our aim is to integrate the computer system as an equal dialogue partner in the communication with sev­eral humans.</p><p>As far as the authors are aware, there is no existent collec­tion of data stressing the designated features important for our research. Thus, we built our own data corpus presented in this paper. The corpus comprises transcribed audio and video data. It emerged from Wizard-of-Oz (WOZ) recordings conducted in 2006 and 2007 in the framework of the research project 'The computer as a dialogue companion - Perception and interaction in multi-user environments'. The obtained data form the basis for our research, to de­velop mechanisms for sophisticated human-computer inter­action, which we will not go into detail here. In this paper we present our corpus which we believe to be of great value for the research community in human computer interaction. The paper is structured as follows. The following section introduces the scenario of the recordings. Section 3 briefly describes the recording setup. Section 4 presents the col­lected data in detail, including the participants, audio and video data. Section 5 concludes the paper.</p></section><section number="2." title="Scenario"><p>The PIT scenario is restaurant selection and composed of three dialogue participants: Two humans are discussing their choice of restaurant, the computer assisting them. The conversation takes place at a desk, one of the users (the main user 'Ul') is sitting in front of the computer. Only this user interacts directly with the system, i.e. addresses it or expresses willingness to communicate by looking at it. The system, however, overhears the complete conversation between the participants.</p><p>The system acts as an independent dialogue partner which becomes active as soon as the users start to speak about the specified domain. At the beginning of the dialogue the users talk about a random topic. As soon as the conversa­tion enters the restaurant domain, the system 'listens' at­tentively. Upon an interaction request by the main user or proactively when required by the conversational situation, the system takes the initiative and gets meaningfully in­volved in the communication. The system is represented by an avatar (not for all dialogues, see below) and synthesised speech output. Further, a restaurant's menu, directions in form of a city map, or bus schedules are occasionally displayed on the screen. An example dialogue excerpt is shown in Figure l <footnote anchor="2"/>. The dialogue partners were randomly<page local="2"/></p><footnote label="1">www.uni-ulm.de/in/pit</footnote><footnote label="2">As the conversation language is German, the example was</footnote><p><b>U1: </b>Computer, could you please give us a Mexican restaurant in the city centre of Ulm.</p><p><b>S: </b><i>The query for Mexican food resulted in two hits. My</i> <i>suggestions are restaurant Flamenco and restau­rant Peppers.</i><i></i></p><p><b>U2:       </b>Ask him where they are.</p><p><b>U1:       </b>Where are the Mexican restaurants?</p><p><b>S: </b><i>Restaurant Flamenco offers Mexican and Spanish</i></p><p><i>food for moderate prices, possesses a terrace and is located in the city centre next to the Danube. Restaurant Peppers offers ...</i></p><p><b>U1:       </b>In which street is the restaurant?</p><p><b>S: </b><i>Flamenco is located at number eight Inselgasse.</i></p><p><b>U1:       </b>Great! How are the prices in this restaurant?</p><p><b>S: [Menu pops up on screen.] ...</b></p><figure caption="Figure 1: Excerpt from one of the recorded dia­logues (Strauß et al., 2007)."></figure><p>assigned different scenarios in form of descriptions of roles and situations that should be adopted in the conversation. The combination of assigned roles included amongst others employer and employee, business colleagues, friends, or a couple in a loving relationship. This procedure not only eased the beginning of the dialogue, it brought in variety in terms of different social contexts. Most times, the par­ticipants followed the role play quite well. In a few cases, however, the scenario was only used to stimulate the con­versation and not referred back to later during the dialogue.</p></section><section number="3." title="Data Collection Setup"><p>The recordings of the data were performed in the Wizard-of-Oz setup described in (Strauß et al., 2006). The wizard simulates the envisioned final system's behaviour as closely as possible, replacing the system components that are not yet functional (such as e.g. the speech recogniser) using the tool described in (Scherer and Strauß, 2008). The setup of the system is shown in Figure 2. The human dialogue partners U1 and U2 interact with the system S which is op­erated by the wizard situated in a different room. U1 is</p><p>directly translated into English.</p><figure caption="Figure 2: Data collection setup (Strauß et al., 2006)"></figure><p>the system's main interaction partner. The dialogues are recorded by three microphones and three cameras as fol­lows. U1 and U2 each wear a lapel microphone (M1 and M2). The signals are sent via wireless transmission to the wizard's computer where they are recorded and played back to the wizard. Room microphone M3 records the entire scene including the system output. Camera C1 is placed as to record from the system's point of view, focusing on the face of U1. C2 is installed behind U1 in order to grasp Ul's point of view, i.e. the screen and the other dialogue partner U2. C3 records the entire scene. Refer to (Strauß et al., 2006) for details on the technical equipment.</p></section><section number="4." title="Wizard Policies"><p>Prior to recording the conversations it was necessary to as­sess several policies the wizard has to follow to assert uni­form system behaviour throughout all dialogues in order to receive unbiased dialogue data. In general, there were two different kinds of targeted data: First, there was general or normal dialogue material and on the other hand emotional data. Both types of interaction followed the same princi­ples: The conversation between the two dialogue partners started without wizard interaction in the beginning when the users acted out the provided scenarios. The differences emerge once the system has joined in the conversation, as pointed out in the following two paragraphs.</p><p><b>Standard Policy. </b>During a standard recording the sys­tem, or wizard respectively, did not interrupt the conver­sations of the users. Speech recognition and language un­derstanding was simulated to be perfect. Correct and help­ful answers were given as frequently and as promptly as possible. The system's first and further interactions were triggered by the following situations:</p><p>• Reactively upon user U1 addressing the system di­rectly</p><p>• Proactively on it's own behalf in order to make a sig­nificant contribution to the conversation (e.g. to report a problem in the task solving process)</p><p>• After a pause in the dialogue exceeding a certain threshold (if a meaningful contribution can be made)</p><p>In the case where U2 addressed the system directly, the ut­terance was recognised, however, no direct response was given. Yet, two different reactions were observed. Either, user U1 instantly took the turn and posed the same or sim­ilar request to the system, or U2's request was followed by a pause which again would mostly lead to an interaction of the wizard. After the users found a restaurant that pleased everyone the wizard generally closed the dialogue. In some cases, however, the participants decided to search for an­other locality, such as a bar for a cocktail after dinner etc. Overall, the recordings adhering to the standard policy re­semble a perfectly working dialogue system.</p><p><b>Emotion Policy. </b>In order to induce emotional behaviour of the users a specific emotion policy was used in several dialogues. During these recordings the wizard occasionally interrupted the users in order to appear rude. Furthermore, repetitive mistakes were made and wrong answers given.</p><page local="3"/><p>Common mistakes included recognition and understanding errors that would sound very similar to the users' actual wishes, e.g. if the user was looking for a "not expensive restaurant" the wizard would return a selection of expensive restaurants. Additionally, the wizard would sometimes just pause to bore the users.</p><p>The emotions obtained using this strategy include anger, boredom and surprise. Emotions such as happiness and sur­prise were also induced by the standard policy, after receiv­ing correct answers and useful informations. Surprise is often induced by the first interaction of the system, for ex­ample after the face of the avatar is shown to the users for the first time. In general, it is to say that the emotions ex­pressed by the users are more moderate than artificial emo­tions played by actors, as in (Burkhardt et al., 2005). How­ever, we consider these moderate emotions as more realistic and common in human computer interaction and therefore very useful for affective computing tasks.</p></section><section number="5." title="Collected Data"><p>The corpus consists of 75 dialogues from three record­ing sessions, refer to Table l for number of dialogues and durations. Session I was performed in 2006. At that stage, the system output consisted of only acoustic output (speech synthesis) and the display of the restaurant's menu in HTML format when required. For the second block of recordings (2007) the system was enhanced. The response time of the wizard was improved. An avatar was integrated to represent the system visually. Furthermore, street maps were included to be shown on the screen. For the third recording session (2007), the system was enhanced to also present bus schedules on the screen. Half (18) of the di­alogues were recorded with the avatar on the screen, half without the avatar. Synthesised and visual output in form of menus and maps was the same for all recordings in this session. The shortest recorded dialogue was 2:43 minutes long (session III), the longest lasted 33:39 minutes (session</p><p>II).</p><table caption="Table 1: Statistical information of the three recording ses­sions."></table><subsection number="5.1." title="Participants"><p>Participants (n=150) were students and employees of the University, who gave written consent to participate in this study. They were between 19 and 51 years of age (on average 24.4 years); 53 of them were female (4 at ses­sion I (10.5%), 18 at session II (45.0%), 31 at session III (43.1%)). Except for six participants, the native language of all participants was German.</p><p>In order to evaluate the dialogue system, several question­naires had to be completed by the participants after the in­teraction with the system. The AttrakDiff (Hassenzahl et al., 2003) questionnaire was used to measure the attractive­ness as well as the pragmatic and hedonic quality of the sys­tem. To evaluate the direct interaction between the human dialogue partner Ul and the computer system, a short ver­sion (n=16 items) of the SASSI (Hone and Graham, 2000) questionnaire was selected. Furthermore, data on partici­pants' technical self assessment was collected. The results show that the evaluation of the system signifi­cantly improved from recording session I to III in terms of attractiveness, usability and acceptance of the system. Fur­ther analysis of the data has to be done in order to reveal the impact of several changes (avatar, response time) on the evaluation of the system.</p></subsection><subsection number="5.2." title="Audio Data"><p>The audio data were recorded using three microphones: One lapel microphone for each participant and a room mi­crophone to record the entire scene including the system output. The audio data were recorded at 16 kilohertz with 16 bit resolution. External sound cards were used to im­prove the recording quality.</p><p>The dialogues all follow a certain pattern marked by three phases: Each dialogue starts with a domain independent chat between the participants. The next phase of the di­alogue is introduced at the point when the conversation switches over to the specified domain and the users start discussing their preferences and aversions in different as­pects of the restaurant domain. The third part is charac­terised by the involvement of the dialogue system in the conversation to achieve the concerted task. The dialogues typically end when the users find a suitable restaurant and thank the system. Some recordings contain various itera­tions of the restaurant search, i.e. after finding one, instead of ending the dialogue, the users started to look for another restaurant (remaining in the third phase).</p></subsection><subsection number="5.3." title="Video Data"><p>Social interaction between humans is not only limited to verbal communication. Also visual communication plays a significant role. In a dialogue scenario, non-verbal commu­nication is particularly characterised by analysing the gaze of a dialogue partner. Directed gaze signalises attention while averted gaze signalises inattentiveness. Therefore it is very interesting to extract and evaluate pose behaviour of individual subjects during the conversation. The dialogues were video recorded from three different an­gles. Figure 3 shows the scene from the viewpoint of cam­eras C3 (long shot) and C1 (face of U1). The goal is now to annotate each video frame (using the data from Cl) with a specific class label to discriminate be­tween video frames where the person (U1) attends to the system and video frames where the person attends to the human dialog partner (U2). This problem is closely linked to the field of automatic image annotation (Cusano et al., 2004), (Jeon and Manmatha, 2004), where a system auto­matically assigns metadata in the form of keywords to an image.</p><p>Here, we train an adaboost classifier (Viola and Jones, 2004) with a small subset of manually annotated image frames in order to automatically extract pose information (direct gaze or averted gaze) from the video data which was acquired during the dialog session.<page local="4"/> We chose adaboost, be­cause this approach is known to be very fast and efficient in detecting faces in images. We trained two different clas­sifiers, one that finds frontal faces against background and one that finds averted faces against background. Finally, both classifiers are then applied to each of the remaining previously unlabeled video frames to determine their pose label. This information now enables us to statistically eval­uate the amount of time user U1 spent focusing on the sys­tem as opposed to on the other user. The video data can again be structured in three interac­tion phases considering the gaze direction of the main user U1. These phases differ from the dialogue phases described above. The first interaction phase is characterised by the conversation between the human dialogue partners before the first system interaction. During this time, there is almost no gaze directed towards the computer screen. The first interaction of the system initiates phase two. During this phase, U1's gaze switches between the computer and user U2, depending on speaker and addressee. The third phase is characterised by an object (other than the avatar) displayed on the screen: Generally, while a restaurant's menu, a street map, or bus schedule is shown on the screen, most of U1's gaze points towards the system. When the object is hidden, the dialogue returns to phase two.</p><table caption="Table 1: Statistical information of the three recording sessions." class="main" frame="box" rules="all" border="1" regular="False"><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Session</p></td><td class="cell"><p>I</p></td><td class="cell"><p>II</p></td><td class="cell"><p>III</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Number of dialogues</p></td><td class="cell"><p>19</p></td><td class="cell"><p>20</p></td><td class="cell"><p>36</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Duration of session</p></td><td class="cell"><p>3:47 h</p></td><td class="cell"><p>4:18 h</p></td><td class="cell"><p>5:40 h</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"><p>Average dialogue duration</p></td><td class="cell"><p>12 min</p></td><td class="cell"><p>13 min</p></td><td class="cell"><p>10 min</p></td><td class="cell"></td></tr><tr class="row"><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td><td class="cell"></td></tr></table></subsection><subsection number="5.4." title="Annotation"><p>The data was transcribed at the utterance level and anno­tated with dialogue acts. Table 2 presents the basic tagset of dialogue acts we used and which proved suitable for our domain and dialogue manager requirements.</p><doubt alpha="100.0" length="6" tooSmall="False" monospace="0.0">Tagset</doubt><p>suggest, request, inform, accept, reject, acknowledge, check stall, greet, other</p><table caption="Table 2: Dialogue act tagset used on PIT Corpus."></table></subsection></section><section number="6." title="Conclusion"><p>In this paper we presented the PIT Corpus, a multi­modal collection of 75 multi-party dialogues recorded in a Wizard-of-Oz setting. Each dialogue involved two human dialogue partners and a computer system, recorded with</p><p>Figure 3: Video recordings from the viewpoint of cameras C3 (left) and C1 (right) (Strauß et al., 2007).</p><p>various microphones and video cameras. The corpus can be found on our website (http://www.uni-ulm.de/in/pit). We hope it to be of great benefit for the HCI research commu­nity.</p></section><section number="7." title="Acknowledgements"><p>This work has been supported by a grant from the Ministry of Science, Research and the Arts of Baden-Württemberg (Az:23-7532.24-13-19/1).</p></section><references><p>F. Burkhardt, A. Paeschke, M. Rolfes, W. Sendlmeier, and B. Weiss. 2005. A database of german emotional speech. In <i>Proceedings of Interspeech 2005.</i></p><doubt alpha="66.7" length="168" tooSmall="False" monospace="0.0">J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronen­thal, G. Lathoud, M. Lincoln, A. Lisowska, I. McCowan,</doubt><doubt alpha="60.0" length="50" tooSmall="False" monospace="0.0">W. Post, D. Reidsma, and P. Wellner. 2005. The AMI</doubt><p>meetings corpus. In <i>Proceedings of the Measuring Be­havior 2005 symposium on "Annotating and measuring Meeting Behavior".</i></p><p>C. Cusano, G. Ciocca, and R. Schettini. 2004. Image anno­tation using SVM. In <i>Internet Imaging IV, </i>volume SPIE.</p><p>M. Hassenzahl, M. Burmester, and F. Koller. 2003. AttrakDiff: Ein Fragebogen zur Messung wahrgenommener hedonischer und pragmatischer Qualität. <i>In J. Ziegler &amp; G. Szwillus (Hrsg.), Mensch &amp; Computer 2003. Interaktion in Bewegung </i>, pages 187-196.</p><p>K. S. Hone and R. Graham. 2000. Towards a tool for the subjective assessment of speech system interfaces (sassi). <i>Nat. Lang. Eng., </i>6(3-4):287-303.</p><p>A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stol-cke, and C. Wooters. 2003. The ICSI meeting corpus. In <i>Proceedings </i><i>ofthe</i><i> IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), </i>pages 364-367.</p><p>J. Jeon and R. Manmatha. 2004. Using maximum entropy for automatic image annotation. In <i>CIVR, </i>pages 24-32.</p><doubt alpha="64.3" length="56" tooSmall="False" monospace="0.0">S. Scherer and P.-M. Strauß. 2008. A Flexible Wizard-of-</doubt><p>Oz Environment for Rapid Prototyping. In <i>Proceedings of the 6th International Conference on Language Re­sources and Evaluation (LREC), </i>Marrakech, Morocco. P.-M. Strauß , H. Hoffmann, W. Minker, H. Neumann, G. Palm, S. Scherer, F. Schwenker, H. Traue, W. Walter, and U. Weidenbacher. 2006. Wizard-of-Oz Data Collec­tion for Perception and Interaction in Multi-User Envi­ronments. In <i>Proceedings </i><i>ofthe</i><i> 5th International Con­ference on Language Resources and Evaluation (LREC), </i>Genova, Italy.</p><p>P.-M. Strauß , H. Hoffmann, and S. Scherer. 2007. Eval­uation and User Acceptance of a Dialogue System Us­ing Wizard-of-Oz Recordings. In <i>3rd IET International Conference on Intelligent Environments, </i>Ulm, Germany.</p><p>P. Viola and M. Jones. 2004. Robust real-time face detection. <i>International Journal of Computer Vision, </i>57(2):137-154.</p></references></body></article>