Dialogue act modeling for automatic tagging and recognition of conversational speech

Andreas StolckeKlaus RiesNoah CoccaroElizabeth ShribergRebecca BatesDaniel JurafskyPaul TaylorRachel MartinCarol Van Ess-DykemaMarie Meteer

article2000CL2,929 citations

Develops a probabilistic framework combining discourse grammar, acoustic prosody, and word n-grams to automatically classify dialogue acts and improve conversational speech recognition on the Switchboard corpus.

Listen

This paper develops a statistical framework for automatically tagging and recognizing dialogue actssuch as statements, questions, backchannels, agreements, and apologiesin spontaneous conversational telephone speech. The work addresses the challenge of modeling discourse structure to support higher-level applications like meeting summarization and conversational agents, while also feeding constraints back to improve automatic speech recognition. The authors treat sequences of dialogue acts as a hidden Markov model whose states emit observations derived from words and prosody, and they integrate this with n-gram discourse grammars, word-based language models, decision trees, and neural networks.

They trained and evaluated the models on 1,155 hand-labeled Switchboard conversations comprising roughly 205,000 utterances. The framework was tested both with perfect word transcripts and with the output of a speech recognizer that had a 41% word error rate, and it incorporated prosodic features such as duration, pitch, and energy extracted directly from the waveform.

The combined model achieved 71% dialogue-act labeling accuracy on word transcripts and 65% accuracy when driven by recognizer output, compared with a 35% chance baseline and 84% human interlabeler agreement. Adding prosody produced consistent gains, especially on ambiguous distinctions such as backchannels versus agreements and questions versus statements; the largest relative benefit appeared when word information was imperfect. The same models yielded a small but statistically reliable reduction in word error rate during speech recognition rescoring, with the greatest improvements on short, lexically constrained acts such as answers and backchannels. Bigram discourse grammars proved nearly optimal, and alternative modeling choices such as maximum-entropy or neural-network discourse grammars added little.

These results demonstrate that a unified probabilistic treatment of lexical, prosodic, and sequential cues can extract useful discourse information from large quantities of noisy conversational data and can modestly tighten language-model constraints for recognition. The gains are limited by the heavy skew toward statements in the Switchboard domain; task-oriented dialogues with more balanced act distributions would likely see larger recognition benefits. The main limitations are the assumption of pre-segmented utterances and the conditional-independence assumptions that decouple evidence across turns and across knowledge sources.

Future work should explore joint lexical-prosodic models that relax those independence assumptions, incorporate turn-overlap timing, and investigate unsupervised discovery of dialogue-act categories optimized for the primary downstream task.

arXiv: cs/0006023
Cover for Dialogue act modeling for automatic tagging and recognition of conversational speech

Abstract

We describe a statistical approach for modeling dialogue acts in conversational speech, i.e., speech-act-like units such as Statement, Question, Backchannel, Agreement, Disagreement, and Apology. Our model detects and predicts dialogue acts based on lexical, collocational, and prosodic cues, as well as on the discourse coherence of the dialogue act sequence. The dialogue model is based on treating the discourse structure of a conversation as a hidden Markov model and the individual dialogue acts as observations emanating from the model states. Constraints on the likely sequence of dialogue acts are modeled via a dialogue act n-gram. The statistical dialogue grammar is combined with word n-grams, decision trees, and neural networks modeling the idiosyncratic lexical and prosodic manifestations of each dialogue act. We develop a probabilistic integration of speech recognition with dialogue modeling, to improve both speech recognition and dialogue act classification accuracy. Models are trained and evaluated using a large hand-labeled database of 1,155 conversations from the Switchboard corpus of spontaneous human-to-human telephone speech. We achieved good dialogue act labeling accuracy (65% based on errorful, automatically recognized words and prosody, and 71% based on word transcripts, compared to a chance baseline accuracy of 35% and human accuracy of 84%) and a small reduction in word recognition error.

Table of Contents

  • 1 Introduction
  • 2 The Dialogue Act Labeling Task
  • 2.1 Utterance Segmentation
  • 2.2 Tag Set
  • 2.3 Major Dialogue Act Types
  • 3 Hidden Markov Modeling of Dialogue
  • 3.1 Dialogue Act Likelihoods
  • 3.2 Markov Modeling
  • 3.3 Dialogue Act Decoding
  • 4 Discourse Grammars
  • 4.1 N-gram Discourse Models
  • 4.2 Other Discourse Models
  • 5 Dialogue Act Classification
  • 5.1 Dialogue Act Classification Using Words
  • 5.1.1 Classification from True Words
  • 5.1.2 Classification from Recognized Words
  • 5.1.3 Results
  • 5.2 Dialogue Act Classification Using Prosody
  • 5.2.1 Prosodic Features
  • 5.2.2 Prosodic Decision Trees
  • 5.2.3 Results with Decision Trees
  • 5.2.4 Neural Network Classifiers
  • 5.2.5 Intonation Event Likelihoods
  • 5.3 Using Multiple Knowledge Sources
  • 5.3.1 Results
  • 5.3.2 Focused Classifications
  • 6 Speech Recognition
  • 6.1 Integrating DA Modeling and ASR
  • 6.2 Computational Structure of Mixture Modeling
  • 6.3 Experiments and Results
  • 7 Prior and Related Work
  • 8 Discussion and Issues for Future Research
  • 9 Conclusions
  • References

Knowls

  1. Knowl 1 — Hidden Markov Model Formulation for Dialogue Act Tagging

    model/method

    The discourse structure of a conversation is modeled as a hidden Markov model (HMM) in which hidden states correspond to dialogue act (DA) labels U=(U1,,Un)U = (U_1, \dots, U_n) across an nn-utterance conversation, and observations correspond to utterance-level evidence E=(E1,,En)E = (E_1, \dots, E_n) (such as words or prosodic features).

    Under the assumption that utterance evidence is conditionally independent given the corresponding dialogue act, the complete conversation likelihood factors by utterance:

    P(EU)=i=1nP(EiUi)P(E \mid U) = \prod_{i=1}^n P(E_i \mid U_i)

    The prior distribution of dialogue acts is represented as a kk-th order Markov discourse grammar:

    P(UiU1,,Ui1)=P(UiUik,,Ui1)P(U_i \mid U_1, \dots, U_{i-1}) = P(U_i \mid U_{i-k}, \dots, U_{i-1})

    To minimize the total number of utterance classification errors, decoding computes the per-utterance posterior probability for each dialogue act candidate uu given all available conversational evidence EE via marginalization over all DA sequences where the ii-th label is uu:

    P(uE)=U:Ui=uP(UE)P(u \mid E) = \sum_{U: U_i = u} P(U \mid E)

    This posterior summation is computed over the whole conversation using the forward-backward dynamic programming algorithm.

  2. Knowl 2 — Dialogue Act Likelihood Estimation Over Speech Recognizer Word Hypotheses

    model/method

    For automatic dialogue act (DA) classification from speech when true word transcriptions are unavailable, the acoustic evidence AiA_i (spectral cepstral features) from an automatic speech recognition (ASR) system is incorporated by marginalizing over candidate word sequences WW:

    P(AiUi)=WP(AiW)P(WUi)P(A_i \mid U_i) = \sum_{W} P(A_i \mid W) P(W \mid U_i)

    This formulation assumes that recognizer acoustic features AiA_i are conditionally independent of the dialogue act UiU_i given the word sequence WW.

    In practice, the sum is approximated over the top NN hypotheses (e.g., N2500N \le 2500) generated by the ASR decoder for utterance ii. To account for acoustic score variances and prevent insertion/deletion imbalances, the individual summand for hypothesis WiW_i with word length Wi|W_i| is scaled using the recognizer language model weight λ\lambda and insertion penalty μ\mu:

    1λlogP(AiWi)+logP(WiUi)μλWi\frac{1}{\lambda} \log P(A_i \mid W_i) + \log P(W_i \mid U_i) - \frac{\mu}{\lambda} |W_i|

    where P(WiUi)P(W_i \mid U_i) is computed using a DA-specific statistical nn-gram language model.

  3. Knowl 3 — Integration of Prosodic Classifiers and Combined Observation Likelihood Scaling

    model/method

    Prosodic features FiF_i (including pitch statistics, pause duration, energy, and speaking rate) are mapped to dialogue acts using posterior classifiers such as CART-style decision trees or neural networks, yielding estimated posteriors P(UiFi)P(U_i \mid F_i).

    To convert classifier posteriors into likelihoods suitable for the HMM discourse framework, Bayes' rule is applied locally:

    P(FiUi)P(UiFi)P(Ui)P(F_i \mid U_i) \propto \frac{P(U_i \mid F_i)}{P(U_i)}

    where P(Ui)P(U_i) is the prior probability of dialogue act UiU_i. Alternatively, classifiers are trained on downsampled data with equal class proportions (P(Ui)P(U_i) uniform), which causes posterior outputs to be directly proportional to likelihoods.

    To integrate recognizer acoustics AiA_i, word hypotheses WiW_i, and prosodic features FiF_i while compensating for differences in model sharpness and correlations with discourse grammars, exponential scaling parameters α\alpha (prosody likelihood weight) and β\beta (overall combined likelihood dynamic range scale) are applied:

    P(Ai,Wi,FiUi){P(Ai,WiUi)P(FiUi)α}βP(A_i, W_i, F_i \mid U_i) \approx \left\{ P(A_i, W_i \mid U_i) P(F_i \mid U_i)^\alpha \right\}^\beta

    The weights α\alpha and β\beta are optimized empirically on held-out validation data.

  4. Knowl 4 — Dialogue-Act-Conditioned Speech Recognition via Mixture-of-LMs and Mixture-of-Posteriors

    model/method

    Automatic speech recognition (ASR) is conditioned on conversational discourse structure by leveraging dialogue act (DA) posterior probabilities inferred from whole-conversation evidence EE.

    In the mixture-of-posteriors formulation, the posterior probability of word sequence WiW_i given acoustic evidence AiA_i and conversational context EE is evaluated by taking a weighted sum of posterior distributions over all DA types:

    P(WiAi,E)=UiP(WiUi)P(AiWi)P(AiUi)P(UiE)P(W_i \mid A_i, E) = \sum_{U_i} \frac{P(W_i \mid U_i) P(A_i \mid W_i)}{P(A_i \mid U_i)} P(U_i \mid E)

    where P(UiE)P(U_i \mid E) is the discourse-derived posterior probability of DA UiU_i, and P(WiUi)P(W_i \mid U_i) is a DA-specific language model (LM).

    In the computationally efficient mixture-of-LMs formulation, the DA posteriors are combined directly into a sentence-level mixture prior before evaluating acoustic scores:

    P(WiAi,E)(UiP(WiUi)P(UiE))P(AiWi)P(Ai)P(W_i \mid A_i, E) \approx \left( \sum_{U_i} P(W_i \mid U_i) P(U_i \mid E) \right) \frac{P(A_i \mid W_i)}{P(A_i)}

    This enables single-pass rescoring or decoding using a dynamic, context-weighted mixture of DA language models without requiring separate recognizer passes for each DA category.

  5. Knowl 5 — Empirical Accuracy of Automatic Dialogue Act Tagging Across Evidence Sources

    data/table

    Automatic dialogue act classification performance on spontaneous conversational telephone speech (Switchboard) was evaluated using word transcripts, speech recognizer output, prosodic features, and nn-gram discourse grammars. The task uses a 42-class tag set with a chance baseline accuracy of 35.0% (predicting the majority class, STATEMENT) and human interlabeler agreement of 84.0%.

    Discourse Grammar True Words (%) Recognized Words (%) Prosody Only (%) Combined Recog.+Prosody (%)
    None (0th-order) 54.3 42.8 38.9 56.5
    Unigram 68.2 61.8 48.3 62.4
    Bigram 70.6 64.3 49.7 65.0
    Trigram 71.0 64.8

    Using recognized words from an ASR system with 41% word error rate increases classification error by 21.4% relative compared to true word transcripts (achieving 64.8% vs. 71.0% with a trigram discourse grammar). Combining prosodic likelihoods with recognized word hypotheses substantially boosts performance without discourse grammar (from 42.8% to 56.5%) and attains 65.0% accuracy with a bigram discourse grammar.

  6. Knowl 6 — Prosodic and Lexical Disambiguation in Focused Binary Dialogue Act Classification

    data/table

    To analyze prosodic classification independently of discourse grammar and corpus class skew, binary classification experiments were conducted on balanced subsets (chance baseline = 50.0%) targeting pairs frequently confused by lexical models: Questions vs. Statements and Agreements vs. Backchannels.

    Classification Task Knowledge Source True Words (%) Recognized Words (%)
    Questions / Statements Prosody only 76.0 76.0
    Words only 85.9 75.4
    Words + Prosody 87.6 79.8
    Agreements / Backchannels Prosody only 72.9 72.9
    Words only 81.0 78.2
    Words + Prosody 84.7 81.7

    Prosody alone provides strong discriminative signal (76.0% and 72.9% accuracy). Combining prosodic features with word likelihoods yields statistically significant accuracy improvements over words alone (p<0.001p < 0.001), with larger relative error reductions on recognized words than on true transcripts (e.g., improving Question vs. Statement accuracy on recognized words from 75.4% to 79.8%).

  7. Knowl 7 — Discourse Grammar Sequence Perplexity and Speaker Turn Dynamics

    data/table

    Statistical discourse grammars model dialogue act (DA) transition dynamics. Perplexities were measured across nn-gram orders under three conditions: DA sequences alone P(U)P(U), joint DA and speaker identity sequences P(U,T)P(U, T) over 84 tags (42 DA tags ×\times 2 conversants), and DA sequences conditioned on known speaker identity P(UT)P(U \mid T).

    Discourse Grammar Order P(U)P(U) P(U,T)P(U, T) P(UT)P(U \mid T)
    None (Uniform, 0th-order) 42.0 84.0 42.0
    Unigram 11.0 18.5 9.0
    Bigram 7.9 10.4 5.1
    Trigram 7.5 9.8 4.8

    Conditioning on speaker turns sharply reduces DA perplexity (e.g., from 7.9 to 5.1 for bigrams). Bigram discourse grammars capture the primary local sequential constraints (such as Question-Answer and Statement-Backchannel adjacency pairs), while trigram models yield only minor incremental gains. Long-distance cache and maximum entropy models offer no significant improvement over standard backoff nn-grams.

  8. Knowl 8 — Speech Recognition Error Rate and Language Model Perplexity Under Dialogue Act Conditioning

    data/table

    ASR rescoring experiments evaluated the impact of dialogue-act-conditioned language modeling on the Switchboard test set (19 conversations, 4,000 utterances, 29,000 words) using NN-best hypothesis lists (N2500N \le 2500).

    Model WER (%) LM Perplexity
    Baseline (standard trigram LM) 41.2 76.8
    1-best LM 41.0 69.3
    Mixture-of-posteriors 41.0 n/a
    Mixture-of-LMs 40.9 66.9
    Oracle LM (hand-labeled DA) 40.3 66.8

    Conditioning on inferred dialogue acts achieves a 13% relative reduction in language model perplexity (from 76.8 to 66.9 for Mixture-of-LMs, matching the Oracle LM perplexity of 66.8). However, the corresponding word error rate (WER) reduction is modest: Mixture-of-LMs reduces WER by 0.3% absolute (from 41.2% to 40.9%), while even the Oracle LM with perfect dialogue act knowledge reduces WER by only 0.9% absolute (to 40.3%).

  9. Knowl 9 — Disproportionate ASR Word Error Reductions Across Dialogue Act Categories

    data/table

    Word error rate (WER) reductions obtained from rescoring with an Oracle DA language model vary dramatically across individual dialogue act categories in the Switchboard test corpus.

    Dialogue Act Type Baseline WER (%) Oracle WER (%) Absolute WER Reduction (%)
    NO-ANSWER 29.4 11.8 -17.6
    BACKCHANNEL 25.9 18.6 -7.3
    BACKCHANNEL-QUESTION 15.2 9.1 -6.1
    ABANDONED/UNINTERPRETABLE 48.9 45.2 -3.7
    WH-QUESTION 38.4 34.9 -3.5
    YES-NO-QUESTION 55.5 52.3 -3.2
    STATEMENT 42.0 41.5 -0.5
    OPINION 40.8 40.4 -0.4

    Short and structurally constrained dialogue acts exhibit large error reductions (e.g., -17.6% for NO-ANSWER, -7.3% for BACKCHANNEL). However, STATEMENTS and OPINIONS, which account for 83% of all words in the corpus, show negligible improvement (-0.5% and -0.4%), explaining why the aggregate WER improvement across the full corpus remains small.

  10. Knowl 10 — Switchboard SWBD-DAMSL Dialogue Act Corpus and Annotation Scheme

    experimental setup

    The Switchboard corpus of spontaneous human-to-human telephone conversations was annotated with dialogue acts using a modified version of the Dialogue Act Markup in Several Layers (DAMSL) standard. The dataset comprises 1,155 conversations totaling 205,000 utterances and 1.4 million words, split into a training set of 1,115 conversations (198,000 utterances, 1.4M words) and a test set of 19 conversations (4,000 utterances, 29,000 words).

    Utterances are defined as sentence-level segmentation units rather than speaker turns. A single turn may comprise multiple utterances, and utterances can span across turns during backchannel interruptions. Approximately 220 fine-grained tag combinations were mapped to a mutually exclusive set of 42 dialogue act classes (e.g., STATEMENT 36%, BACKCHANNEL/ACKNOWLEDGE 19%, OPINION 13%, ABANDONED 6%, AGREEMENT/ACCEPT 5%, YES-NO-QUESTION 2%). Interlabeler agreement on the 42-tag set was 84%, yielding a chance-normalized Kappa statistic of κ=0.80\kappa = 0.80.

Coverage note — Omitted secondary exploratory prosodic modeling variants (intonation event continuous HMMs and alternative neural network hidden layer configurations) because they yielded no improvements over decision trees and did not alter the main findings.

References

  1. 1.Alexandersson, Jan and Norbert Reithinger. 1997. Learning dialogue structures from a corpus. In G. Kokkinakis, N. Fakotakis, and E. Dermatas, editors, Proceedings of the 5th European Conference on Speech Communication and Technology, volume 4, pages 2231–2234, Rhodes, Greece, September.
  2. 2.Anderson, Anne H., Miles Bader, Ellen G. Bard, Elizabeth H. Boyle, Gwyneth M. Doherty, Simon C. Garrod, Stephen D. Isard, Jacqueline C. Kowtko, Jan M. McAllister, Jim Miller, Catherine F. Sotillo, Henry S. Thompson, and Regina Weinert. 1991. The HCRC Map Task corpus. Language and Speech, 34(4):351–366.
  3. 3.Austin, J. L. 1962. How to do Things with Words. Clarendon Press, Oxford.
  4. 4.Bahl, Lalit R., Frederick Jelinek, and Robert L. Mercer. 1983. A maximum likelihood approach to continuous speech recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 5(2):179–190, March.
  5. 5.Bard, Ellen G., Catherine Sotillo, Anne H. Anderson, and M. M. Taylor. 1995. The DCIEM Map Task corpus: Spontaneous dialogues under sleep deprivation and drug treatment. In Isabel Trancoso and Roger Moore, editors, Proceedings of the ESCA-NATO Tutorial and Workshop on Speech under Stress, pages 25–28, Lisbon, September.
  6. 6.Baum, Leonard E., Ted Petrie, George Soules, and Norman Weiss. 1970. A maximization technique occurring in the statistical analysis of probabilistic functions in Markov chains. The Annals of Mathematical Statistics, 41(1):164–171.
  7. 7.Berger, Adam L., Stephen A. Della Pietra, and Vincent J. Della Pietra. 1996. A maximum entropy approach to natural language processing. Computational Linguistics, 22(1):39–71.
  8. 8.Bourlard, Hervé and Nelson Morgan. 1993. Connectionist Speech Recognition. A Hybrid Approach. Kluwer Academic Publishers, Boston, MA.
  9. 9.Breiman, L., J. H. Friedman, R. A. Olshen, and C. J. Stone. 1984. Classification and Regression Trees. Wadsworth and Brooks, Pacific Grove, CA.
  10. 10.Bridle, J. S. 1990. Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition. In F. Fogleman Soulie and J. Herault, editors, Neurocomputing: Algorithms, Architectures and Applications. Springer, Berlin, pages 227–236.
  11. 11.Brill, Eric. 1993. Automatic grammar induction and parsing free text: A transformation-based approach. In Proceedings of the ARPA Workshop on Human Language Technology, Plainsboro, NJ, March.
  12. 12.Carletta, Jean. 1996. Assessing agreement on classification tasks: The Kappa statistic. Computational Linguistics, 22(2):249–254.
  13. 13.Carlson, Lari. 1983. Dialogue Games: An Approach to Discourse Analysis. D. Reidel.
  14. 14.Chu-Carroll, Jennifer. 1998. A statistical model for discourse act recognition in dialogue interactions. In Jennifer Chu-Carroll and Nancy Green, editors, Applying Machine Learning to Discourse Processing. Papers from the 1998 AAAI Spring Symposium. Technical Report SS-98-01, pages 12–17. AAAI Press, Menlo Park, CA.
  15. 15.Church, Kenneth Ward. 1988. A stochastic parts program and noun phrase parser for unrestricted text. In Second Conference on Applied Natural Language Processing, pages 136–143, Austin, TX.
  16. 16.Core, Mark and James Allen. 1997. Coding dialogs with the DAMSL annotation scheme. In Working Notes of the AAAI Fall Symposium on Communicative Action in Humans and Machines, pages 28–35, Cambridge, MA, November.
  17. 17.Dermatas, Evangelos and George Kokkinakis. 1995. Automatic stochastic tagging of natural language texts. Computational Linguistics, 21(2):137–163.
  18. 18.Eckert, Wieland, Florian Gallwitz, and Heinrich Niemann. 1996. Combining stochastic and linguistic language models for recognition of spontaneous speech. In Proceedings of the IEEE Conference on Acoustics, Speech, and Signal Processing, volume 1, pages 423–426, Atlanta, May.
  19. 19.Finke, Michael, Maria Lapata, Alon Lavie, Lori Levin, Laura Mayfield Tomokiyo, Thomas Polzin, Klaus Ries, Alex Waibel, and Klaus Zechner. 1998. Clarity: Inferring discourse structure from speech. In Jennifer Chu-Carroll and Nancy Green, editors, Applying Machine Learning to Discourse Processing. Papers from the 1998 AAAI Spring Symposium. Technical Report SS-98-01, pages 25–32. AAAI Press, Menlo Park, CA.
  20. 20.Fowler, Carol A. and Jonathan Housum. 1987. Talkers’ signaling of “new” and “old” words in speech and listeners’ perception and use of the distinction. Journal of Memory and Language, 26:489–504.
  21. 21.Fukada, Toshiaki, Detlef Koll, Alex Waibel, and Kouichi Tanigaki. 1998. Probabilistic dialogue act extraction for concept based multilingual translation systems. In Robert H. Mannell and Jordi Robert-Ribes, editors, Proceedings of the International Conference on Spoken Language Processing, volume 6, pages 2771–2774, Sydney, December. Australian Speech Science and Technology Association.
  22. 22.Godfrey, J. J., E. C. Holliman, and J. McDaniel. 1992. SWITCHBOARD: Telephone speech corpus for research and development. In Proceedings of the IEEE Conference on Acoustics, Speech, and Signal Processing, volume 1, pages 517–520, San Francisco, March.
  23. 23.Grosz, B. and C. Sidner. 1986. Attention, intention, and the structure of discourse. Computational Linguistics, 12(3):175–204.
  24. 24.Hirschberg, Julia B. and Diane J. Litman. 1993. Empirical studies on the disambiguation of cue phrases. Computational Linguistics, 19(3):501–530.
  25. 25.Iyer, Rukmini, Mari Ostendorf, and J. Robin Rohlicek. 1994. Language modeling with sentence-level mixtures. In Proceedings of the ARPA Workshop on Human Language Technology, pages 82–86, Plainsboro, NJ, March.
  26. 26.Jefferson, Gail. 1984. Notes on a systematic deployment of the acknowledgement tokens ‘yeah’ and ‘mm hm’. Papers in Linguistics, 17:197–216.
  27. 27.Jekat, Susanne, Alexandra Klein, Elisabeth Maier, Ilona Maleck, Marion Mast, and Joachim Quantz. 1995. Dialogue acts in VERBMOBIL. Verbmobil-Report 65, Universität Hamburg, DFKI GmbH, Universitẗ Erlangen, and TU Berlin, April.
  28. 28.Jurafsky, Dan, Rebecca Bates, Noah Coccaro, Rachel Martin, Marie Meteer, Klaus Ries, Elizabeth Shriberg, Andreas Stolcke, Paul Taylor, and Carol Van Ess-Dykema. 1997. Automatic detection of discourse structure for speech recognition and understanding. In Proceedings IEEE Workshop on Speech Recognition and Understanding, pages 88–95, Santa Barbara, CA, December.
  29. 29.Jurafsky, Daniel, Rebecca Bates, Noah Coccaro, Rachel Martin, Marie Meteer, Klaus Ries, Elizabeth Shriberg, Andreas Stolcke, Paul Taylor, and Carol Van Ess-Dykema. 1998. Switchboard discourse language modeling project final report. Research Note 30, Center for Language and Speech Processing, Johns Hopkins University, Baltimore, January.
  30. 30.Jurafsky, Daniel, Elizabeth Shriberg, and Debra Biasca. 1997. Switchboard-DAMSL Labeling Project Coder’s Manual. Technical Report 97-02, University of Colorado, Institute of Cognitive Science, Boulder, CO. http://www.colorado.edu/ling/jurafsky/manual.august1.html.
  31. 31.Jurafsky, Daniel, Elizabeth E. Shriberg, Barbara Fox, and Traci Curl. 1998. Lexical, prosodic, and syntactic cues for dialog acts. In Proceedings of ACL/COLING-98 Workshop on Discourse Relations and Discourse Markers, pages 114–120. Association for Computational Linguistics.
  32. 32.Katz, Slava M. 1987. Estimation of probabilities from sparse data for the language model component of a speech recognizer. IEEE Transactions on Acoustics, Speech, and Signal Processing, 35(3):400–401, March.
  33. 33.Kita, Kenji, Yoshikazu Fukui, Masaaki Nagata, and Tsuyoshi Morimoto. 1996. Automatic acquisition of probabilistic dialogue models. In H. Timothy Bunnell and William Idsardi, editors, Proceedings of the International Conference on Spoken Language Processing, volume 1, pages 196–199, Philadelphia, October.
  34. 34.Kompe, Ralf. 1997. Prosody in speech understanding systems. Springer, Berlin.
  35. 35.Kowtko, Jacqueline C. 1996. The Function of Intonation in Task Oriented Dialogue. Ph.D. thesis, University of Edinburgh, Edinburgh.
  36. 36.Kuhn, Roland and Renato de Mori. 1990. A cache-base natural language model for speech recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(6):570–583, June.
  37. 37.Levin, Joan A. and Johanna A. Moore. 1977. Dialogue games: Metacommunication structures for natural language interaction. Cognitive Science, 1(4):395–420.
  38. 38.Levin, Lori, Klaus Ries, Ann Thymé-Gobbel, and Alon Lavie. 1999. Tagging of speech acts and dialogue games in Spanish CallHome. In Towards Standards and Tools for Discourse Tagging (Proceedings of the Workshop at ACL’99), pages 42–47, College Park, MD, June.
  39. 39.Linell, Per. 1990. The power of dialogue dynamics. In Ivana Marková and Klaus Foppa, editors, The Dynamics of Dialogue. Harvester, Wheatsheaf, New York, London, pages 147–177.
  40. 40.Mast, M., R. Kompe, S. Harbeck, A. Kießling, H. Niemann, E. Nöth, E. G. Schukat-Talamazzini, and V. Warnke. 1996. Dialog act classification with the help of prosody. In H. Timothy Bunnell and William Idsardi, editors, Proceedings of the International Conference on Spoken Language Processing, volume 3, pages 1732–1735, Philadelphia, October.
  41. 41.Menn, Lise and Suzanne E. Boyce. 1982. Fundamental frequency and discourse structure. Language and Speech, 25:341–383.
  42. 42.Meteer, Marie, Ann Taylor, Robert MacIntyre, and Rukmini Iyer. 1995. Dysfluency annotation stylebook for the Switchboard corpus. Distributed by LDC, ftp://ftp.cis.upenn.edu/pub/treebank/swbd/doc/DFL-book.ps, February. Revised June 1995 by Ann Taylor.
  43. 43.Morgan, Nelson, Eric Fosler, and Nikki Mirghafori. 1997. Speech recognition using on-line estimation of speaking rate. In G. Kokkinakis, N. Fakotakis, and E. Dermatas, editors, Proceedings of the 5th European Conference on Speech Communication and Technology, volume 4, pages 2079–2082, Rhodes, Greece, September.
  44. 44.Munk, Marcus. 1999. Shallow Statistical Parsing for Machine Translation. Diploma thesis, Carnegie Mellon University.
  45. 45.Nagata, Masaaki. 1992. Using pragmatics to rule out recognition errors in cooperative task-oriented dialogues. In John J. Ohala, Terrance M. Nearey, Bruce L. Derwing, Megan M. Hodge, and Grace E. Wiebe, editors, Proceedings of the International Conference on Spoken Language Processing, volume 1, pages 647–650, Banff, Canada, October.
  46. 46.Nagata, Masaaki and Tsuyoshi Morimoto. 1993. An experimental statistical dialogue model to predict the speech act type of the next utterance. In Katsuhiko Shirai, Tetsunori Kobayashi, and Yasunari Harada, editors, Proceedings of the International Symposium on Spoken Dialogue, pages 83–86, Tokyo, November.
  47. 47.Nagata, Masaaki and Tsuyoshi Morimoto. 1994. First steps toward statistical modeling of dialogue to predict the speech act type of the next utterance. Speech Communication, 15:193–203.
  48. 48.Ohler, Uwe, Stefan Harbeck, and Heinrich Niemann. 1999. Discriminative training of language model classifiers. In Proceedings of the 6th European Conference on Speech Communication and Technology, volume 4, pages 1607–1610, Budapest, September.
  49. 49.Pearl, Judea. 1988. Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. Morgan Kaufmann, San Mateo, CA.
  50. 50.Power, Richard J. D. 1979. The organisation of purposeful dialogues. Linguistics, 17:107–152.
  51. 51.Rabiner, L. R. and B. H. Juang. 1986. An introduction to hidden Markov models. IEEE ASSP Magazine, 3(1):4–16, January.
  52. 52.Reithinger, Norbert, Ralf Engel, Michael Kipp, and Martin Klesen. 1996. Predicting dialogue acts for a speech-to-speech translation system. In H. Timothy Bunnell and William Idsardi, editors, Proceedings of the International Conference on Spoken Language Processing, volume 2, pages 654–657, Philadelphia, October.
  53. 53.Reithinger, Norbert and Martin Klesen. 1997. Dialogue act classification using language models. In G. Kokkinakis, N. Fakotakis, and E. Dermatas, editors, Proceedings of the 5th European Conference on Speech Communication and Technology, volume 4, pages 2235–2238, Rhodes, Greece, September.
  54. 54.Ries, Klaus. 1999a. HMM and neural network based speech act classification. In Proceedings of the IEEE Conference on Acoustics, Speech, and Signal Processing, volume 1, pages 497–500, Phoenix, AZ, March.
  55. 55.Ries, Klaus. 1999b. Towards the detection and description of textual meaning indicators in spontaneous conversations. In Proceedings of the 6th European Conference on Speech Communication and Technology, volume 3, pages 1415–1418, Budapest, September.
  56. 56.Rose, R. C., E. I. Chang, and R. P. Lippmann. 1991. Techniques for information retrieval from voice messages. In Proceedings of the IEEE Conference on Acoustics, Speech, and Signal Processing, volume 1, pages 317–320, Toronto, May.
  57. 57.Sacks, H., E. A. Schegloff, and G. Jefferson. 1974. A simplest semantics for the organization of turn-taking in conversation. Language, 50(4):696–735.
  58. 58.Samuel, Ken, Sandra Carberry, and K. Vijay-Shanker. 1998. Dialogue act tagging with transformation-based learning. In Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics, volume 2, pages 1150–1156, Montreal.
  59. 59.Schegloff, E. A. 1968. Sequencing in conversational openings. American Anthropologist, 70:1075–1095.
  60. 60.Schegloff, Emanuel A. 1982. Discourse as an interactional achievement: Some uses of ‘uh huh’ and other things that come between sentences. In Deborah Tannen, editor, Analyzing Discourse: Text and Talk. Georgetown University Press, Washington, D.C., pages 71–93.
  61. 61.Searle, J. R. 1969. Speech Acts. Cambridge University Press, London-New York.
  62. 62.Shriberg, Elizabeth, Rebecca Bates, Andreas Stolcke, Paul Taylor, Daniel Jurafsky, Klaus Ries, Noah Coccaro, Rachel Martin, Marie Meteer, and Carol Van Ess-Dykema. 1998. Can prosody aid the automatic classification of dialog acts in conversational speech? Language and Speech, 41(3-4):439–487.
  63. 63.Shriberg, Elizabeth, Andreas Stolcke, Dilek Hakkani-Tür, and Gökhan Tür. 2000. Prosody-based automatic segmentation of speech into sentences and topics. Speech Communication, 32(1-2):to appear, September. Special Issue on Accessing Information in Spoken Audio.
  64. 64.Siegel, Sidney and N. John Castellan, Jr. 1988. Nonparametric Statistics for the Behavioral Sciences. McGraw-Hill, New York, second edition.
  65. 65.Stolcke, Andreas and Elizabeth Shriberg. 1996. Automatic linguistic segmentation of conversational speech. In H. Timothy Bunnell and William Idsardi, editors, Proceedings of the International Conference on Spoken Language Processing, volume 2, pages 1005–1008, Philadelphia, October.
  66. 66.Suhm, B. and A. Waibel. 1994. Toward better language models for spontaneous speech. In Proceedings of the International Conference on Spoken Language Processing, volume 2, pages 831–834, Yokohama, September.
  67. 67.Taylor, Paul, Simon King, Stephen Isard, and Helen Wright. 1998. Intonation and dialog context as constraints for speech recognition. Language and Speech, 41(3-4):489–508.
  68. 68.Taylor, Paul, Simon King, Stephen Isard, Helen Wright, and Jacqueline Kowtko. 1997. Using intonation to constrain language models in speech recognition. In G. Kokkinakis, N. Fakotakis, and E. Dermatas, editors, Proceedings of the 5th European Conference on Speech Communication and Technology, volume 5, pages 2763–2766, Rhodes, Greece, September.
  69. 69.Taylor, Paul A. 2000. Analysis and synthesis of intonation using the tilt model. Journal of the Acoustical Society of America, 107(3):1697–1714.
  70. 70.Van Ess-Dykema, Carol and Klaus Ries. 1998. Linguistically engineered tools for speech recognition error analysis. In Robert H. Mannell and Jordi Robert-Ribes, editors, Proceedings of the International Conference on Spoken Language Processing, volume 5, pages 2091–2094, Sydney, December. Australian Speech Science and Technology Association.
  71. 71.Viterbi, A. 1967. Error bounds for convolutional codes and an asymptotically optimum decoding algorithm. IEEE Transactions on Information Theory, 13:260–269.
  72. 72.Warnke, V., R. Kompe, H. Niemann, and E. Nöth. 1997. Integrated dialog act segmentation and classification using prosodic features and language models. In G. Kokkinakis, N. Fakotakis, and E. Dermatas, editors, Proceedings of the 5th European Conference on Speech Communication and Technology, volume 1, pages 207–210, Rhodes, Greece, September.
  73. 73.Warnke, Volker, Stefan Harbeck, Elmar Nöth, Heinrich Niemann, and Michael Levit. 1999. Discriminative estimation of interpolation parameters for language model classifiers. In Proceedings of the IEEE Conference on Acoustics, Speech, and Signal Processing, volume 1, pages 525–528, Phoenix, AZ, March.
  74. 74.Weber, Elizabeth G. 1993. Varieties of Questions in English Conversation. John Benjamins, Amsterdam.
  75. 75.Witten, Ian H. and Timothy C. Bell. 1991. The zero-frequency problem: Estimating the probabilities of novel events in adaptive text compression. IEEE Transactions on Information Theory, 37(4):1085–1094, July.
  76. 76.Woszczyna, M. and A. Waibel. 1994. Inferring linguistic structure in spoken language. In Proceedings of the International Conference on Spoken Language Processing, volume 2, pages 847–850, Yokohama, September.
  77. 77.Wright, Helen. 1998. Automatic utterance type detection using suprasegmental features. In Robert H. Mannell and Jordi Robert-Ribes, editors, Proceedings of the International Conference on Spoken Language Processing, volume 4, pages 1403–1406, Sydney, December. Australian Speech Science and Technology Association.
  78. 78.Wright, Helen and Paul A. Taylor. 1997. Modelling intonational structure using hidden Markov models. In Intonation: Theory, Models and Applications. Proceedings of an ESCA Workshop, pages 333–336, Athens, September.
  79. 79.Yngve, Victor H. 1970. On getting a word in edgewise. In Papers from the Sixth Regional Meeting of the Chicago Linguistic Society, pages 567–577, Chicago, April. University of Chicago.
  80. 80.Yoshimura, Takashi, Satoru Hayamizu, Hiroshi Ohmura, and Kazuyo Tanaka. 1996. Pitch pattern clustering of user utterances in human-machine dialogue. In H. Timothy Bunnell and William Idsardi, editors, Proceedings of the International Conference on Spoken Language Processing, volume 2, pages 837–840, Philadelphia, October.

Citation

MLA
Stolcke, A., et al. “Dialogue Act Modeling for Automatic Tagging and Recognition of Conversational Speech”. Computational Linguistics, 2000, pp. 339–74, https://aclanthology.org/J00-3003/.
APA
Stolcke, A., Ries, K., Coccaro, N., Shriberg, E., Bates, R., Jurafsky, D., Taylor, P., Martin, R., Ess-Dykema, C. V., & Meteer, M. (2000). Dialogue act modeling for automatic tagging and recognition of conversational speech. Computational Linguistics, 339–374. https://aclanthology.org/J00-3003/
Chicago
Stolcke, A., K. Ries, N. Coccaro, et al. 2000. “Dialogue Act Modeling for Automatic Tagging and Recognition of Conversational Speech”. Computational Linguistics, 339–74. https://aclanthology.org/J00-3003/.
Harvard
Stolcke, A. et al. (2000) “Dialogue act modeling for automatic tagging and recognition of conversational speech”, Computational Linguistics. Association for Computational Linguistics, pp. 339–374. Available at: https://aclanthology.org/J00-3003/.
Vancouver
1. Stolcke A, Ries K, Coccaro N, Shriberg E, Bates R, Jurafsky D, Taylor P, Martin R, Ess-Dykema CV, Meteer M (2000) Dialogue act modeling for automatic tagging and recognition of conversational speech. In: Computational Linguistics. Association for Computational Linguistics, pp 339–374

BibTeX

@article{stolcke-etal-2000-dialogue,
    title = "Dialogue act modeling for automatic tagging and recognition of conversational speech",
    author = "Stolcke, Andreas  and
      Ries, Klaus  and
      Coccaro, Noah  and
      Shriberg, Elizabeth  and
      Bates, Rebecca  and
      Jurafsky, Daniel  and
      Taylor, Paul  and
      Martin, Rachel  and
      Van Ess-Dykema, Carol  and
      Meteer, Marie",
    journal = "Computational Linguistics",
    volume = "26",
    number = "3",
    year = "2000",
    address = "Cambridge, MA",
    publisher = "MIT Press",
    url = "https://aclanthology.org/J00-3003/",
    pages = "339--374"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF