Dissertations, Theses, and Capstone Projects

Date of Degree

9-2026

Document Type

Doctoral Dissertation

Degree Name

Doctor of Philosophy

Program

Computer Science

Advisor

Rebecca Levitan

Committee Members

Julia Hirschberg

Raj Korpan

Kyle Gorman

Subject Categories

Artificial Intelligence and Robotics

Keywords

entrainment, discourse analysis, machine learning, dialogue systems

Abstract

When people converse, they adjust their speech in response to one another in a complex set of behaviors known as entrainment. Its presence leads to positive conversational outcomes, but it is difficult to model in spoken dialogue systems because people entrain in highly variable ways, and fundamental limiting assumptions constrain studies to predetermined behaviors. Consequently, experimental systems entrain to their users uniformly with simpler behaviors than in real dialogue. Recent deep learning models show promise in capturing complex nonlinear entrainment, but most do not account for individual variation. Additionally, they retain assumptions limiting them to turn exchanges, despite growing evidence that entrainment occurs across broad units of discourse related to the functional role of speech.

In this thesis, we present a deep learning architecture that models arbitrary conversational behaviors without traditional limiting assumptions. Our model uses an attention layer to identify previous dialogue turns most relevant for predicting upcoming acoustic-prosodic speech features. We demonstrate that the attention scores are meaningful, interpretable, and can segment conversations around topic shifts in a completely unsupervised way. Furthermore, we use our segmentation to test and confirm hypotheses about the relationship between entrainment and discourse structure. We introduce an unsupervised soft clustering module which captures individual interaction styles and can influence model behavior. Finally, we perform a comprehensive evaluation of several model variants to understand how each change impacts predictive performance and dialogue segmentation.

This work is embargoed and will be available for download on Wednesday, March 31, 2027

Share

COinS