Dissertations, Theses, and Capstone Projects
Date of Degree
9-2026
Document Type
Doctoral Dissertation
Degree Name
Doctor of Philosophy
Program
Computer Science
Advisor
Sarah Ita Levitan
Committee Members
Julia Bell Hirschberg
Raj Korpan
Rivka Levitan
Abstract
As virtual agents become increasingly integrated into everyday communication, trust has become essential for effective human-AI interaction. Although recent advances in natural language generation and speech synthesis have significantly improved fluency and intelligibility, most systems are not designed to control how trustworthy they are perceived. As a result, a mismatch can emerge between perceived trustworthiness and actual system capability, leading to over-trust or under-trust. While trust is critical for meaningful interaction, current text-to-speech (TTS) systems are primarily optimized for naturalness and intelligibility rather than calibrated trustworthiness.
This dissertation investigates the perceptual foundations of trust in both human and synthesized speech through the study of acoustic prosody, lexical style, and speaker/listener dynamics. Across multiple studies, the findings show that trust perception is highly context-dependent and shaped by the interaction of vocal, linguistic, and multimodal cues. Rather than treating trust as a property to be uniformly maximized, this work frames trust as a perceptual signal that should be calibrated to align with a system’s reliability, uncertainty, and communicative intent.
Building on these insights, this dissertation introduces computational frameworks for modeling and generating trustworthy speech. The work first characterizes trust-related acoustic patterns in human speech, then extends the analysis to synthesized speech to identify both shared and modality-specific trust cues. Finally, a trust-aware speech generation framework is proposed that combines representation learning and reinforcement learning to generate speech conditioned on target trust levels, enabling adaptive control of speaking style and prosody.
Overall, this dissertation advances a computational framework for understanding and generating calibrated trust in speech-based virtual agents. By connecting perceptual trust modeling with controllable speech generation, this work contributes toward intelligent virtual agents capable of dynamically adapting their communicative strategies to interact in ways that are not only natural and expressive, but also appropriately trustworthy across diverse users and interaction contexts.
Recommended Citation
Yu, Yuwen, "Modeling and Generating Trustworthy Speech" (2026). CUNY Academic Works.
https://academicworks.cuny.edu/gc_etds/6841
