Zackary Rackauckas

zcr2105@columbia.edu

NLP + Speech Researcher — Converstional Agents for Learning

My research focuses on speech-first AI systems for language learning. I create adaptive agents that combine ASR, expressive TTS, dialogue modeling, retrieval, and user-centered interaction design. I am especially interested in how spoken interfaces can support language acquisition, pronunciation feedback, and learner engagement by making linguistic variation, affect, register, and feedback audible in interaction. I am currently a research associate at UC Irvine working with Dr. Mark Warschauer and a software engineer at Harvard GSE working with Ying Xu, where I develop child-facing AI systems for speech, phoneme, and pronunciation feedback. I also created Jouzu, a Japanese learning platform built around voiced conversational agents, expressive character speech, and in-context learner scaffolding.

I received my M.S. in Computer Science (Thesis Track) at Columbia Engineering, advised by Julia Hirschberg. I was also a special graduate research student at the University of Tokyo where I collaborated with Nobuaki Minematsu and Yuka Akiyama. I received my B.A. from Swarthmore College where I worked with John Bundschuh. I also collaborated with Daniele Struppa and Erik Linstead at Chapman University.

My work focuses on speech-to-speech conversational systems, expressive multilingual text-to-speech, and multimodal retrieval and generation for adaptive, affective dialogue agents.

Pronunciation: /rəˈkɔkəs/

You can find me on LinkedIn, Google Scholar, and GitHub. You can also take a look at my CV.

Selected Publications

  • VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering
    MAGMaR @ ACL 2025.
    Paper

  • Re:Member: Emotional Question Generation from Personal Memories
    HCI+NLP @ EMNLP 2025.
    Paper

  • Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
    IEEE UEMCON 2025.
    Paper

  • Evaluating RAG-Fusion with RAGElo: An Automated Elo-Based Framework
    LLM4Eval @ SIGIR 2024.
    Paper

See the full list on the Publications page.

Selected Systems

  • Jouzu — LLM-driven character learning platform.
    Investigates how stylized, voice-enabled conversational agents influence learner engagement and dialogue-based second-language acquisition.
    Website

  • VoxRAG — Transcription-free spoken question answering (ACL 2025).
    Explores retrieval-augmented generation directly over speech representations, reducing reliance on full ASR pipelines for spoken QA.
    GitHub

  • AdaptLingo — Adaptive speech-to-speech dialogue agent.
    Studies proficiency-aware conversational generation by integrating prosody detection, constrained decoding, and end-to-end speech modeling.
    GitHub

  • RAG Product Assistant (Infineon) — Enterprise retrieval system.
    Examines how retrieval-augmented LLM systems can improve technical knowledge access and decision efficiency in real-world engineering workflows.
    (Code proprietary)

See the full technical breakdown on the Systems page.

Contact