
Chibuzor Okocha
I am a PhD student in Computer Science at the University of Florida and a member of the UF DataStudio Lab. I am currently on the job market and open to Research Scientist roles in Speech and Audio AI.
My research centers on Speech and Audio AI, with extensive experience in training multimodal and tool-calling audio models. I focus on advancing the reasoning capabilities of Audio Language Models and developing robust, inclusive systems for accented and multilingual speech processing.
I am passionate about building open science communities, mentoring aspiring AI researchers, and democratizing access to cutting-edge speech technology across diverse languages and cultures.
Recent News
Excited to share that I have 2 papers accepted to IEEE SLT 2026!
Busy month! Gave a workshop talk at the ACM Tapia conference and had a paper accepted to Interspeech.
Successfully completed my summer research internship at the Adobe Speech AI lab in San Francisco, capping it off with a talk to the CAVA group!
Received best paper award at ASRU 2025 Workshop on Childspeech for "Can large audio language models understand child stuttering speech?"
AfriVox paper accepted at EACL 2026!
Presenting poster "Can Large Audio Language Models Understand Child Stuttering Speech? Speech Summarization, and Source Separation" at ASRU 2025 Satellite Workshop in Hawaii.
Excited to present my research at the TTIC Summer Workshop on Foundations of Speech and Audio Foundation Models in Chicago.
Excited to present our AfriSpeech-Dialog work at NAACL 2025.
Research Areas
Speech and Audio AI
Developing advanced AI systems for speech and audio processing applications
Audio Language Models
Researching reasoning capabilities and cognitive processes in audio language models
Accented Speech Recognition
Building robust speech recognition systems for diverse accents and dialects
Multilingual Audio AI
Creating inclusive AI systems that work across multiple languages and cultures
Recent Publications
Afrispeech Semantics: Evaluating Audio–Semantic Reasoning in Spoken Language Models Across Domains and Accents
Presented in San Diego, July 2026 | Chibuzor Okocha, et al.
Evaluating audio-semantic reasoning capabilities of spoken language models across different domains and African accents.
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond
NAACL 2025 | Mardhiyah Sanni, Tassallah Abdullahi, Devendra D. Kayande, Emmanuel Ayodele, Naome A. Etori, Michael S. Mollel, Moshood Yekini, Chibuzor Okocha, et al.
A comprehensive dataset for evaluating ASR and summarization on African-accented speech conversations.
AfriVox: Probing Multilingual and Accent Robustness of Speech LLMs
EACL 2026 | OpenReview
Open-source benchmark across 20 African languages and 100+ African English accents, evaluating multimodal speech LLMs vs traditional ASR/AST models.
How Robust Are Neural Audio Codecs for African Speech? A Multi-Task Benchmark and the Limits of Perceptual Quality
IEEE SLT 2026 | Chibuzor Okocha, Christan Earl Grant
Comprehensive evaluation framework for neural audio codecs on African speech data and low-resource language settings.
Beyond Word Error Rate: A Switch-Aware Evaluation of ASR and Audio Language Models on English–Yoruba Code-Switched Speech
IEEE SLT 2026 | Chibuzor Okocha, Christan Earl Grant
An evaluation of ASR and Audio Language Models on English-Yoruba code-switched speech, moving beyond traditional WER metrics.
Reasoning Beyond Transcription: Audio Language Models on Child Stuttering Speech
Interspeech 2026 | Chibuzor Okocha
Exploring reasoning capabilities and cognitive processes in audio language models applied to child stuttering speech.
Can large audio language models understand child stuttering speech?
ASRU 2025 Workshop on Childspeech (Best Paper Award) | Chibuzor Okocha, Maya Bakri, Christan Grant | arXiv
Evaluating LALMs on disfluent child speech for source separation and summarization tasks.
Domain-Aware Speaker Diarization On African-Accented English
arXiv preprint (Under Review) | Chibuzor Okocha, Kelechi Ezema, Christan Grant | arXiv
Examining domain effects in speaker diarization for African-accented English across general and clinical dialogues.