ILM AI
← Back to research
ResearchJul 08, 2026PROJECT SONUS · SPEECH TO SPEECH PEDAGOGY

Real Time Conversational Phonetics

A voice-native tutor that listens, speaks, and coaches in real time

/θ//ʃ//aɪ/
ILM AI for Speech
“A tutor you can talk to, and it talks back instantly.”

Why this name

“Conversational phonetics” captures both halves of the work: a conversation fast enough to feel human (sub-second speech-to-speech), and phonetics, the science of how speech sounds, applied as live coaching on pronunciation, rhythm, and clarity.

The problem

Typing is a barrier. Younger learners, language learners, and anyone revising on the move find it far more natural to speak, and explaining an idea aloud is one of the most powerful ways to learn. Yet most AI tutors are text-only, and existing voice tools feel slow, robotic, and stumble on real, imperfect student speech.

What we're building

A voice-first tutor built as a low-latency speech-to-speech pipeline. Learners hold a natural back-and-forth conversation with roughly 300 ms of perceived delay, practise reading and explaining aloud, and receive gentle, specific feedback on pronunciation, fluency, and delivery, down to the individual phoneme (/θ/, /ʃ/, /aɪ/).

Key capabilities

Sub-second conversation: speak naturally and hear an answer with barely any delay.

Accent-robust recognition: understands regional accents, hesitations, and classroom-quality audio.

Speak-to-learn: learners explain concepts aloud and the tutor coaches their reasoning.

Phoneme-level feedback: actionable tips on pronunciation, rhythm, intonation, and pace.

Hands-free and accessible: revision on the go, and a lifeline for learners who struggle with typing.

Research focus

Low-latency speech-to-speech architectures that keep the conversation human: latency ≤ 300 ms.

Recognition that stays accurate on the messy, spontaneous speech of real students: P(phoneme | audio, context).

Prosody-aware assessment, rhythm and intonation as well as words, that encourages rather than corrects.

Where it fits at ILM AI

A spoken mode for Ilmino's AI tutor: revision by talking, oral exam practice, spoken languages, and accessibility-first learning.

Sonus: just say it, and learn out loud.

≤ 300 ms round tripaccent robust ASRprosody feedback

More research