The Problem: Dialogue That Actually Sounds Like a Conversation
English proficiency tests - the ones immigrants need to pass for immigration, education, and employment in Canada - hinge on listening comprehension. Candidates are tested on their ability to follow natural spoken dialogue: two people talking in a café, a customer resolving an issue with a bank representative, a tenant speaking with a landlord.
Producing that kind of audio at professional quality is the bottleneck that has kept high-quality ESL listening materials out of reach for independent course builders. You either pay a voice studio thousands of dollars, or you compromise on quality.
Kris and his team at Anglophonic tried the obvious alternative first: ElevenLabs. They generated voice lines individually and stitched the dialogue together. The result fell apart at exactly the moment that mattered most - it didn't sound like two people in conversation. It sounded like two people reading separate scripts in the same room. For a listening exercise where the whole point is natural comprehension, that's a fundamental failure.
The core challenge: ESL listening exercises require dialogue that sounds genuinely conversational - natural turn-taking, appropriate pacing, realistic intonation - not a sequence of individual narrations. Stitching audio together line-by-line produces the latter, not the former.
The question was whether AI audio had reached the point where it could generate a real conversation, not simulate one.
The Workflow: From Test Materials to Audio in One Day
Anglophonic built a clean four-step pipeline around ZenMic Studio. It's repeatable, fast, and requires no audio production skills whatsoever:
Feed the source test prep documents into NotebookLM or Gemini to analyse the structure, question types, and scenario requirements - café dialogues, phone calls, workplace exchanges.
Use the AI to generate realistic dialogue scripts shaped around the test scenarios - grounded in the vocabulary levels, conversational contexts, and comprehension targets of the actual exam.
Translate each dialogue script into a ZenMic Studio prompt - structured to trigger natural two-speaker audio generation with the right tone, pacing, and setting for each scenario.
Run three dialogues at a time directly in ZenMic Studio - building a full batch of ready-to-use listening exercises in parallel, not sequentially.
One day of prompting produced a complete, usable batch of test-prep listening dialogues - the kind of output that would have taken weeks and thousands of dollars to produce through traditional voice production.
The Result: Clearer Than the Professional Originals
The output from ZenMic is even better than the original test materials in terms of clarity and sound quality… I can see this platform as a real educational platform.
The bar Kris is comparing against isn't amateur content - it's the listening audio produced by professional English-language test publishers, the organisations whose materials have been used in IRCC-certified settlement programmes for years. Beating that bar on clarity and naturalness is significant.
What ZenMic delivered wasn't just technically proficient audio - it was dialogues that sounded like real conversations. The critical difference between a line-by-line TTS approach and a full-dialogue generation is exactly this: naturalness. Candidates preparing for English proficiency tests need to hear the rhythm, intonation, and flow of actual spoken English - not a simulation of it.
Anglophonic now has a scalable, repeatable production pipeline for building out its 2026 course library - one that doesn't require a recording studio, voice talent, or an audio editor.
Why ZenMic Where ElevenLabs Fell Short
- ✕ Generate each speaker line separately
- ✕ Manual stitching in audio editor
- ✕ No conversational flow between speakers
- ✕ Unnatural pacing and turn-taking
- ✕ Hours of manual production per dialogue
- ✕ Fails the listening exercise quality bar
- ✓ Generate complete two-speaker dialogues
- ✓ No audio editing required
- ✓ Natural conversational rhythm and flow
- ✓ Realistic pacing, intonation, turn-taking
- ✓ Batch of three dialogues per prompt run
- ✓ Clearer than professional test-publisher audio
What's Next for Anglophonic
Anglophonic is continuing to build out its 2026 course library using the ZenMic pipeline and plans to share this workflow publicly on LinkedIn - so other EdTech builders and settlement organisations can replicate it.
Their roadmap includes working directly with IRCC-certified settlement organisations - the federally funded bodies that help newcomers to Canada navigate immigration, employment, and language requirements - so those organisations can refer clients straight into Anglophonic's courses.
One open wish: Background ambience for scene-based dialogues - café noise for café conversations, office sounds for workplace exchanges. Noted as a nice-to-have rather than a blocker. The core quality of the dialogue audio is already there.