🎙️ ZenMic
Open Studio

Start generating podcasts for free.

🎓 EdTech 🎧 ESL / Language Learning 🇨🇦 Canada 📖 Case Study

How Anglophonic Built English Test-Prep Audio That Outperforms the Original Test Materials

A Canadian EdTech startup building English proficiency courses for immigrants and refugees used ZenMic Studio to produce listening dialogues clearer and more natural than materials published by professional test-prep organisations - in a single day.

Website
anglophonic.ca
Industry
EdTech / ESL / Immigration
Based
Canada
Published
July 2026
1 day
To produce a full working batch of test-prep dialogues
3
Dialogues generated per prompt, in parallel
Better
Clarity than original professional test materials
$0
Recording studio or voice talent required
1

The Problem: Dialogue That Actually Sounds Like a Conversation

English proficiency tests - the ones immigrants need to pass for immigration, education, and employment in Canada - hinge on listening comprehension. Candidates are tested on their ability to follow natural spoken dialogue: two people talking in a café, a customer resolving an issue with a bank representative, a tenant speaking with a landlord.

Producing that kind of audio at professional quality is the bottleneck that has kept high-quality ESL listening materials out of reach for independent course builders. You either pay a voice studio thousands of dollars, or you compromise on quality.

Kris and his team at Anglophonic tried the obvious alternative first: ElevenLabs. They generated voice lines individually and stitched the dialogue together. The result fell apart at exactly the moment that mattered most - it didn't sound like two people in conversation. It sounded like two people reading separate scripts in the same room. For a listening exercise where the whole point is natural comprehension, that's a fundamental failure.

The core challenge: ESL listening exercises require dialogue that sounds genuinely conversational - natural turn-taking, appropriate pacing, realistic intonation - not a sequence of individual narrations. Stitching audio together line-by-line produces the latter, not the former.

The question was whether AI audio had reached the point where it could generate a real conversation, not simulate one.

2

The Workflow: From Test Materials to Audio in One Day

Anglophonic built a clean four-step pipeline around ZenMic Studio. It's repeatable, fast, and requires no audio production skills whatsoever:

1
Analyse existing test materials

Feed the source test prep documents into NotebookLM or Gemini to analyse the structure, question types, and scenario requirements - café dialogues, phone calls, workplace exchanges.

2
Generate dialogue scripts

Use the AI to generate realistic dialogue scripts shaped around the test scenarios - grounded in the vocabulary levels, conversational contexts, and comprehension targets of the actual exam.

3
Convert scripts to ZenMic prompts

Translate each dialogue script into a ZenMic Studio prompt - structured to trigger natural two-speaker audio generation with the right tone, pacing, and setting for each scenario.

4
Generate in batches of three

Run three dialogues at a time directly in ZenMic Studio - building a full batch of ready-to-use listening exercises in parallel, not sequentially.

One day of prompting produced a complete, usable batch of test-prep listening dialogues - the kind of output that would have taken weeks and thousands of dollars to produce through traditional voice production.

3

The Result: Clearer Than the Professional Originals

"

The output from ZenMic is even better than the original test materials in terms of clarity and sound quality… I can see this platform as a real educational platform.

The bar Kris is comparing against isn't amateur content - it's the listening audio produced by professional English-language test publishers, the organisations whose materials have been used in IRCC-certified settlement programmes for years. Beating that bar on clarity and naturalness is significant.

What ZenMic delivered wasn't just technically proficient audio - it was dialogues that sounded like real conversations. The critical difference between a line-by-line TTS approach and a full-dialogue generation is exactly this: naturalness. Candidates preparing for English proficiency tests need to hear the rhythm, intonation, and flow of actual spoken English - not a simulation of it.

Anglophonic now has a scalable, repeatable production pipeline for building out its 2026 course library - one that doesn't require a recording studio, voice talent, or an audio editor.

4

Why ZenMic Where ElevenLabs Fell Short

Line-by-line TTS (ElevenLabs)
  • Generate each speaker line separately
  • Manual stitching in audio editor
  • No conversational flow between speakers
  • Unnatural pacing and turn-taking
  • Hours of manual production per dialogue
  • Fails the listening exercise quality bar
Full-dialogue generation (ZenMic)
  • Generate complete two-speaker dialogues
  • No audio editing required
  • Natural conversational rhythm and flow
  • Realistic pacing, intonation, turn-taking
  • Batch of three dialogues per prompt run
  • Clearer than professional test-publisher audio
5

What's Next for Anglophonic

Anglophonic is continuing to build out its 2026 course library using the ZenMic pipeline and plans to share this workflow publicly on LinkedIn - so other EdTech builders and settlement organisations can replicate it.

Their roadmap includes working directly with IRCC-certified settlement organisations - the federally funded bodies that help newcomers to Canada navigate immigration, employment, and language requirements - so those organisations can refer clients straight into Anglophonic's courses.

One open wish: Background ambience for scene-based dialogues - café noise for café conversations, office sounds for workplace exchanges. Noted as a nice-to-have rather than a blocker. The core quality of the dialogue audio is already there.

Frequently Asked Questions

Can ZenMic generate realistic multi-speaker ESL dialogue for listening exercises?
Yes. ZenMic Studio generates natural-sounding two-speaker dialogues - café conversations, customer-employee exchanges, and contextual everyday scenes - designed for ESL listening comprehension practice. The dialogues sound like real conversations, not stitched-together voice lines, which is the critical quality marker for any listening exercise.
How does ZenMic compare to ElevenLabs for ESL audio production?
ElevenLabs requires generating each voice line separately and stitching dialogue together manually - producing audio that doesn't sound like a natural conversation. ZenMic generates full multi-speaker dialogues in a single pass. The result sounds like two people actually talking, which is essential for listening comprehension exercises.
Can I use AI-generated audio for English proficiency test prep courses?
ZenMic produces professional-quality audio that meets or exceeds the clarity and naturalness of commercially published test prep materials. Anglophonic found the ZenMic output clearer than the original test materials they were building from. The audio is suitable for IRCC settlement organisations, ESL courses, IELTS/CELPIP preparation, and English proficiency test prep.
How fast can I produce a batch of ESL listening dialogues?
Anglophonic produced a full working batch of test-prep listening dialogues in a single day using ZenMic Studio. By pairing ZenMic with NotebookLM or Gemini for script generation, you can go from raw test materials to finished audio in hours - not weeks.

Ready to Build Your Own ESL Audio Pipeline?

Join educators, course creators, and EdTech founders already using ZenMic to produce professional listening audio - without studios, voice talent, or audio editors.