miolingo

Overview

miolingo is a multi-language pronunciation trainer with real-time AI feedback across seven languages. It started as a way to get phonemic transcriptions and reproduce them via IPA-to-speech, and grew from there into a full scoring and feedback pipeline.

Why and how it was built

The scoring pipeline pairs espeak-ng’s phonemic reference IPA against a Whisper transcription of the learner’s actual speech, then compares them phone-by-phone using panphon articulatory feature distances — rather than naive text-edit distance — so a near-miss vowel scores differently from a genuinely wrong consonant. Per-language “fold maps,” extracted from espeak’s own allophone rules, tolerate accent variation without missing real errors. The project is mid-rewrite from a Streamlit/MySQL app into a local-first Svelte SPA over a stateless FastAPI phonetics sidecar, built directly from a formal Wolfram Language state-machine spec that keeps three parallel ports (web, Swift, Electron) honest against the same test tables.

Key Features

  • Phone-by-phone scoring via articulatory feature distances, not text-edit distance
  • Real ASR in the loop: Whisper transcribes the learner’s actual speech for comparison
  • Per-language accent tolerance via fold maps derived from espeak-ng’s own allophone rules
  • Seven languages, with a formal state-machine spec keeping web/Swift/Electron ports in sync

Status

Mid-rewrite: the actively developed tree is a feature branch, not the repository’s stale default branch. Nearly every real feature (/api/attempt scoring, /api/tts, /api/g2p, /api/materials, /api/minimal-pairs, /api/translate) is served by a FastAPI sidecar that imports the Python phonetics stack (espeak-ng, Whisper, panphon) directly — there’s no client-side phonetics, so a live “try it” deploy would need that backend hosted, not just static files. No deploy here for that reason.