Matrix Hanzi · Pinyin Station Interactive pronunciation for beginners

Pronunciation foundation · No login needed

Hear it. Dissect it. Repeat it.
Every Pinyin syllable, nailed.

The station breaks Mandarin phonetics into three digestible layers: initials (onsets), finals (rimes) and tones — then reassembles them in a complete syllable matrix. Click any cell to open its "dissection space": dual-speed audio, mouth-shape guidance, memory hooks and a one-click minimal speaking popup that scores your pronunciation right on the page.

  • 21 initials
  • 41 finals
  • 5 tones
  • 411 syllables
  • 1,307 toned variants
  • audio in 2 speeds
Free — all content & audioNo login — progress stays on your deviceNo personal data collected
Line-art profile of a face emitting sound waves — the pronunciation station

Curriculum flow

Four discovery stages, from familiar sounds to hard ones

No dumping the whole Pinyin table at once. Each stage is a small "micro-milestone": hear a few sounds, dissect the mouth shape, shadow them — then move on. Click a stage to highlight its syllables on the matrix.

Full combination table · Full syllable matrix

The syllable matrix: initials × finals

Every cell is a real syllable of standard Mandarin. Click a cell to open its dissection space: pick a tone, hear both speeds, see the representative character, the mouth shape, and practice it in the minimal speaking popup. Empty cell = the combination does not exist. Soft-gold cells are common syllables (frequent in everyday conversation) — learn these first; switch on "Prioritize common syllables" to dim the rarer ones.

Stage 1 — easy, familiar sounds Stage 2 — aspiration & curled tip Stage 3 — ü, er & spelling rules Common — learn these first

Outside the matrix

The "loner" syllables

A few syllables are not built from any initial + final in the table — they are interjections or special cases. Rare, but worth knowing so they do not surprise you in a story or a drama.

Layer 1 · Initials (声母)

21 onsets — six mouth-shape groups

Initials decide the "launch" of a syllable. Learners typically confuse three axes: aspirated / unaspirated (b–p, d–t, g–k…), tongue place (z–zh–j, c–ch–q, s–sh–x) and voicing (r). Every card carries dual-speed audio of the call-name sound (呼读音), the mouth shape and a memory hook. Soft-gold cards mark initials that appear in common syllables — the same color convention as the matrix.

Illustration of a hand held in front of the mouth to feel the puff of air

The hand test — your aspiration detector

Hold your palm 10 cm from your mouth. Say b, d, g, j, zh, z: almost nothing reaches your palm. Say p, t, k, q, ch, c: a puff of wind hits it. Those pairs differ by exactly that puff — and untrained ears simply do not hear it. Let your hand hear it for a few weeks and your ears will catch up.

Layer 2 · Finals (韵母)

41 finals — from simple vowels to nasal rimes

Finals are the "body" of a syllable, where the tone lives. Three classic traps: ü (say "i" with rounded lips), ing/eng (they end in ng, not n) and ian (pronounced "ien"). Hear each final at both speeds and check the mouth shape before combining it on the matrix. Soft-gold cards mark finals that appear in common syllables — the same color convention as the matrix.

Layer 3 · Tones (声调)

Four tones + the neutral on a five-level pitch scale

Mandarin is tonal: the same "ma" with four different pitch contours is four different words. The chart on the left draws each tone's pitch path on a 1 (low) → 5 (high) scale. Use the slow audio to feel the glide.

54321 1 2 3 4 neutral
Five-level pitch scale: tone 1 flat at 5; tone 2 climbs 3→5; tone 3 dips 2→1→4; tone 4 falls 5→1; the neutral tone is short with no contour of its own.

Tone sandhi (变调)

Three sandhi rules worth memorizing

In flowing speech some tones change so the mouth does not stumble. Hear each pair first — let your ear catch the pattern before reading the rule.

Reflex loop · Shadowing loop

Hear the model → shadow it → get scored

Every dissection space has a button that opens a minimal speaking popup right on the page: pick Free (browser recognition, unlimited) or AI-beta (detailed pronunciation scoring), read the current word, then get your score, what the machine heard and replay buttons — for both the model audio and your own voice. For deeper sentence-level analysis, the popup links to the full Speaking Studio.

1

Hear the standard

Press "Standard" to load the natural-speed native model audio.

2

Inspect the slow version

The slow take exposes the puff of air, the vocal-fold buzz and the pitch glide.

3

Shadow it

Speak along with the audio as soon as you hear it — do not wait for the end. Your mouth learns faster than your memory.

4

Practice in the popup

The minimal popup opens right on the page with the current word: pick Free or AI-beta, tap the mic, read aloud, get your score and replay your own voice.

Illustration of a microphone and the shadowing loop

Shadowing tips for beginners

Do not translate in your head. Imitate the music of the syllable: its height, its length, its puff — like mimicking a line of a song. For tone 3, exaggerate the dip (sink really low, then flick up); for tone 4, be decisive, like giving an order. Ten sounds a day and in two weeks your ear separates z/zh/j on its own.

Open Speaking Studio →