From Tone Marks on the Page to Tones in Real Speech

Pinyin gives us stable citation forms, not a complete pitch script for connected speech. To understand why accurate tone-mark reading can still sound unnatural, we need to trace the roles of tone change, first-language background, perception, and production.

10 min readChapter 3mandarin tonespinyintone sandhi

A learner looks at nǐ hǎo, identifies both third-tone marks correctly, and pronounces two careful third tones. Every symbol has been read as printed, yet the result sounds stiff or unnatural. We often explain this gap by saying that “tones are difficult,” but that label does not identify what went wrong or what to practise next. A more useful explanation follows a causal chain from written citation forms to connected speech, then through the learner’s first-language system, and finally into the separate tasks of hearing and producing pitch patterns.

Pinyin is an input to speech, not a recording of it

The strongest starting point is a property of pinyin itself: it normally marks tones in their citation forms, not every form those tones take inside continuous speech. Citation form is the conventional form we use when identifying a syllable or word in isolation. Connected speech is different because neighbouring tones can trigger systematic changes.

The familiar pair nǐ hǎo illustrates the distinction directly. On the page, we write nǐ hǎo. In ordinary pronunciation, the first syllable changes under the relevant tone-change rule, producing a pattern conventionally represented as ní hǎo. The learner therefore has to perform an operation that is not fully displayed in the spelling. AllSet Learning’s guide to tone-change rules describes this practical problem clearly, while a 2026 study in Frontiers in Psychology supports the broader distinction between pinyin orthography and the forms that emerge in the speech stream.

This matters because pinyin is often treated as if it were an exact pitch score. It is not. It is a stable writing system that tells us which lexical tones underlie the syllables, but it does not rewrite every word to reflect every contextual adjustment made in speech. That stability is useful: readers should not have to learn a new spelling whenever the phonetic environment changes. The cost is that learners must know how to transform the written sequence into an appropriate spoken sequence.

Once we recognize that transformation, a common pronunciation problem becomes easier to diagnose. A learner may identify every printed mark correctly and still produce a sequence that is too literal. The issue is not necessarily ignorance of the individual tones. It may be failure to apply tone change across syllable boundaries. Practice aimed only at isolated syllables will not automatically solve a problem that appears when syllables are combined.

“Difficult” describes a relationship, not a sound

The next link changes how we talk about pronunciation difficulty. A sound or contrast is not simply easy or difficult in itself. Its difficulty depends partly on its relationship to the categories and habits already established in the learner’s first language, or L1.

A contrast that resembles something familiar in one L1 may be unfamiliar to speakers of another. Research on second-language speech acquisition, including work by Antoniou and Chin in Frontiers in Psychology and a 2019 study in the Australian Journal of Linguistics, treats learners’ prior language experience as relevant to tone perception and learning. This does not mean L1 determines the final result. It means that statements such as “Mandarin tone X is difficult” are incomplete unless we also ask: difficult for whom, and in what task?

That shift has immediate practical value. “Mandarin tones are hard” gives us no observable target. A better diagnosis asks whether the learner has difficulty distinguishing pitch patterns, producing a controlled contour, applying tone change, or coordinating all three processes while speaking. These are related problems, but they are not identical. A learner may hear a contrast reliably yet fail to produce it. Another may produce isolated tones accurately but lose them in a multisyllabic sequence. A third may know the tone-change rule explicitly but fail to apply it at conversational speed.

L1 background can help us predict possible directions of transfer, but it cannot guarantee an individual outcome. Exposure, training history, attention, and time spent using the language remain important variables. We should use L1 as a diagnostic lens, not as a verdict on what a learner can or cannot achieve.

Vietnamese offers support, interference, and no guarantees

For learners whose L1 is Vietnamese, the relational view of difficulty becomes especially important. Vietnamese speakers already know from daily language use that pitch-related contrasts can distinguish words. That experience may provide an initial perceptual framework that learners from non-tonal L1 backgrounds do not possess.

Yet familiarity with lexical tone is not the same as having Mandarin tone categories ready for transfer. Vietnamese is commonly analysed with six tones, while Standard Mandarin is commonly taught with four lexical tones. The mapping between the two systems is not mechanical. Learners cannot simply assign each Mandarin tone to a Vietnamese equivalent and expect the resulting pitch patterns to be natural. Potential gaps may also occur in consonant contrasts, which means that a learner’s pronunciation profile cannot be reduced to tone alone.

Several studies provide relevant pieces of this picture. A 2023 article in Global Chinese examined 30 Vietnamese-L1 learners using a set of 80 disyllabic words. A 2022 study in the Journal of Phonetics followed 33 advanced Vietnamese-L1 learners. A 2025 article in the educational science journal of Ho Chi Minh City University of Education discussed the use of Sino-Vietnamese knowledge in vocabulary learning. That last source concerns lexical learning rather than serving as direct evidence about Mandarin tone production, but it illustrates the broader point that an existing language background can function as a resource without transferring automatically or uniformly.

We should therefore handle the proposed “Vietnamese tone advantage” carefully. It is a plausible hypothesis with indirect support, not an independently verified guarantee. The learner samples in the cited studies are limited, and the evidence for interference or mismatch is currently denser than the evidence for a general advantage. The most defensible conclusion is narrower: prior experience with a tonal language may create favourable starting conditions, but it does not allow us to predict the speed or quality of progress for every Vietnamese learner.

This distinction also matters for teaching. If we assume that Vietnamese speakers “already understand tones,” we may neglect the specific Mandarin categories and contextual rules they still need to learn. If we assume that their existing tone system is only a source of error, we overlook a potentially useful perceptual resource. A sound approach treats L1 knowledge as both support and possible interference, then tests what an individual learner can actually perceive and produce.

Hearing a tone and producing it are different jobs

The final link takes us from description to training. Tone perception and tone production interact, but they are not the same skill. Perception practice asks learners to distinguish, identify, or categorize what they hear. Production practice asks them to generate an intended pitch pattern with adequate control. Connected-speech practice adds another problem: adjusting a tone when its context requires a changed form.

The evidence is strongest for the value of structured perception training. Antoniou and Chin’s 2018 work addresses tone perception in inexperienced listeners. The 1999 study “Training American Listeners to Perceive Mandarin Tones,” published in the Journal of the Acoustical Society of America, reported improvement from perceptual training, and a 2003 follow-up in the same journal examined longer-term training effects. Together, these studies support the proposition that listeners can improve their perception of Mandarin tones and that perceptual learning can extend beyond the exact items used during training, including into production-related performance.

A stronger comparative claim sometimes follows: production training yields much smaller gains than perception training. The sources cited here do not independently establish that quantitative comparison, and the available evidence does not warrant turning it into a general teaching law. It remains an open question rather than a settled fact. We should not conclude that speaking practice is ineffective, nor should we present a precise ranking of training methods that the evidence does not support.

The safer instructional consequence is to separate the jobs without isolating them. We can train listening to improve discrimination, train speaking to develop control over pitch movement, and train tone change to connect syllables in realistic sequences. Each task addresses a different link in the chain:

  • Perception practice tests whether we can hear the intended contrast.
  • Production practice tests whether we can create the intended contour.
  • Tone-change practice tests whether we can transform citation forms inside a sequence.
  • Connected-speech practice tests whether these operations remain available when attention is shared across words, meaning, and timing.

This division gives teachers and learners something more actionable than “practise tones.” It also prevents two opposite mistakes. One is to treat an unsettled comparison as proof that speaking practice contributes little. The other is to dismiss well-supported perception training because claims about its superiority have been overstated. The evidence supports structured listening work; production remains necessary; and contextual tone change must be practised because it is required by the gap between orthography and speech.

We can now answer the original question without falling back on the vague claim that tones are simply hard. Correctly reading tone marks does not ensure a natural pitch sequence because the written marks represent citation forms. Speech may require tone changes that pinyin does not display. Those changes must then pass through the learner’s existing perceptual categories and production habits, which are partly shaped by L1.

The four links also imply four different diagnostic questions. Are we reading the underlying tones correctly? Do we know which contextual changes apply? Can we perceive the resulting contrast? Can we produce it while combining syllables into speech? If we skip any of these questions, we risk prescribing the wrong task. More isolated repetition will not necessarily repair an unrecognized tone-change error, and more rule explanation will not necessarily repair an unstable pitch contour.

For Vietnamese learners, the same chain cautions against both pessimism and overconfidence. A tonal L1 may provide useful experience, but Mandarin does not inherit its categories from Vietnamese. Transfer may help in one part of the task and interfere in another. The only responsible way to use that background is to treat it as a hypothesis about where to look, then evaluate the learner’s actual performance.

Conclusion and limits

Reading pinyin accurately and speaking with a natural pitch sequence are not equivalent achievements. Pinyin supplies stable citation forms; connected speech requires contextual transformation; and the learner must perceive and produce the result through a system influenced by prior language experience. That causal chain turns a broad complaint about “difficult tones” into separate, observable learning tasks.

This article does not settle three larger questions. First, it does not independently verify a general Vietnamese advantage in learning Mandarin tones. Second, it does not establish that production training produces substantially smaller gains than perception training; that comparison remains contested and is not independently verifiable from the sources cited here. Third, it offers no timetable for improvement, no ranking of first-language backgrounds, and no formula for converting research findings into a fixed training dose. The evidence supports a clearer diagnosis, not a promise of outcomes.

Sources cited

  • AllSet Learning, “Tone Change Rules”
  • Frontiers in Psychology, volume 17, article 1856709, 2026
  • Antoniou and Chin, Frontiers in Psychology, 2018
  • Australian Journal of Linguistics, volume 39, issue 3, 2019
  • Global Chinese, volume 9, issue 2, 2023
  • Journal of Phonetics, volume 95, article 101197, 2022
  • Ho Chi Minh City University of Education Journal of Science, 2025
  • “Training American Listeners to Perceive Mandarin Tones,” Journal of the Acoustical Society of America, volume 106, issue 6, 1999
  • Journal of the Acoustical Society of America, volume 113, issue 2, 2003
An Anatomy of Chinese14 chapters · from naming to digital life
Browse the series