Skip to main content icon/video/no-internet

Segmentation of speech refers to the human ability to identify the boundaries of discrete language units (e.g., phonemes, syllables, words) in the continuous speech signal. Although the speech signal has many silent parts, these are not necessarily related to the boundaries of language units or pauses between boundaries. For instance, some of these silent portions can be due to speech articulation (e.g., stops closure). This identification process, which is automatic and fast for adults when listening to a familiar language, relies on several cues that work in a probabilistic way, merging information from different linguistic sources: lexico-semantic, syntactic, and prosodic. However, word segmentation cues are not completely reliable and they are language-specific. Native language experience plays an important role in defining relevant cues for speech segmentation in each language. Adults perform speech segmentation very easily in their own language provided that conversation or the speech signal is not affected by contextual conditions, such as a noisy environment. This entry discusses speech segmentation as a relevant issue for speech perception, language acquisition, language learning, and automatic speech recognition (ASR).

Identifying word forms in the oral stream and mapping them onto meaningful units is a complex task because words in speech are not produced by how they are in isolation, and they have different characteristics according to the phonetic and phonotactic context in which they occur. For example, in the speech stream, words are produced sequentially, and sometimes, concatenation and coarticulation occur between segments or even syllables. Both concatenation and coarticulation create an impact on the acoustic properties of words, changing the signal significantly. This means that the way a word sounds can be highly variable.

Speech segmentation is also a challenging task in language acquisition because infants do not have access to a mental lexicon, at least in their first year of life, so they cannot use it to identify language units (i.e., words) in the oral stream. The way infants deal with speech segmentation has been a research issue with contradictory results. Early studies proposed that infants perform speech segmentation first by learning words in isolation and then by identifying them within the speech continua. This claim has been attacked because the speech that infants hear is not produced word by word in isolation: Words are embedded in the speech stream. So infants must have another way to identify meaningful units. Another strong claim supports the hypothesis that infants rely on their ability to segment speech by using sublexical segmentation, that is, by using patterns provided by metrical cues, acoustic cues, and phonotactic cues. Metrical cues are related to strong and weak syllables and how they are ordered in a language. Researchers tested the idea that strong syllables in continuous speech could be considered word onsets by the listener, and results of studies on speech perception supported this view. Acoustic cues, such as longer segment duration or initial word segment aspiration, are cues that indicate the location of word beginnings in some languages. Phonotactic cues indicate possible word boundaries. Disallowed phonemic sequences in the language are key factors in aiding word segmentation.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading