Skip to main content icon/video/no-internet

In digital signal processing, sampling refers to the reduction of a continuous signal such as speech into a discrete signal (i.e., a series of numbers or digits). Similarly, sampling rate is the rate at which samples of a continuous signal are taken and converted into a digital form, typically expressed in samples per second (Hz). This entry discusses the process of sampling in the context of the digital recording of speech.

In the field of speech and language pathology, it is often important to analyze the spoken output of an individual with a speech disorder for both clinical and research purposes. However, speech is a fleeting event; the airborne acoustic waves produced by the speaker’s articulators, and received by the listener’s ears, quickly vanish. In order to be able to store, study, and analyze the speech signal, these continuous changes in air pressure are converted into another form that can be easily stored and manipulated. This can be accomplished through the digital signal processing of speech. By transforming speech into a digital form, the speech signal can be stored. Specialized acoustic software can be used to analyze it in an efficient and comprehensive manner. Sampling is one of the basic operations required for the process of digitizing speech signals.

Although the digital recording of speech requires a highly sophisticated set of processes, the basic steps involved are straightforward. Initially, the speech output to be recorded is received by a microphone. The sound waves generated by the speaker hit the microphone’s diaphragm and set it into vibration. These vibrations are carried into the other components of the microphone and converted into an electric current. This conversion of the original sounds into electrical energy is the first step toward making speech signal storable.

However, even in this form, the signal is still an analog or continuous signal; its amplitude varies continuously with time. Since digital computers operate on a sequence of numbers, the speech signal must be further converted into a series of discrete points. This is achieved by obtaining a sequence of discrete samples of time and amplitude from the continuously variable waveform. This process is called analog-to-digital conversion. It primarily involves three operations: filtering, sampling, and quantization. If these operations are performed properly, the resulting digital speech signal will accurately represent the original acoustic wave without any loss of information.

The sampling operation is arguably the most important aspect of the digitization process. The sampling operation chops the analog signal at certain time points that are periodically spaced, resulting in a number of equal intervals. The number of samples used depends on the sampling rate of the digitization process. A sampling rate of 20000 Hz, for example, means that every second of the continuous speech signal is converted into a sequence of 20,000 samples and thus the interval between each sampling point is 0.00005 s.

By definition, the process of sampling involves reducing the original signal; only a certain number of points are sampled, and the energy between the sampling points is discarded. As a result, the acoustic or electrical energy of the analog signal, with its infinite values, is reduced to a finite number of samples. For the digitized version of the speech signal to be an accurate and unambiguous reconstruction of the original signal, it is necessary to select an appropriate sampling rate. According to Harry Nyquist’s sampling theorem, if the sampling rate chosen is at least twice the frequency range (bandwidth) of the speech signal being digitized, the analog-to-digital conversion will be carried out without any loss of information. This sampling rate is called the Nyquist frequency. Nyquist’s theorem is based on the premise that two sampling points are enough to capture the highest frequency of interest. For instance, if one wishes to sample a speech signal that has a bandwidth of 12000 Hz (i.e., a range of 0–12000 Hz), then this signal should be sampled at a rate of at least 24000 Hz in order to obtain a digital signal equivalent to the original analog signal. It is not necessary to sample at a rate higher than the Nyquist frequency, and this, in fact, should be probably avoided because it requires additional computational power without any significant gains in terms of the accuracy of the speech signal. In any case, sampling rates higher than about 50000–60000 kHz cannot supply more usable information for human listeners. On the other hand, undersampling (i.e., using a sampling rate lower than the Nyquist frequency) results in a digital signal that is a distorted representation of the original speech event. Thus, any subsequent analysis would contain serious errors.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading