Skip to main content icon/video/no-internet

Reliability

This entry focuses on reliability, which is a desired property of the scores obtained from measurement instruments. Therefore, reliability is relevant in a variety of contexts, such as scores on multiple-choice questions on an achievement test and Likert-type scale responses on a survey questionnaire. The concept of reliability is premised on the idea that observed responses are imperfect representations of an unobserved, hypothesized latent variable. In the social sciences, it is typically this unobserved characteristic, rather than the observed responses themselves, that is of interest to researchers. Specifically, reliability serves to quantify the precision of measurement instruments over numerous consistent administration conditions or replications and, thus, the trustworthiness of the scores produced with the instrument. This entry discusses several frameworks for estimating reliability.

Reliability as Replication

While it is assumed that items (also called questions, tasks, prompts, or stimuli) on an instrument measure the same theoretical construct, this assumption always needs to be tested with empirical data. Reliability does not provide direct information about the meaningfulness of score interpretations, which is the concern of validation studies, but serves as an empirical prerequisite to validity.

Reliability provides information on the replicability of observed scores from instruments. Consequently, at least two sets of scores are required to obtain information about reliability. These can be multiple administrations of an instrument to the same group of respondents (test–retest reliability or score stability), the administration of two comparable versions of the instrument to the same group of respondents (parallel-forms reliability), or the comparison of scores from (at least) two random halves of an instrument (split-half reliability and internal consistency).

To estimate test–retest reliability, the original instrument has to be administered twice. The length of the time interval is not prespecified but should be chosen to allow variance in performance without the undue influence of developmental change; ideally, the time interval should be reported along with the reliability estimate. To estimate parallel-forms reliability, two forms of the instrument have to be developed according to the same set of test specifications and have to be administered to the same group of respondents. As with test–retest reliability, the time interval between administrations is important and should be reported.

To estimate the split-half reliability or internal consistency of an instrument, only one instrument needs to be administered to one group of respondents at one point in time. Technically, the instrument needs to be split into at least two randomly parallel halves. Of course, there are many ways to divide the items on the instrument into two halves. To overcome the limitation imposed by the arbitrariness of the split, coefficient alpha has been developed. It is theoretically equivalent to the average of all potential split-half reliability estimates and is easily computed. Estimating the internal consistency of an instrument effectively measures the homogeneity of the items on the instrument.

Reliability as a Theoretical Quantity

To understand the definitions of different reliability coefficients, it is necessary to understand the basic structure of the measurement framework of classical test theory (CTT). In CTT, observed scores (X) are decomposed into two unobserved components, true score (T) and error (E; i.e., X = T + E). While true score is assumed to be stable across replications and observations, error represents a random component that is uncorrelated with true score. Additional restrictions can be placed on the means and variances of the individual components as well as their correlations in order to make them identified and estimable in practice.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading