Skip to main content icon/video/no-internet

Reliability refers to the consistency or stability of a measurement. A test or instrument with good reliability means that the respondent will obtain the same score on repeated testing as long as no other extraneous factors affect the score. In actuality, a respondent will rarely obtain the exact same score over repeated testing because repeated assessments of any phenomenon will likely be affected by chance errors. Thus, the goal of testing is to minimize chance errors and maximize the reliability of the measurement with the recognition that a perfectly reliable measure is rarely attainable. Although a highly reliable test will not yield identical scores for a participant from Time 1 to Time 2, the scores will tend to be similar if the test is reliable. Thus, for example, on a 7-point Likert-type scaled item (1 = strongly false to 7 = strongly true), a participant whose score is high on an item at Time 1 will tend to score high on that same item at Time 2 if the item is reliable. This entry discusses the different components of reliability and the relationship between reliability and validity.

Reliability and Validity

Reliability is extremely important because evidence of reliability is necessarily the first step in establishing the scientific acceptance and usefulness of a test. Indeed, good reliability is a prerequisite for validity of a test, which is defined as the extent to which a test accurately measures the construct that it purports to measure. It is possible for a test to have good reliability but poor validity. For example, if a group of people all agree that an older adult who suddenly became unable to hear or speak suffers from a fictitious disorder, alphabetitis, that would provide evidence of reliability of the diagnosis because all of the people in the group agree. However, this symptom “unable to hear or speak” may not actually measure alphabetitis (because it is fictitious). Thus, the diagnosis would lack validity even though it has reliability. As another example, there is general agreement from reports about UFOs that extraterrestrial aliens have enlarged heads; small slits for noses; and eyes that are large, opaque, black, and tear-shaped. Although agreement provides evidence of reliability, there is no scientific evidence that extraterrestrial aliens exist, or supposing they did, that this description is accurate. Thus, reliability is a necessary but not sufficient precursor for validity. More specifically, in test construction and development, evidence for reliability is required before the validity of a test can subsequently be evaluated. The two most common forms of reliability are test–retest reliability and scale reliability.

Test–Retest Reliability

Test–retest reliability is a measure of a test’s consistency over a period of time. Test–retest reliability assumes that the construct being measured is relatively stable over time, such as personality characteristics or intelligence. A good test manual should specify the sample, the test–retest interval (typically about 1 week to several months), and the reliability coefficient (typically Pearson’s correlation). If the construct is likely to change over time (e.g., perceived stress), then test makers generally choose a shorter interval (e.g., 1 week). Test–retest reliabilities values are reported as correlation coefficients, and the values are considered excellent if they are .90 or higher and good if they are .80 or higher. If a construct is thought to be relatively stable, but the test–retest reliability coefficient for a test of that construct is around .50, it most likely means that the test is unreliable. Some possible explanations to consider for the low reliability are that there are too few questions on the test or that some questions are confusing, too long, or poorly worded. Another possibility is that some extraneous variable affected the construct during the interval (like further stress or a financial windfall), causing the scores to change, although the change was not due to item unreliability. One final concern regarding the interpretation of test–retest reliabilities is that they may be spuriously high because of practice or memory effects. A respondent may do better on the second testing because the trait being assessed improves with practice. Some people may respond similarly to a second administration of a test because they remember many of the answers that they provided previously. One possible solution to this problem is the use of alternate forms of the same test.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading