Skip to main content icon/video/no-internet
Measurement, Theories of

Measurement is the process of estimating psychometric properties of variables or constructs. There are a number of measurement theories that represent different assumptions about and approaches to the estimation process, including the classical test theory (CTT), the generalizability theory (G-theory), the congeneric theory, and the item response theory (IRT).

Classical Test Theory

The most influential theory in psychometric measurement is the CTT. The theory relies on the test score to measure a person's ability or other psychometric properties. Central to the theory is the concept of reliability, first used by Charles Spearman in 1904. According to Spearman, the observed score of a variable or construct consists of two separate components: the true score and the error of measurement. The true score represents the perfect measurement of a construct, for instance, the true ability of a person to accomplish a task. In theory, if a person takes the same test an infinite number of times, the average of the scores will equal the true score. The measurement error, which accounts for the difference between an observed score and the true score, can be divided into systematic error and random error. The former is related to either the examinee or the measure, although not to the construct being measured; the latter results from chance happenings such as guessing and distractions in test administration. The presence of measurement errors makes a measure unreliable. Mathematically, the coefficient of reliability is the ratio of variance of the true score to that of the observed score. The estimation of reliability coefficient is based on the scores of a group of examinees. There are three methods for estimating reliability: (1) the test-retest method is based on repeated administrations of a measure to the same group of examinees and correlates the results from the different administrations, (2) the alternate form approach uses equivalent forms of the same test to be administered to a group of examinees and examines the correlation between the results, and (3) the most frequently used method for reliability estimation is the internal consistency approach that looks at the homogeneity of items within the same instrument. Related to this approach is the split half method of estimating reliability that divides an instrument into two halves and correlates the results from the halves. Estimation of internal consistency for dichotomous items is based on the Kuder-Richardson formulas, while the estimation for polytomous items is based on the Cronbach's coefficient alpha.

Another important index of reliability, the standard error of measurement, is based on individual examinee's scores. It is the standard deviation of the differences between observed scores and the true score for an individual. Standard error of measurement can be used to determine the score band for an individual examinee, the range of scores where the person's true score is most likely located. The standard error of measurement is equal to the standard deviation times the square root of 1 minus the reliability of the measurement.

The second important CTT concept is validity, which is defined as the extent to which an instrument measures what it is supposed to measure. This concept, however, is not considered a feature unique to CTT. Traditionally, there are three types of validity. Content validity concerns whether the sample of test items adequately represents the content of a subject tested. Criterion-related validity uses an external measure as a criterion to validate the results of an instrument. A subcategory of criterion-related validity is concurrent validity, which can be established if high correlation is found between results of the test and results of a concurrent measure of the same variable. When the criterion is a measure available in the future rather than concurrent, the correlation between results of the test and results of the criterion can be used as evidence for predictive validity of the test, the other subcategory of criterion-related validity. Construct validity concerns whether a construct functions in a way consistent with a relevant theory. It can be established through convergent analysis or discriminant analysis. Evidence for convergent validity is established if high correlation is found between the construct in question and a theoretically related variable. Evidence for discriminant validity is established when a low correlation is found between the construct and a theoretically unrelated variable. Construct validity can also be established if the internal structure of a construct is confirmed through a statistical procedure called factor analysis.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading