Skip to main content icon/video/no-internet

Correlation is a synonym for association. Within the framework of statistics, the term correlation refers to a group of indices that are employed to describe the magnitude and nature of a relationship between two or more variables. As a measure of correlation, which is commonly referred to as a correlation coefficient, is descriptive in nature, it cannot be employed to draw conclusions with regard to a cause–effect relationship between the variables in question. To infer cause and effect, it is necessary to conduct a controlled experiment involving an experimenter-manipulated independent variable in which subjects are randomly assigned to experimental conditions. Typically, data for which a correlation coefficient is computed are also evaluated with regression analysis. The latter is a methodology for deriving an equation that can be employed to estimate or predict a subject’s score on one variable from the subject’s score on another variable. This entry discusses the history of correlation and measures for assessing correlation.

History

Although in actuality a number of other individuals had previously described the concept of correlation, Francis Galton, an English anthropologist, is generally credited with introducing the concept of correlation in a lecture on heredity he delivered in Great Britain in 1877. In 1896 Karl Pearson, an English statistician, further systematized Galton’s ideas and introduced the now commonly employed product-moment correlation coefficient, which is represented by the symbol r, for interval-level or continuous data. In 1904 Charles Spearman, an English psychologist, published a method for computing a correlation for ranked data. During the 1900s numerous other individuals contributed to the theory and methodologies involved in correlation. Among the more notable contributors were William Gossett and Ronald Fisher, both of whom described the distribution of the r statistic; Udney Yule, who developed a correlational measure for categorical data, as well as working with Pearson to develop multiple correlation; Maurice Kendall, who developed alternative measures of correlation for ranked data; Harold Hotelling, who developed canonical correlation; and Sewall Wright, who developed path analysis.

The Pearson Product-Moment Correlation

The Pearson product-moment correlation is the most commonly encountered bivariate measure of correlation. A bivariate correlation assesses the degree of relationship between two variables. The product-moment correlation describes the degree to which a linear relationship (the linearity is assumed) exists between one variable designated as the predictor variable (represented symbolically by the letter X) and a second variable designated as the criterion variable (represented symbolically by the letter Y). The product-moment correlation is a measure of the degree to which the variables covary (i.e., vary in relation to one another). From a theoretical perspective, the product-moment correlation is the average of the products of the paired standard deviation scores of subjects on the two variables. The equation for computing the unbiased estimate of the population correlation is r = (zx zy)/(n1).

The value r computed for a sample correlation coefficient is employed as an estimate of ρ (the lowercase Greek letter rho), which represents the correlation between the two variables in the underlying population. The value of r will always fall within the range of –1 to +1 (i.e., –1 r+1). The absolute value of r (i.e., |r|) indicates the strength of the linear relationship between the two variables, with the strength of the relationship increasing as the absolute value of r approaches 1. When r=±1, within the sample for which the correlation was computed, a subject’s score on the criterion variable can be predicted perfectly from his or her score on the predictor variable. As the absolute value of r deviates from 1 and moves toward 0, the strength of the relationship between the variables decreases, such that when r = 0, prediction of a subject’s score on the criterion variable from his or her score on the predictor variable will not be any more accurate than a prediction that is based purely on chance.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading