Skip to main content icon/video/no-internet

Do people with higher education earn more money? If I spend more hours tutoring my 8-year-old daughter, will she do better in math? If I run more, will it lower my blood pressure? We ask questions such as these every day, and researchers invest their time in finding the answers—a way in which to measure the degree to which variables are related. The correlation statistic, or coefficient (r), is often used to assess whether there is a predictable association between two sets of scores. This entry briefly describes the properties of a correlation, the types of variables measured and corresponding correlational index, and the notion of correlation and causality.

Properties of the Correlation Statistic

Correlation statistics indicate whether there is an association between two variables. The nature of this association can be explained in terms of both strength and direction. Correlation statistics range from –1.0 to 1.0, indicating the strength of the relationship. An absolute value of 1.0 indicates a perfect correspondence and a value of 0 indicates that there is no relationship between the two variables.

The value of the correlation coefficient also indicates the direction of the relationship. A correlation can be positive, with an increase in one variable associated with an increase in the other (e.g., the more education one receives, the higher their earnings), or it can be negative, with an increase in one variable associated with a decrease in the other (e.g., the more miles one runs a day, the more his or her blood pressure decreases).

Types of Variables and Statistics

The examples just mentioned describe a type of variable called a continuous or interval variable. Interval variables represent a normal distribution of continuous values so that a score could have any fractional value (e.g., 0 to 100), there is an inherent order to the variables (lower to higher), and there is an exact difference between each score (e.g., the distance between 1 and 2 miles is the same as the distance between 3 and 4 miles). Moreover, the distribution of the scores in the population follows a bell-shaped curve, with the average score at the peak of the curve, the majority of scores close to the average, and a few extreme scores at either end of the curve.

A ratio variable is also an interval variable and contains all the properties of an interval variable but without a true value of zero. For example, temperature is a ratio variable; degrees range from low to high, the distance between degrees is the same, but 0 degrees does not indicate an absence of (or no) temperature. The correlation statistic calculated for this type of distribution of scores is the most widely used and known as Pearson product-moment correlation coefficient (r).

Variables can also be characterized as nominal or ordinal. Nominal variables are also known as categorical variables because they represent categories or labels that are mutually exclusive, meaning the value cannot fall into more than one category; the score would be either yes or no, not both. A dichotomous variable is a nominal variable composed of only two categories (e.g., male/female, yes/no). A nominal variable that is composed of more than two categories (e.g., hair color, ethnicity, political affiliation) is termed multidichotomous.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading