Skip to main content icon/video/no-internet

When data are clustered (i.e., unit observations nested within clusters), researchers often wish to know how units within clusters are correlated. For example, it may be useful for rater reliability studies to know how the ratings of judges (units) are correlated within targets (clusters). Another example from survey research is how individuals’ responses (units) are correlated within neighborhoods (clusters). These questions can be answered with intraclass correlation (ICC or ρ) coefficients, which are ratios of the variance associated with the cluster variable in question to the total variance. These ratios can be interpreted as a correlation coefficient (i.e., how units are correlated within clusters).

In addition to answering substantive questions, ICCs are important parameters in a priori power calculations for experiments or precision estimates for surveys where the sampling design involves clusters. This entry primarily focuses on the one-way random effects model, as it is the most common. For simplicity, this entry also makes the assumption of balanced data (i.e., all clusters have the same number of cases).

Overview

In 1925, R. A. Fisher called the correlation of units within a cluster an intraclass correlation. This parameter is different from the more familiar interclass correlation because the ICC makes the assumption that the observations within a cluster share a common mean and standard deviation. To illustrate the difference, consider a population of 2 units per j = 1,2,3,…m clusters where each cluster is a row with two variables, one for each i = 1,2 units in the cluster, y1 and y2, with means

y¯1=Σy1m

and

y¯2=Σy2m,

and standard deviations

s1=(y1y¯1)2m

and

s2=(y2y¯2)2m.

The estimate of the interclass correlation is defined as

rinterclass=Σ(y1y¯1)(y2y¯2)ms1s2.

However, the ICC relies on a common mean

y¯c=Σ(y1+y2)2m,

and a common standard deviation,

sc=(y1y¯c)2+(y2y¯c)22m.

The ICC then takes a similar form as the interclass correlation and can be estimated from

rintraclass=Σ(y1y¯c)(y2y¯c)msc2.

Of course, there are often more than two units per cluster, so this formula must be extended to handle additional units, but (5) and (8) illustrate the relationship between the typical interclass correlation and the ICC. Note that in practice, ICCs are estimated in a way that incorporates degrees of freedom and will not equal (8).

The ICC and Analysis of Variance Tables

Analysis of variance (ANOVA) tables are often used to estimate ICCs. In the one-way random effects ANOVA, the outcome y for units i = 1,2,3,…n in each of j = 1,2,3,…m clusters is generated by the following linear model

yij=y¯+aj+eij,

where y¯ is the overall average of y. The average of y for the jth cluster is y¯j, and so aj=y¯jy¯ is a random variable associated with cluster j with a mean of 0 and variance σa2. Finally, eij=yijy¯j is a random within cluster error term with a mean of 0 and variance σe2. Since the terms aj and eij are not from the exhaustive population, but instead from randomly selected samples of clusters and units, they are noted as random effects and thus compose a random effects model. The values σe2 and σa2 are called “variance components” and the one-way ANOVA ICC is defined as

ρintraclass=σa2σa2+σe2.

The ICC can be estimated from an ANOVA table using the mean squares (MSs). To show the link between the ICC and the F-test, recall the F-test of the null hypothesis that all cluster means are equal is defined as the ratio of the MS between clusters (MSB) and the MS within clusters (MSW) with mn and mnm degrees of freedom,

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading