Skip to main content icon/video/no-internet

Multiplicity Problem

Typically, a research project involves testing multiple research hypotheses. These research hypotheses could be evaluated using, for example, comparisons of means, bivariate correlations, or regressions. In fact, most studies consist of a mixture of several different types of test statistics. An important consideration when conducting multiple tests of significance is how to deal with the increased likelihood (relative to conducting a single test of significance) of falsely declaring one (or more) hypothesis(es) statistically significant (i.e., a Type I error). This is termed the multiplicity problem. This multiplicity problem is especially relevant to the topic of research design because the issues associated with the multiplicity problem relate directly to designing studies (i.e., number and nature of variables to include) and deriving a data analysis strategy (e.g., number of relationships to investigate) for the study. This entry introduces the multiplicity problem, in addition to discussing some of the strategies that have been proposed for dealing with the problem.

Introduction to the Multiplicity Problem

To help clarify the multiplicity problem, imagine a soldier who needed to cross fields containing land mines in order to obtain supplies. It is clear that the more fields the individual crosses, the greater the probability that they would activate a land mine; likewise, researchers conducting many tests of significance have an increased chance of erroneously finding tests that are statistically significant.

For instance, imagine that a researcher is interested in determining whether overall course ratings differ for lecture, seminar, or computer-mediated instruction formats. In this type of experiment, researchers are often interested in whether significant differences exist between any pair of formats. For example, do the ratings of students in lecture-format classes differ from the ratings of students in seminar-format classes? The multiplicity problem in this situation is that in order to compare each format in a pairwise manner, three tests of statistical significance need to be conducted (i.e., comparing the means of lecture vs. seminar, lecture vs. computer-mediated, and seminar vs. computer-mediated instruction). There are numerous ways of addressing the multiplicity problem (i.e., dealing with the increased likelihood of falsely declaring tests statistically significant). It is hoped that this entry will help clarify many of the issues surrounding the multiplicity issue.

Common Multiple Testing Situations

There are many different settings in which researchers conduct null hypothesis significance testing (NHST). The following are just a few of the more common settings where multiplicity issues arise:

  • conducting pairwise and/or complex contrasts in a linear model with categorical variables;
  • conducting multiple main effect and interaction tests in a factorial analysis of variance (ANOVA) or multiple regression setting;
  • analyzing multiple simple effect, interaction contrast, or simple slope tests when analyzing interactions in linear models;
  • analyzing multiple univariate ANOVAs following a significant multivariate analysis of variance;
  • analyzing multiple correlation coefficients;
  • assessing the significance of multiple factor loadings or factor correlations in factor analysis;
  • analyzing multiple dependent variables separately in linear models;
  • evaluating multiple parameters simultaneously in a structural equation model; and
  • analyzing multiple brain voxels for stimulation in functional MRI research.

Further, as stated previously, most studies involve a mixture of many different types of test statistics.

An important factor in understanding the multiplicity problem is understanding the different ways in which a researcher can “group” their tests of significance. For example, suppose in the study looking at whether student ratings differ across instruction formats that there was also another independent variable, the sex of the instructor. There would now be two “main effect” variables (instruction format and sex of the instructor) and potentially a hypothesis regarding the interaction between instruction format and sex of the instructor. On one hand, the researcher might want to “group” the hypotheses tested under each of the main effect (e.g., pairwise comparisons) and interaction (e.g., simple effect tests) hypotheses into separate “families,” where a family is defined as a group of related hypotheses that are considered simultaneously in the decision process. Therefore, control of the Type I error rate might be imposed separately for each family, or in other words, the Type I error rate for each of the main effect and interaction “families” is maintained at α (the maximum permissible rate of Type I errors for a specific test). On the other hand, the researcher may prefer to treat the entire set of tests for all main effects and interactions as one family, depending on the nature of the analyses and the way in which inferences regarding the results will be made. The central point here is that when researchers conduct multiple tests of significance, they must make important decisions about how their tests are related, and these decisions will directly affect the power and Type I error rates for both the individual tests and for the group of tests conducted in the study.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading