Skip to main content icon/video/no-internet

Mixture models were created to decompose nonnormal univariate or multivariate distributions into distinct (i.e., a mixture of) normal distributions. In the social sciences, mixture models are typically referred to as person-centered analyses and used to identify distinct subpopulations of participants, referred to as profiles, differing from one another quantitatively and qualitatively on a series of indicators and/or relations among indicators. Quantitative research typically relies on statistics (e.g., correlations, regressions) obtained from a sample to infer what happens in the population from which the sample was extracted. Mixture models relax this assumption of population homogeneity (i.e., that the results can be generalized to all members of the population) to extract subpopulations (i.e., profiles) characterized by a distinct set of results. After briefly reviewing the history of mixture models, this entry describes the defining characteristics of these models, discusses several types of mixture models, and examines the main current limitations of these models.

History

Mixture models have been around since the end of the 1950s, although early developments can be traced back to Karl Pearson in 1894. The estimation of these models was greatly simplified in the late 1970s with the development of the expectation-maximization algorithm. However, the popularity of these models increased since the beginning of the 21st century following the development of user-friendly statistical packages (Mplus, Latent Gold, Stata-GLLAMM, and SAS-ProcTraj) and the creation of connections between mixture models and the structural equation modeling (SEM) framework. The resulting statistical framework is referred to as generalized SEM (GSEM).

Defining Characteristics of the GSEM Framework

The GSEM framework allows researchers to study relations occurring among any number of observed or latent continuous or categorical variables and to decompose these relations across different levels of analysis (e.g., occasions, individual, group). In SEM, a latent variable is used to infer the presence of unobservable constructs (e.g., self-esteem) from a series of indicators (e.g., responses to a questionnaire). These latent variables are corrected for the imprecise nature (i.e., random measurement error) of the indicators, resulting in the ability to estimate relations among latent variables corrected for unreliability. GSEM adds to SEM the ability to estimate categorical latent variables, reflecting latent (unobserved) subpopulations (profiles) differing from one another on any, or all, parts of a SEM model. These profiles present three defining characteristics.

Typological

Latent profiles are typological, resulting in a classification system whereby participants are assumed to come from distinct subpopulations characterized by a distinct set of results. This classification ability is an important advantage of person-centered analyses given its alignment with humans’ natural tendency to think in terms of categories.

Prototypical

Latent profiles are prototypical, resulting in a classification system whereby each participant is assigned a probability of membership into all profiles. This probability reflects the degree of similarity between the participant and the latent prototypes reflected by the profiles. Just as SEM decomposes observed variance into latent factors corrected from measurement error and unique components incorporating this error, this prototypical nature results in profiles corrected for classification error.

Exploratory

SEM typically involves the estimation of an a priori model, matching theoretical expectations, which is assessed in terms of model fit information, and can be contrasted with other a priori models. Mixture models involve a more exploratory, or inductive, process. In mixture models, the optimal solution is typically selected after consideration of alternative solutions including increasing numbers of profiles, which are then contrasted on various statistical indicators and in terms of theoretical, heuristic, and practical meaningfulness. For this reason, providing evidence of replicability (over groups or time points) and construct validity (in terms of associations with predictors or outcomes) is critical to the demonstration of the meaningfulness of a solution. This inductive nature does not mean that mixture models cannot be used for confirmatory (deductive) purposes. Any exploratory method can be used for confirmatory or exploratory purposes. However, confirmatory applications of mixture models are also available for research areas where theory and research are sufficient to support a priori hypotheses. Yet, even then, the a priori model would have to be contrasted with unconstrained models to document its value following a more exploratory methodology.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading