Skip to main content icon/video/no-internet

Bootstrapping is a computerized simulation operation that involves random resampling from the data set one is using to produce perhaps thousands of new data sets that have similar participant compositions to the original data set. For example, if a researcher had a data set containing 500 participants, the bootstrapping procedure would sample from the original data set to create new data sets, each of 500 participants. It is used to determine the statistical significance of parameter estimates whose significance cannot be tested using established statistical distributions (e.g., t, F, and chi-square distributions). One might use bootstrapping when data are nonnormal or the distribution of a parameter is difficult to anticipate (e.g., indirect effects in path analysis). Bootstrapping estimates the standard error (standard deviation of a set of statistical values such as a set of means from multiple samples) for a given statistical analysis, as well as confidence intervals for a parameter. Standard errors and confidence intervals are key to determining statistical significance.

This entry describes how and why bootstrapping samples are created, and how bootstrap samples indicate statistical significance of results. Other resampling techniques that are alternatives to bootstrapping are also introduced. The entry concludes with a review of software that can be used for bootstrapping.

How Bootstrap Samples Are Drawn

The key to bootstrapping is that the new bootstrap simulation samples are drawn with replacement. Imagine that a researcher places five ping-pong balls into a hat, the number “1” having been painted onto one ball, a “2” onto another, and the same for “3,” “4,” and “5.” Suppose further that the researcher plans to draw a random sample of three balls. The term with replacement refers to the fact that the ball that is drawn first is put back into the hat to possibly be drawn again. The same is done with the ball drawn second. Hence, the final three-ball sample conceivably could consist of balls 2, 2, and 4, or balls 1, 3, and 3, for example. If sampling is done without replacement, a ball selected into the sample is not returned to the hat; hence, the same number cannot appear more than once in the sample. Possible samples might, therefore, include balls 1, 2, and 3, or balls 2, 4, and 5.

Without replacement, drawing 500 people to create new data sets of 500 (a more realistic example for bootstrapping) would simply keep reproducing the original sample of the same 500 people. With replacement, however, the same participant in the original data set could be selected into a new sample, put back into the original sampling pool, then selected again into the same new sample. Other participants in the original sample may not be selected at all into one of the new samples. Suppose a researcher who has a sample of 500 participants wishes to draw 1,000 new bootstrap samples, each with the same sample size of 500. One of these bootstrap samples might include participants with the identification (ID) number 1, ID number 2, ID number 2 again (so that his or her data are used twice), ID number 5, ID number 6, and so forth, all the way to a total of 500 participants. Another bootstrap sample might include participants with ID number 3, ID number 3 again, ID number 4, ID number 7, ID number 7 again, ID number 7 a third time, ID number 10, and so forth.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading