Skip to main content icon/video/no-internet

Big data is a loose concept referring to relatively large collections of data as well as the research that builds on them. Whereas in other domains big data is defined in terms of digital storage as data taking up several terabytes, within lifespan human development, it is sometimes used for data sets that are simply considerably larger than those typically used and require data analysis tools that are unusual for the relevant subfield. This entry introduces some examples of big data within lifespan human development research, then links big data to other relevant and recent developments; it ends with a few key controversies.

Types of Big Data

The big data approach of using advanced computational techniques to extract structured information from massive data sets found early adopters within the broader scientific community among biologists (e.g., the GenBank). Early success stories led to more organized initiatives, including the 2012 call by the National Science Foundation for big data in science. Within lifespan human development, four types of big data approaches can be defined.

The first concerns correlational studies with large cohorts of participants, mostly as part of government-funded projects relevant to physical and cognitive development, which can be traced back to the spread of standardized testing in the early 20th century (e.g., the Army classification test used on American soldiers in World War I). A more recent example is the Canadian Longitudinal Study on Aging, which gathers data from about 50,000 participants, about half providing regular physical samples (e.g., blood) in addition to completing questionnaires and other types of assessments. This research is conceptually continuous with previous smaller scale approaches, in that the research question and method are defined in advance but in recent years with a substantial increase in the sample sizes considered feasible.

A second type of research leading to large data sets involves more intensive data collection and annotation from a potentially smaller number of individuals. For instance, in the Speechome project, created around 2006 by Deb Roy at the Massachusetts Institute of Technology, a single child’s development was studied via audio and video recordings gathered in all rooms of the child’s house throughout all his waking hours between birth and 3 years of age. In this type of project, the data are often too plentiful, requiring preprocessing or analyses via semiautomatic procedures drawn from computer science.

A third type of big data initiative involves cumulating multiple data sets in comparable formats, such as via public repositories and cumulative meta-analyses, which virtually create big data by aggregation. For instance, the Child Language Development Exchange System (CHILDES) is a repository of child language samples that was set up in 1984 and has since received contributions from dozens of researchers. As of the end of 2015, the transcriptions contained a total of 59 million words (often accompanied by the raw audio or video data). Unlike the previous two types, big data from repositories allows reuse of the data to test hypotheses beyond the research questions that originally motivated the research.

Finally, one big data research approach not yet widespread in the field involves using naturalistic data that humans produce in their everyday technological life, as explained in the next subsection. Unlike all previous types, these data are not created with a specific research goal, and they may require substantial preprocessing to be informative.

...

  • Loading...
locked icon

Sign in to access this content

Get a 30 day FREE TRIAL

  • Watch videos from a variety of sources bringing classroom topics to life
  • Read modern, diverse business cases
  • Explore hundreds of books and reference titles

Sage Recommends

We found other relevant content for you on other Sage platforms.

Loading