Lost in a random forest: Using Big Data to study rare events

Journal Article (Journal)

Sudden, broad-scale shifts in public opinion about social problems are relatively rare. Until recently, social scientists were forced to conduct post-hoc case studies of such unusual events that ignore the broader universe of possible shifts in public opinion that do not materialize. The vast amount of data that has recently become available via social media sites such as Facebook and Twitter—as well as the mass-digitization of qualitative archives provide an unprecedented opportunity for scholars to avoid such selection on the dependent variable. Yet the sheer scale of these new data creates a new set of methodological challenges. Conventional linear models, for example, minimize the influence of rare events as “outliers”—especially within analyses of large samples. While more advanced regression models exist to analyze outliers, they suffer from an even more daunting challenge: equifinality, or the likelihood that rare events may occur via different causal pathways. I discuss a variety of possible solutions to these problems—including recent advances in fuzzy set theory and machine learning—but ultimately advocate an ecumenical approach that combines multiple techniques in iterative fashion.

Full Text

Duke Authors

Cited Authors

  • Bail, CA

Published Date

  • December 27, 2015

Published In

Volume / Issue

  • 2 / 2

Electronic International Standard Serial Number (EISSN)

  • 2053-9517

Digital Object Identifier (DOI)

  • 10.1177/2053951715604333

Citation Source

  • Scopus