Open educational resource that introduces how to observe and study human behavior in the digital age using digital trace (often “big”) data. It mentions what makes these data powerful, what can go wrong when repurposing records created for non-research purposes, and how to select research strategies that fit the data and the question. The slides provide a compact framework based on Salganik’s ten common characteristics of big data—three frequent strengths (big, always-on, nonreactive) and seven recurring risks (incomplete, inaccessible, nonrepresentative, drifting, algorithmically confounded, dirty, sensitive). For each characteristic, the slides highlight the main implication for inference, measurement, and governance (privacy, access, re-identification risk). The slides also outline Salganik’s three practical research strategies with digital trace data: counting things (including scalable measurement workflows such as “label a little, model a lot”), forecasting/nowcasting (including baseline comparisons and performance decay under drift), and approximating experiments (natural experiments, matching, and core assumptions for causal claims). Intended audience: students and researchers in psychology, sociology, computational social science, and related behavioral sciences; practitioners who use platform or administrative data for analysis and policy.
David Abián (Thu,) studied this question.