PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 10, 2019Social Psychological and Personality Science1,052 citations

An MTurk Crisis? Shifts in Data Quality and the Impact on Study Results

View Full Paper
MCMichael S. ChmielewskiSKSarah C. Kucker

Key Points

  • To evaluate whether data quality on Amazon's Mechanical Turk decreased after summer 2018 and assess the consequences on psychological study findings and measurement validity.
  • Implemented a four-wave naturalistic experimental design assessing data quality pre-, during, and post-summer 2018.
  • Evaluated participant failure rates on response validity indicators alongside the reliability and construct validity of a standard personality measure.
  • Tested the replicability of well-established psychological findings across waves and evaluated data-screening mitigation strategies.
  • Participant failure rates on response validity indicators increased significantly during and after summer 2018.
  • Reliability and validity of a widely used personality measure declined, leading to failed replications of well-established empirical findings.
  • Implementing response validity indicators and systematic data screening successfully mitigated the observed quality deficits.

Abstract

Amazon’s Mechanical Turk (MTurk) is arguably one of the most important research tools of the past decade. The ability to rapidly collect large amounts of high-quality human subjects data has advanced multiple fields, including personality and social psychology. Beginning in summer 2018, concerns arose regarding MTurk data quality leading to questions about the utility of MTurk for psychological research. We present empirical evidence of a substantial decrease in data quality using a four-wave naturalistic experimental design: pre-, during, and post-summer 2018. During and to some extent post-summer 2018, we find significant increases in participants failing response validity indicators, decreases in reliability and validity of a widely used personality measure, and failures to replicate well-established findings. However, these detrimental effects can be mitigated by using response validity indicators and screening the data. We discuss implications and offer suggestions to ensure data quality.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chmielewski et al. (2019) studied this question.

synapsesocial.com/papers/69d97f092a25b240b7a3ca5ahttps://doi.org/10.1177/1948550619875149
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Handbook of Personality Theory and Research.1970 · 2,176 citations
  2. 2Tripartite model of anxiety and depression: Psychometric evidence and taxonomic implications.1991 · 3,586 citations
  3. 3Crowdsourcing user studies with Mechanical Turk2008 · 1,971 citations
  4. 4Personality and Psychopathology1999 · 101 citations
  5. 5An Evaluation of Amazon’s Mechanical Turk, Its Rapid Rise, and Its Effective Use2018 · 792 citations