PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 6, 2026Scientific Reports0 citationsOpen Access

A score based likelihood ratio framework for deepfake image identification in forensic science

TGTianli GuoJLJisong LiYTYunqi Tang

Key Points

  • The study aims to create a reliable system for identifying deepfake images in forensic contexts.
  • Developed a score-based likelihood ratio system for deepfake identification.
  • Utilized the FaceForensics++ dataset with video-level data splitting.
  • Employed kernel density estimation for modeling score distributions.
  • Applied calibration techniques for optimal performance evaluation.
  • Tested generalization across multiple unseen datasets.
  • Capsule detector achieved the highest performance with AUC of 0.983.
  • Model showed low misleading evidence rates (RMEP = 0.069, RMED = 0.092).
  • Error control was effective with an EER of 0.0804.
  • Generalization testing yielded AUCs between 0.621 and 0.783 on unseen datasets.
  • System demonstrates potential but needs further validation for real-world use.

Abstract

Abstract This paper proposes a score-based likelihood ratio system for forensic identification of deepfake images, addressing challenges in digital media identification due to rapid deepfake development. Built on the FaceForensics + + dataset, the system prevents data leakage via video-level splits (training, validation, selection, calibration, and test sets). Among six candidate models, the Capsule detector demonstrates the most robust performance (AUC = 0. 983). Score distributions of real and fake samples are modeled using kernel density estimation, with optimal bandwidths selected through a two-stage search (real: 0. 004, fake: 0. 003). Extreme LR values are bounded using the empirical lower and upper bounds method (− 2. 3634 to 1. 9933), and PAV calibration is applied to optimize the calibration performance of the LR system. On the FF + + test set, the system exhibits favorable performancewith forensic practice expectations: low misleading evidence rates (RMEP = 0. 069, RMED = 0. 092), good error control (EER = 0. 0804), and reduced decision loss after calibration (the cost log-likelihood ratio from 0. 2899 to 0. 1625). Generalization tests on five unseen datasets (Celeb-DF-v1/v2, DFDCPₘethodA/B, UADFV) yield AUCs between 0. 621 and 0. 783—highest on UADFV (0. 783), stable on DFDCP, weaker on Celeb-DF. The results show that at the moment, the technique shows potential for forensic-oriented deepfake identification, but requires further validation across diverse real-world scenarios before practical forensic application.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Guo et al. (2026) studied this question.

synapsesocial.com/papers/69aa701a531e4c4a9ff598fchttps://doi.org/10.1038/s41598-026-42176-w
Ask AI
Helpful
Bookmark
Share
View Full Paper