PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 2026Journal of Network and Computer Applications3 citationsOpen Access

MLRan: A behavioural dataset for ransomware analysis and detection

View Full Paper
FOFaithful Chiagoziem OnwuegbucheSASunday Olaoluwa AdelodunAJAnca Delia Jurcut

Key Points

  • The aim is to introduce MLRan, a comprehensive dataset for effective ransomware detection that is reproducible and diverse.
  • Introduced a behavioural ransomware dataset, MLRan, with over 4800 samples across 64 families from 2006 to 2024.
  • Employed a two-stage feature selection process that reduced an initial 6.4 million features to 483 informative ones.
  • Evaluated various machine learning models on dataset accuracy, precision, and recall.
  • Achieved accuracy of 98.7%, precision of 98.9%, and recall of 98.5% with selected machine learning models.
  • Identified key indicators of malicious behaviour, such as registry tampering and API misuse, using SHAP and LIME.
  • Provided an open-source pipeline that supports dataset creation, feature extraction, and model training.

Abstract

Ransomware remains a critical threat to cybersecurity, yet publicly available datasets for training machine learning-based ransomware detection models are scarce and often have limited sample size, diversity, and reproducibility. In this paper, we introduce MLRan, a behavioural ransomware dataset, comprising over 4800 samples across 64 ransomware families and a balanced set of goodware samples. The samples span from 2006 to 2024 and encompass the four major types of ransomware: locker, crypto, ransomware-as-a-service, and modern variants. We also propose guidelines (GUIDE-MLRan), inspired by previous work, for constructing high-quality behavioural ransomware datasets, which informed the curation of our dataset. We evaluated the ransomware detection performance of several machine learning (ML) models using MLRan. For this purpose, we performed feature selection by conducting mutual information filtering to reduce the initial 6.4 million features to 24,162, followed by recursive feature elimination, yielding 483 highly informative features. The ML models achieved an accuracy, precision and recall of up to 98.7%, 98.9%, 98.5%, respectively. Using SHAP and LIME, we identified critical indicators of malicious behaviour, including registry tampering, strings, and API misuse. The dataset and source code for feature extraction, selection, ML training, and evaluation are available publicly to support replicability and encourage future research, which can be found at https://github.com/faithfulco/mlran . • MLRan: largest open-source behavioural ransomware dataset (64 families, 4.8K+ samples). • GUIDE-MLRan provides standardised guidelines for reproducible dataset creation. • Two-stage feature selection reduced 6.4M features to 483 without accuracy loss. • SHAP and LIME reveal key ransomware behaviours: strings, registry, and API. • Fully open-source pipeline: sandboxing, code for feature extraction and selection, ML training, and dataset.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Onwuegbuche et al. (2026) studied this question.

synapsesocial.com/papers/6a16d8b50f965e9c137bacfehttps://doi.org/10.1016/j.jnca.2026.104475
Ask AI
Helpful
Bookmark
Share
View Full Paper