PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 30, 2026Electronics0 citationsOpen Access

A Comparative Performance Study of Host-Based Intrusion Detection Using TextRank-Based System Call Preprocessing and Deep Learning Models

View Full Paper
HYHyunwook YouCPChulgyun ParkDSDongkyoo Shin

Key Points

  • The aim is to compare different models for host-based intrusion detection using a deep learning approach.
  • Developed a TextRank-based preprocessing pipeline using the LID-DS 2021 dataset.
  • Compared five deep learning models: RF, LSTM, CNN + LSTM, BiLSTM, and CNN + BiGRU.
  • Selected three scenarios covering various attack categories for analysis.
  • Deep learning models achieved F1-scores ranging from 0.90 to 0.94.
  • Random Forest showed lower performance with F1-scores of 0.55 to 0.63.
  • CNN + BiGRU produced the strongest results among the evaluated models.

Abstract

Host-based intrusion detection systems (HIDSs) can address the limitations of network-based detection by analyzing system calls and other low-level events. Many existing benchmark datasets remain inadequate for evaluating modern attacks because they were built in outdated environments and cover only a limited set of attack behaviors. To address this gap, this study builds a TextRank-based preprocessing pipeline on the LID-DS 2021 dataset and compares five end-to-end pipelines: Random Forest (RF), Long Short-Term Memory (LSTM), Convolutional Neural Network(CNN) + LSTM, LSTM, Bidirectional LSTM (BiLSTM), and CNN + Bidirectional Gated Recurrent Unit (BiGRU). Of the 15 scenarios in the dataset, six multi-stage attacks were excluded, and three representative scenarios were selected based on attack-category coverage and suitability for single-chunk host-level detection. Within these three selected scenarios and same-scenario file-level splits, the deep learning pipelines achieved F1-scores of 0.90–0.94, whereas RF ranged from 0.55 to 0.63. Among the evaluated pipelines, CNN + BiGRU produced the strongest overall results. These findings indicate that, under this constrained evaluation setting, sequential deep learning pipelines can be effective for scenario-specific system-call-based HIDS; however, broader generalization to unseen attacks or to the full LID-DS 2021 scenario set remains unverified.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

You et al. (2026) studied this question.

synapsesocial.com/papers/69f2a4f18c0f03fd67764106https://doi.org/10.3390/electronics15091856
Ask AI
Helpful
Bookmark
Share
View Full Paper