PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 5, 2026Scientific Reports1 citationsOpen Access

Hybrid models of sparse and robust regression to solve heterogeneity problem in black pepper big data

PKPaavithashnee Ravi KumarOIOlayemi Joshua IbidojaMAMajid Khan Majahar Ali

Key Points

  • The study aims to address the heterogeneity problem in black pepper data using hybrid regression models to improve moisture content removal accuracy.
  • Utilized hybrid models integrating sparse regression techniques (elastic net, ridge, LASSO) and robust regression estimators (M Bi Square, M Hampel, M Huber, S and MM).
  • Identified the top-ranking variables influencing moisture content removal for black pepper.
  • Evaluated model performance under 2-sigma and 3-sigma limits before and after managing heterogeneity.
  • The hybrid Ridge model with M Bi-Square performed best before addressing heterogeneity.
  • After heterogeneity removal, the LASSO model with S estimator was the most effective across both 2-sigma and 3-sigma limits.

Abstract

Data analytics is increasingly important in agriculture, particularly in smart farming, enhancing decision-making and sustainability. Research on factors affecting moisture content removal in black pepper drying using solar dryers is crucial for cost reduction and improving product quality and quantity. This drying process involves numerous parameters, resulting in big data complexity. Heterogeneity among these parameters can introduce bias, leading to incorrect inferences, while multicollinearity and outliers impact model validation and interpretation. This study proposes hybrid models of sparse and robust regression to solve the heterogeneity problem using black pepper big data. Sparse regression techniques such as elastic net, ridge and LASSO are used to identify the 25, 35, 45, 55 and 100 highest-ranking variables for black pepper moisture content removal. These models are hybridized with robust regression estimators (M Bi Square, M Hampel, M Huber, S and MM) for handling outliers. Results indicate that before heterogeneity, the hybrid Ridge model with M Bi-Square performs best under both 2-sigma and 3-sigma limits. After heterogeneity removal, LASSO model with S estimator proves to be the most effective across both limits.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kumar et al. (2026) studied this question.

synapsesocial.com/papers/69d1fe07a79560c99a0a4845https://doi.org/10.1038/s41598-026-39290-0
Ask AI
Helpful
Bookmark
Share
View Full Paper