Large datasets of fuel properties are indispensable for predictive combustion modeling and next-generation fuel design. However, resource-intensive experiments restrict existing databases to 200–500 compounds, capturing an infinitesimal fraction of the C1-20 hydrocarbon space. Furthermore, conventional rule-based and supervised learning extraction methods are constrained by poor scalability, domain-specific nomenclature, and weak contextual inference. To address these limitations, we introduce IgnitionGPT, a large language model fine-tuned on GPT-4.1-mini for the automated, concurrent extraction of three ignition metrics: Research Octane Number, Motor Octane Number, and Cetane Number. The model was trained on a human-annotated JSONL dataset of 304 sources (263 peer-reviewed articles, 41 patents) encompassing 581 diverse compounds. By evaluating IgnitionGPT directly against its zero-shot foundation, we isolate the impact of domain-specific fine-tuning. The model overcomes baseline overgeneralization (47.8% F1) to achieve saturated extraction accuracy on unseen data (i.e., 100% for the best model). Remarkably, it reaches this saturation on an 85% held-out test split using a mere 10% of the data for fine-tuning, demonstrating true robustness across heterogeneous literature. Ultimately, by open-sourcing our data and methods, this fine-tuning framework transitions chemical information retrieval from fragmented, rule-based heuristics to unified, concurrent extraction towards bridging the gap between experimental limitations and data-driven molecular design and modeling.
Abdulelah S. Alshehri (Sun,) studied this question.