PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 15, 20260 citationsOpen Access

Almost Correct, Almost Useless? A Probabilistic-Formal Synthesis for the LLM Era

View Full Paper
ASAlfredo Sepulveda-Jimenez

Key Points

  • The research aims to address deficiencies in the mathematical foundations of using large language models with formal verification in software systems.
  • Assessment of three main claims in Meyer's argument regarding system reliability, correctness, and hallucinations in AI.
  • Development of a correlated-failure model for system architecture and a definition of ε-correctness.
  • Creation of a conformal-verification framework for finite-sample correctness in LLM outputs.
  • Introduction of a correlated-failure model that enhances architectural fault containment.
  • Establishment of a new definition of ε-correctness, broadening the scope of traditional Hoare logic.
  • Demonstration that hallucination rates can be controlled under verifier-conditioned generation, providing empirical support for the framework.

Abstract

In his Communications of the ACM article and the May 2026 ACM TechTalk Software Ver-ication in the Age of Articial Intelligence, Mey26 argues that since modern generativeAI is statistical, while professional software requires categorical correctness, the only viablefuture is a marriage of large language models (LLMs) with formal specication and proof.We endorse the directional claim but show that three of the load-bearing premises of theargument are mathematically defective. First, the Dijkstra-style compounding argumentpN, which Meyer deploys to motivate exponential decay of system reliability with modulecount, presupposes per-module independence of failure events, an assumption empiricallyrefuted by KL86 and structurally incompatible with the way modules in real systems areconstructed, deployed, and defended. Second, the binary works/useless dichotomy thatanchors the A/B/C taxonomy conates per-input total correctness in the sense of Hoa69with distributional reliability on a non-trivial input measure μ; once disambiguated, al-most correct is not only meaningful but is the dominant operating regime for the entireB class. Third, the claim that hallucinations are essential, not accidental to modern AIbegs the question by xing the generator's training objective and decoding regime; underverier-conditioned generation, hallucination rates are by construction zero relative to theverier's logic, as demonstrated empirically by Alp25, Yan+25, and SYA25. We replacethe defective scaolding with (i) a correlated-failure model with architectural fault contain-ment, (ii) a denition of ε-correctness extending Hoare logic to distributional reliability, and(iii) a conformal-verication framework that delivers nite-sample correctness guaranteeson LLM-generated artifacts. The synthesis preserves Meyer's central insight while quanti-fying the genuine open problem: the correlated stochastic verication that arises when bothspecication and implementation are generated by statistically dependent models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Alfredo Sepulveda-Jimenez (2026) studied this question.

synapsesocial.com/papers/6a06b81ce7dec685947aab06https://doi.org/10.5281/zenodo.20162450
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Algorithm Almost Correctness, Almost Uselessness? A Probabilistic-Formal Synthesis for the LLM Era2026
  2. 2Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks2025
  3. 3Natural Language Has No Formal Semantics. Verify the Decision Anyway.2026
  4. 4The Forensic Standard: Deterministic Verification in Legal AI2026
  5. 5When Verification Explores Too Far: Semantic Coverage and Validity in LLM-Generated Code Checks2026