RMUNN: Even a hypothetical 10% hallucination rate is too high for reliability-seeking readers; next-token models may not guarantee logic
The Gist
RMUNN says these models still make things up too often. Even if they only messed up one time in ten someday, that is still too shaky if you want something you can basically count on, and he doubts pure next-word predictors will ever be rock-solid reasoners. This packet is a steelman reconstruction of RMUNN's Hacker News reply for LogicFirst import. It is not an endorsement of RMUNN, andrewdb, or any position in the thread.
Conclusion
Hallucination risk remains material for reliability-seeking readers even under a charitable 10% future rate, and next-token statistical architecture supplies a residual reason not to expect guaranteed logical reliability from present-style LLMs.
Premises
- Present models remain too likely to hallucinate for a reader who wants high reliability rather than "often good enough."
- Even if, at some distant future date, the hallucination rate fell to about 10% (so that an AI-constructed argument were about 90% likely to be well constructed), that residual error rate would still be too high for a reader seeking near-certainty.
- RMUNN doubts that statistically next-token language models will reach guaranteed logical reliability while they remain probability-of-next-token engines rather than systems organized around facts and logical reasoning.
- That architecture doubt is offered as a residual reason to expect lasting non-guarantee, not as a proof that no future AI system of any kind can be reliable.
- Because correctness is not guaranteed and expected error remains material under the stated hypothetical, AI origin continues to mark elevated hallucination risk relative to a reliability-seeking reader's preferred baselines.
Assumptions
- The 10% figure is RMUNN's illustrative ceiling for a distant future, not a measured present rate for all tasks. "Guaranteed logic" means reliability suitable for a reader seeking near-100% trust allocation, not mathematical undecidability. Research residual: measured hallucination and factuality rates vary sharply by task and benchmark. Grounded summarization leaderboards sometimes report low single-digit or sub-2% rates for strong models on easier setups, while harder factuality and definitive-answer suites often show much higher error, and Vectara's harder multi-domain summarization set has…
- Differs: Easy-bench rates and tool grounding can be far below 10%; do not treat 10% as the unique present rate.
- Differs: Easy-bench and tooling can undercut a blanket present 10% claim.
Analysis
Overall strength: Moderate. Argument type: Inductive.
Premise Strength
- Present models remain too likely to hallucinate for a reader who wants high reliability rather than "often good enough." (Moderate) — Plausible and consistent with known task-dependent hallucination variance, but rests on a single informal (Hacker News) testimonial source and an unspecified reliability threshold, limiting its evidentiary weight beyond a personal risk-tolerance claim.
- Even if, at some distant future date, the hallucination rate fell to about 10%... that residual error rate would still be too high for a reader seeking near-certainty. (Moderate) — Nearly analytic once the reader's goal is stipulated as near-100% trust, which makes it internally consistent but low in independent diagnostic content; its persuasive force depends heavily on accepting an unquantified 'near-certainty' standard as reasonable.
- RMUNN doubts that statistically next-token language models will reach guaranteed logical reliability... (Weak) — A plausible but contested and largely unfalsifiable architectural intuition; it does not engage with counter-evidence from retrieval-augmented, tool-using, or hybrid neurosymbolic systems that may decouple reliability from the base generative mechanism.
- That architecture doubt is offered as a residual reason... not as a proof that no future AI system of any kind can be reliable. (Strong) — A well-executed epistemic hedge that appropriately limits the scope of the architectural claim, avoiding overreach and improving the argument's overall calibration.
- Because correctness is not guaranteed and expected error remains material... AI origin continues to mark elevated hallucination risk... (Moderate) — Follows reasonably from the prior premises once granted, but largely restates the conclusion in definitional terms conditioned on an unspecified reader baseline rather than adding independent support.
Potential Fallacies
- Unfalsifiable threshold framing (P2, P5) — Because 'too high' and 'near-certainty' are never quantified, essentially any nonzero error rate can be declared insufficient. This makes the core claim resistant to being satisfied by any conceivable future improvement, since the reader-preference bar is never fixed. The argument's own stipulation (a reader who wants near-100% trust) makes this near-tautological rather than a claim that could be empirically falsified.
- Weak generalization from architecture to outcome (P3) — Inferring a durable ceiling on 'guaranteed logical reliability' from the fact that a system is built to predict the next token risks conflating an implementation-level mechanism with a functional capability ceiling. This inference does not directly engage hybrid systems (retrieval augmentation, tool use, verification layers) that already decouple practical reliability from the raw generative mechanism.
- Domain-specific benchmark generalized beyond its scope (P2 and its supporting assumption) — The supporting evidence (Vectara-style grounded summarization benchmarks) measures a narrower task than 'AI-constructed arguments' or general logical reasoning. Using hard-task summarization error rates to anchor a broader claim about argument reliability risks overstating how representative that evidence is.
- Unstated asymmetric standard (implicit special pleading) (P5 and the overall framing) — The argument never compares the demanded reliability standard to human error rates on equivalent reasoning or argument-construction tasks, even though no human source offers logical guarantees either. Without this comparison, it is unclear whether AI is being held to a distinctive standard not applied elsewhere.
Counterarguments
- Conclusion / P5 (High impact) — No human expert, peer-review process, or scientific method offers guaranteed logical reliability either, yet these are routinely trusted; applying a 'guaranteed reliability' standard uniquely to AI without a comparative human baseline risks an asymmetric or double standard rather than a substantive critique.
- P3 (High impact) — Deployed AI systems increasingly combine next-token generation with retrieval grounding, tool use, and verification layers; reliability may be better understood as an emergent property of the full system stack rather than an intrinsic ceiling of the base architecture, which would undercut the architectural doubt's practical relevance.
- P2 / supporting assumptions (Medium impact) — The 10% anchor is drawn from a harder multi-domain benchmark while the same assumptions acknowledge sub-2% rates on easier tasks; foregrounding the higher figure risks selectively emphasizing the most alarming data point rather than reflecting the full distribution of evidence.
- P1, P2, P5 (Medium impact) — Pushed to its logical extreme, a standard where any nonzero error disqualifies a system as 'reliable enough' would also disqualify calculators, compilers, and human institutions, none of which are error-free — suggesting the demanded standard, if applied consistently, leads to an implausible universal skepticism.
- Conclusion (Medium impact) — The 'reliability-seeking reader' is treated as a fixed, near-universal category, but most real-world LLM use occurs on a 'good enough' spectrum for lower-stakes tasks; the argument's scope may be much narrower in practice than its general framing suggests.
Suggested Improvements
- Quantify the reliability threshold — Specify what error rate, if any, would satisfy a 'reliability-seeking reader,' or explicitly frame the standard as task- and stakes-dependent. Without an operational threshold, the central claim cannot be falsified by any future improvement in model performance, which weakens its evidentiary and persuasive value.
- Add a comparative human baseline — Compare the cited AI hallucination/error rates to human error rates on equivalent reasoning or argument-construction tasks. This would clarify whether the demanded reliability standard is uniquely applied to AI or reflects a generally applicable bar, addressing the argument's most significant unaddressed vulnerability.
- Engage hybrid and tool-augmented architectures — Explicitly discuss retrieval-augmented generation, verification layers, and tool use as potential mitigations to the architectural doubt in P3, rather than treating raw next-token prediction as the sole determinant of reliability. This would strengthen the architectural claim by addressing the most obvious and technically informed counterargument rather than leaving it unaddressed.
- Match benchmark evidence to the claim's scope — Use benchmarks that measure general argument construction or multi-step reasoning validity, not solely grounded summarization factuality, when supporting claims about 'AI-constructed arguments.' This would reduce the risk of generalizing from a narrower, task-specific benchmark to a broader claim about logical reliability.
- Upgrade or contextualize sourcing — Supplement the single informal (Hacker News) testimonial source with more systematic references, or more prominently flag its epistemic status within the premises themselves rather than only in supporting notes. This would improve the evidentiary weight of P1 and P3 for audiences unfamiliar with the informal discourse context from which the claims originate.
Scenario Tests
- Retrieval-augmented and verification-layered systems reduce effective end-to-end error well below the raw model's next-token hallucination rate. (Challenges) — Would render the architectural doubt in P3 practically less relevant even if theoretically valid, since realized reliability would depend more on the surrounding system than on the base generative mechanism.
- Future benchmarks show hallucination rates converging toward or below 10% broadly across hard, multi-domain reasoning tasks. (Supports) — Would corroborate that 10% is a realistic 'charitable ceiling' worth taking seriously, lending empirical weight to P2's hypothetical.
- The relevant reader is a casual or exploratory user rather than one demanding near-certainty. (Challenges) — Would show the argument's conclusion applies only to a narrow subset of users, despite its general framing, limiting its practical scope.
- The same 'guaranteed logical reliability' standard is applied to human expert reasoning and peer review. (Challenges) — Would reveal that no real-world epistemic source meets this standard, suggesting the argument's demand may be either vacuous or selectively applied to AI.
Coherence & Relevance
The argument is internally coherent and notably well-hedged: it explicitly avoids overclaiming (P4), acknowledges task-dependent variance in hallucination rates (A1–A3), and clearly distinguishes a normative reliability-risk claim from a more speculative architectural claim. Its principal weaknesses are an unquantified reliability threshold that renders the core claim difficult to falsify, an absence of comparison to human error baselines, and limited engagement with hybrid or tool-augmented architectures that could bear on whether the architectural doubt translates into practical unreliability. These gaps do not amount to formal invalidity but do limit the argument's persuasive force beyond readers who already share its stipulated near-certainty standard.
- Present models remain too likely to hallucinate for a reader who wants high reliability rather than "often good enough." (Strong) — Relevant to establishing baseline concern, but the underlying reliability threshold is undefined, leaving the connection to the conclusion partly rhetorical rather than fully evidential.
- Even if... the hallucination rate fell to about 10%... that residual error rate would still be too high for a reader seeking near-certainty. (Strong) — Tightly connected to the conclusion's risk claim given the stipulated reader goal, though its force depends on accepting an unquantified 'near-certainty' bar as the correct standard.
- RMUNN doubts that statistically next-token language models will reach guaranteed logical reliability... (Moderate) — Connects to the conclusion's architectural claim but does not engage hybrid/tool-augmented systems that could decouple reliability from raw architecture, leaving a gap between the premise and the conclusion's forward-looking scope.
- That architecture doubt is offered as a residual reason... not as a proof... (Strong) — Functions well as a scope-limiter, tightly and appropriately connecting P3 to the conclusion without overreach.
- Because correctness is not guaranteed and expected error remains material... AI origin continues to mark elevated hallucination risk... (Moderate) — Largely restates the conclusion given the prior premises; it adds little independent support and inherits the unresolved threshold and baseline-comparison gaps from P1/P2.