RMUNN: Even a hypothetical 10% hallucination rate is too high for reliability-seeking readers; next-token models may not guarantee logic

The Gist

RMUNN says these models still make things up too often. Even if they only messed up one time in ten someday, that is still too shaky if you want something you can basically count on, and he doubts pure next-word predictors will ever be rock-solid reasoners. This packet is a steelman reconstruction of RMUNN's Hacker News reply for LogicFirst import. It is not an endorsement of RMUNN, andrewdb, or any position in the thread.

Conclusion

Hallucination risk remains material for reliability-seeking readers even under a charitable 10% future rate, and next-token statistical architecture supplies a residual reason not to expect guaranteed logical reliability from present-style LLMs.

Premises

  1. Present models remain too likely to hallucinate for a reader who wants high reliability rather than "often good enough."
  2. Even if, at some distant future date, the hallucination rate fell to about 10% (so that an AI-constructed argument were about 90% likely to be well constructed), that residual error rate would still be too high for a reader seeking near-certainty.
  3. RMUNN doubts that statistically next-token language models will reach guaranteed logical reliability while they remain probability-of-next-token engines rather than systems organized around facts and logical reasoning.
  4. That architecture doubt is offered as a residual reason to expect lasting non-guarantee, not as a proof that no future AI system of any kind can be reliable.
  5. Because correctness is not guaranteed and expected error remains material under the stated hypothetical, AI origin continues to mark elevated hallucination risk relative to a reliability-seeking reader's preferred baselines.

Assumptions

Analysis

Overall strength: Moderate. Argument type: Inductive.

Premise Strength

Potential Fallacies

Counterarguments

Suggested Improvements

Scenario Tests

Coherence & Relevance

The argument is internally coherent and notably well-hedged: it explicitly avoids overclaiming (P4), acknowledges task-dependent variance in hallucination rates (A1–A3), and clearly distinguishes a normative reliability-risk claim from a more speculative architectural claim. Its principal weaknesses are an unquantified reliability threshold that renders the core claim difficult to falsify, an absence of comparison to human error baselines, and limited engagement with hybrid or tool-augmented architectures that could bear on whether the architectural doubt translates into practical unreliability. These gaps do not amount to formal invalidity but do limit the argument's persuasive force beyond readers who already share its stipulated near-certainty standard.

View this argument on LogicFirst.ai