Haupz Blog

... still a totally disordered mix

On LLMs Being Wrong

2026-10-11 — Michael Haupt

(Note: I originally wrote this text in 2024, over two years ago. While things have changed somewhat in the meantime, LLMs can still be seen committing errors like the one described here. The point is not the error in particular, it’s what it represents. Thus, I’m simply reusing the example from back then.)

Somewhere on LinkedIn, I came across one of those comical inaccuracies LLMs like ChatGPT are so notorious for producing. I tried for myself - behold.

The first image shows that ChatGPT got the simple request wrong the first time. Upon asking it to count again, it produced the right answer and a snippet of Python code it had generated to correctly compute the result. In the second image, the LLM responded with the correct solution repeatedly, until it suddenly lapsed to giving the correct answer for an entirely wrong reason.

Now, it’s tempting to brush that off as just another example of an LLM producing nonsense. However, and hear me out, I believe there’s reason for real, deep, lasting concern. I’ll reflect on this in two parts.

Part 1. The problem I asked the LLM to solve is dead simple. It also showed it can solve it correctly by recognising the question as a counting problem it could solve by generating and running some Python code. That it did not do so in the first place but instead gradient-descended on obvious nonsense expressed with great confidence is, as I see it, a huge problem.

We’re used to relying on technology that surrounds us. We expect that operating a light switch turns the light on (or off). We expect the car to start when we turn the key. We expect the calculator (or spreadsheet) to produce the correct result when entering numbers and operations. At the meta level, we expect the technology in question to work as expected, most of the time. In turn, if the expectations aren’t met, that indicates something is broken. This is an unwritten contract between humans and technology.

This LLM technology, as exemplified, violates that contract. It doesn’t work as expected, way too often. Exaggerating just a bit, it radiates a sense of being broken by default. This also shows in how it eventually makes up an entirely wrong reason for eventually giving the right answer: now “strawbrerry” is supposed to contain the letter “r” three times.

That’s in-your-face, visible-from-space wrong. That’s alternative facts, presented with the confidence we’ve come to know from bad people.

Part 2. Apologetics abound. Generative AI experts hasten to explain that more specialised models have a much smaller failure rate, or that AIs are simply irrational and can’t be compared to humans, or that “strawberry” might be too rare a word to have the statistical approximation of a correct result for the question generated.

One by one, in reverse order, here I go.

So “Strawberry” is more rare than other words. You gotta be kidding me, what a lame excuse. The response is still trivially, obviously wrong and not what a user would expect. Pointing to the statistical nature of how LLMs work is misleading - if the statistics at work generate wrong responses so easily, the statistics are wrong.

The point that AIs are irrational is a category error. I would even go as far as saying that the “I” in AI is a misnomer. This is statistical inference at a large scale, there is no ratio involved. The anthropomorphisation is all too easy and leads us to think about LLMs in the wrong categories. They’re software, and should be judged as such.

That more specialised models have a smaller failure rate is great. Seriously, thank Goodness for that. It means that this technology is able fulfil the aforementioned contract. It also means that LLMs are not there yet. From that, it follows that they should be used only with great care.

Summary. Bear with me, and forgive the snark. I’m basically just an optimistic skeptic. I thoroughly believe Generative AI has considerable - disruptive - potential to become a new generation of “brain extension” or tool for humans. The human/technology contract is important, though, and as long as we can’t trust LLMs to fulfil its part, we shouldn’t use them for serious business where critical matters are expected to just work.

Tags: the-nerdy-bit