Can you trust the confidence?

October 10, 2026

A confidence is only useful if you can act on it. Set a threshold, handle the sure answers automatically, and send the unsure ones to a person. Decision models now all return a confidence with their answers. But can you trust it?

So we drew a digit that slowly turns from a 4 into a 9, and asked which one it is.

Four handwritten drawings: a 4 with a closed top, drawings A and B whose top is half round, and a 9.

Four is still a 4, and Nine is already a 9. Drawings A and B sit between them, where the 4's top has closed and only half rounded into a loop. They differ by a few percent of curve; side by side, they are hard to tell apart.

The question

Each model saw one drawing and one question, with two options:

response = client.generate(
    context="A handwritten digit.",
    images=["a.png"],
    questions={"digit": {
        "type": "string", "enum": ["4", "9"], "return_probabilities": True,
        "instructions": "Which digit is this handwritten character?",
    }},
)
response.result["digit"]
# {"value": "9", "probabilities": {"4": 0.41, "9": 0.59}, "confidence": 0.18}

We compared TypeLLM with the decision APIs that take image input: OpenAI's Decisions API (gpt-6-luna), and Cloudflare's Clef and Clef-flash through OpenRouter. Each got the same question, and each call ran at least five times.

The answers

Each row below is one model, and each column one drawing. The two bars are the model's probabilities for 4 and for 9, and the line under them its answer and confidence:

Four models' probabilities for 4 and 9 on four drawings. TypeLLM gives A 41% and 59%, B 35% and 65%. OpenAI Decisions gives A 73% and 27%, B 5% and 95%. Clef and Clef-flash are also shown.

TypeLLM reads A and B the way a person would: both unclear, close to an even split, leaning a little further to 9 on B. Its confidence stays low for both, 0.18 and 0.30, because neither drawing is clearly a 4 or a 9.

The OpenAI Decisions API calls A a 4, and B a 9 with a confidence of 0.90: two nearly identical drawings, two opposite answers, one of them sure. Clef-flash flips the same way. Clef reads the two alike, but as a fairly sure 9.

What it does to a threshold

Say you act on answers with a confidence of 0.8 or more, and send the rest to a person:

  • TypeLLM sends both drawings to a person, which is what you want: they are hard to read, and a person settles them.
  • OpenAI Decisions sends A to a person, and files B as a 9 on its own.
  • Clef-flash does the same: A to a person, B filed as a 9.

The drawings are nearly the same; whether one is handled automatically should not hang on a few percent of curve. With TypeLLM it does not: inputs that look alike get alike confidence, so a threshold means what you set it to mean.

So, can you trust it?

Before you act on a confidence, check two things. An unclear input should get a low one, and inputs that look alike should get about the same one. On these drawings, TypeLLM does both: it is unsure of A and B, and about as unsure of each. That is a confidence you can set a threshold on.

Try it

The drawings and every call are in a notebook: open it in Colab, paste your keys, and run it. Or draw your own, or take a photo of a digit, and ask for probabilities: any enum or boolean field takes "return_probabilities": true and returns a confidence with its answer. A when condition can then act on it in the same call. See Probabilities and confidence, or the Refund only when sure example.

← all posts