Confidence
TypeLLM does not return one fixed confidence score. With return_probabilities, an enum or boolean field returns the probability of every candidate, and you define the confidence your application needs from that distribution: how sure the model is of its answer, how close the runner-up came, or how spread out the whole distribution is.
response = client.generate(
context=ticket,
questions={
"category": {
"type": "string",
"enum": ["billing", "bug", "account", "feature_request"],
"return_probabilities": True,
},
},
)
answer = response.result["category"]
# {"value": "billing",
# "probabilities": {"billing": 0.81, "bug": 0.02, "account": 0.15, "feature_request": 0.02}}Ways to measure it
Each of these reads the same probabilities map; pick the one that matches the decision you make with the answer.
import math
p = answer["probabilities"]
ranked = sorted(p.values(), reverse=True)
top = ranked[0] # how likely the chosen answer is
margin = ranked[0] - ranked[1] # how far ahead of the runner-up it is
spread = -sum(x * math.log(x) for x in ranked if x > 0) / math.log(len(ranked))
# 0: all on one answer, 1: evenly spread- Top probability is the simplest: the chance the model gives its own answer. Use it when any wrong answer costs the same.
- Margin catches close calls between two answers, such as
billingagainstaccountabove, even when the top probability looks high. - Spread (normalized entropy) tells you whether the model is unsure across many candidates, not just two.
- For a boolean, the probability of
Trueis itself a score: set the threshold that trades false positives against false negatives the way your use needs.
Act on it
A common pattern accepts confident answers automatically and sends the rest to a person, or to a second pass with thinking turned on.
if top >= 0.9 and margin >= 0.5:
route(ticket, answer["value"]) # confident: act on it
else:
send_to_review(ticket, p) # unsure: let a person decideWhen several fields must all be right, combine them, for example by taking the lowest top probability among them, so one uncertain field is enough to send the item for review.
Choose thresholds on your data
- Probabilities describe how the model weighs the candidates you gave it, not a guarantee that an answer is factually correct. Check what a threshold means on your own data before you rely on it.
- Label a sample of real inputs, run them, and for each threshold count how many answers it accepts and how many of those are right. Choose the threshold that gives the accuracy you need with as many accepted answers as possible.
- Set thresholds per field: an easy field and a hard one rarely need the same.
- Option order can shift probabilities. Permutation averaging removes that effect, at the cost of more input tokens for the field.
See Probabilities for the return format and sampling.