Think only when it is hard

October 2, 2026
With thinking on auto, hello is answered at once and strawberry is thought through.

Jev is a System One model: every field is answered at once, the way you answer 2 + 2 without stopping to think. That is fast, and it is right as long as the question is easy. Some questions are not.

In TypeLLM, a field can think before it answers. Until now you had to decide which fields think, and how much, when you wrote the schema. But whether a question is hard often depends on the input, not on the field. With "thinking": "auto", TypeLLM decides that on each call.

Two words, one question

Ask how many "r"s are in two words, and let each field choose:

response = client.generate(context='Counting "r"', questions={
    "hello":      {"type": "integer", "enum": list(range(10)),
                   "instructions": 'How many times does the letter "r" appear in "hello"?',
                   "thinking": "auto"},
    "strawberry": {"type": "integer", "enum": list(range(10)),
                   "instructions": 'How many times does the letter "r" appear in "strawberry"?',
                   "thinking": "auto"},
})

The API returns:

response.result           # {"hello": 0, "strawberry": 3}
response.thinking_effort  # {"hello": "none", "strawberry": "low"}
response.thinking         # {"strawberry": "Let's spell: s t r a w b e r r y. ... So 3 r's."}

hello has no "r" to find, so it answered at once. strawberry is the word models famously miscount, because they read tokens, not letters. It got a low effort, spelled the word out and counted three. The call spent 148 thinking tokens, all of them on strawberry.

How it decides

Before an "auto" field answers, TypeLLM asks the model one quick question about it, for this input: how much reasoning does it need? The answer is one of four levels:

Effort Thinks for up to
none answers at once
low 512 tokens
medium 2,048 tokens
high 4,096 tokens

The field then answers at that level, and response.thinking_effort reports the level each "auto" field got. All of this happens inside the same call.

Why not choose the effort yourself?

You still can: thinking_effort sets a fixed level for a field. But a fixed level has to be chosen in advance, for every input the field will ever see.

  • The same field can be easy or hard. Asking which of two dates is earlier is trivial for most pairs, and a trap for "June 8, 2026" against "07/06/2026" written day first, where models tend to read July 6. In our date example, "auto" gave that pair a low effort, and the model read 07/06/2026 as 7 June before it answered.
  • Thinking costs time and tokens. Reasoning tokens are billed and make the call slower. Thinking on every field pays that price on every call; "auto" pays it where the input needs it.
  • Easy fields stay fast. A field that gets none answers at once, like any field without thinking.

The model judges the difficulty, and it can misjudge. If you know a field is always hard, give it a fixed thinking_effort instead.

Try it

"thinking": "auto" is in TypeLLM 0.4.0 and later, and in the TypeLLM API. Run the letter count and date examples in the playground, or read the thinking docs.

← all posts