Documentation

Changelog

Install the latest release with pip install -U typellm.

0.3.0 (29 September 2026)

The hosted API: TypeLLMClient(api_key=...), or the TYPELLM_API_KEY environment variable, calls it with no GPU of your own. See Quick start.

Hosted calls retry up to twice on HTTP 429, 5xx other than 504 and connection errors, waiting as the service asks (Retry-After). max_retries=0 turns this off.

A string that reaches text_max_tokens is cut off there instead of failing the call. An input too long for thinking raises SGLangError with .status 400.

0.2.6 (28 September 2026)

Breaking: generate() returns a Generation

generate() returns .result, .thinking and .usage, like the HTTP API. client.last_usage, last_thinking and last_prompts are removed. See After a call.

Breaking: removed options

minimum, maximum and maxLength are no longer supported. TypeLLMClient(thinking_budget=...) is removed; set thinking_budget per field.

0.2.5 (28 September 2026)

temperature alone chooses decoding: 0, the default, gives the most likely answer, and above 0 samples. mode is deprecated.

0.2.4 (28 September 2026)

A call retries once when SGLang drops an idle connection.

0.2.3 (27 September 2026)

last_usage.input_tokens counts what you sent once: the context, the questions and the images. See the client reference.

0.2.2 (27 September 2026)

Breaking: thinking is set per field

TypeLLMClient(thinking=...) and run_schema(thinking=...) are removed and raise TypeError. Add "thinking": True, and optionally "thinking_budget", to the fields that should reason; the others answer at once. client.last_thinking returns each field's reasoning and last_usage.thinking_tokens counts it. See Thinking.

Breaking: shorter default text limit

text_max_tokens now defaults to 128 instead of 512. Pass a larger value to the client for longer answers.

0.2.1 (27 September 2026)

Calls with number and string fields make fewer requests.

0.2.0 (27 September 2026)

One client can be shared by many threads. generate() takes timeout, cancel and seed per call, and client.last_usage reports each call's requests and tokens. See the client reference.

Breaking: numeric token tables are removed

numeric_cache_dir and the typellm.numeric module are removed.

0.1.8 (26 September 2026)

permutations="auto" averages a balanced set of option orderings: K or 2K for K options, instead of K!. See permutation averaging.

0.1.7 (26 September 2026)

Breaking: execution= is removed

Starting with 0.1.7, execution= is no longer supported, and passing it raises TypeError. This applies to TypeLLMClient, generate() and run_schema(); the CLI no longer accepts --execution. Fields run in parallel by default. When a field declares depends_on, the request follows the dependency graph. Sequential execution and client.last_prompt are also removed; use client.last_prompts instead.

0.1.6 (25 September 2026)

Nullable fields, JSON answers with prefilled keys, and one prompt format for every field type.

0.1.5 (24 September 2026)

Image input, with numeric decoding and thinking batched across fields.

0.1.4 (23 September 2026)

Per-enum permutation averaging.

0.1.2 (22 September 2026)

The modules moved into a single typellm package, with no library logic changes.