API reference
The hosted API has one endpoint that does the work, POST /v1/generate. The Python client calls it for you; any language can call it directly.
Authentication
The base URL is https://api.typellm.ai. Send your key from Dashboard, API keys as a bearer token on every request:
curl https://api.typellm.ai/v1/models \
-H "Authorization: Bearer $TYPELLM_API_KEY"POST /v1/generate
Answers every question about the context in one call.
curl https://api.typellm.ai/v1/generate \
-H "Authorization: Bearer $TYPELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"context": "Receipt from Hilton London. Total: £324.50",
"questions": {
"total": {"type": "number", "instructions": "Total in GBP."},
"expense_type": {"type": "string", "enum": ["meal", "travel", "equipment"]},
"policy_ok": {"type": "boolean", "instructions": "Within the travel policy?",
"thinking": true, "thinking_budget": 1024}
},
"options": {"temperature": 0, "seed": 42},
"timeout": 60
}'Request body
| Field | Type | Description |
|---|---|---|
context | string, required | The text to answer from. |
questions | object, required | Field names mapped to question definitions (below). At most 64. |
images | array of strings | Base64 data:image/... URIs; no URLs. At most 8, each up to 10 MB. |
model | string | A public model; typellm-latest when left out. |
options.temperature | number ≥ 0 | 0 (the default) gives each question its most likely answer; above 0 samples at that temperature. |
options.seed | integer | Fixes the service's random choices. Results can still vary slightly with batching. |
timeout | number | Seconds for the call: 60 by default, at most 90. |
Question definition
Each question is a JSON Schema field with a few TypeLLM keys. maxLength, minLength, pattern, format, minimum and maximum are not accepted; string answers stop at 128 tokens.
| Key | Type | Description |
|---|---|---|
type | string | string, number, integer or boolean; ["string", "null"] allows no answer. See Output types. |
enum | array | The only answers allowed. |
instructions | string | What to answer; description is read when it is missing. |
thinking | boolean | Reason before answering. See Thinking. |
thinking_budget | integer | Cap on that reasoning: 1,024 tokens by default, at most 4,096. |
return_probabilities | boolean | On enum and boolean questions: answer with {"value", "probabilities"}. See Probabilities. |
permutations | boolean | On enum questions: average over a balanced set of option orders. Off by default. |
depends_on | array of strings | Questions answered first and shown to this one. See Dependencies. |
Response
{
"id": "gen_...",
"model": "typellm-latest",
"result": {"total": 324.5, "expense_type": "travel", "policy_ok": true},
"thinking": {"policy_ok": "The hotel is in London, where the cap is..."},
"usage": {"input_tokens": 410, "thinking_tokens": 204},
"elapsed": 3.1
}| Field | Type | Description |
|---|---|---|
id | string | The call's id; quote it when you contact us. |
model | string | The public model that answered. |
result | object | One typed answer per question. |
thinking | object | The reasoning of each question that thought. |
usage.input_tokens | integer | What you sent, counted once: the context, the questions as JSON, and the images. Billed. |
usage.thinking_tokens | integer | Reasoning tokens. Billed. Answers are free. |
elapsed | number | Seconds the service spent, queueing included. |
GET /v1/models
The public models your key can use: {"models": ["typellm-latest"]}.
Errors
Errors come with an HTTP status and a body like:
{"error": {"type": "insufficient_balance", "message": "the balance has run out; ..."}}| Status | Type | When |
|---|---|---|
400 | invalid_request, limit_exceeded, invalid_image, context_length_exceeded | A malformed body or question, a limit exceeded, an unreadable image, or an input over 128K tokens. |
401 | unauthorized | A missing, unknown or revoked key. |
402 | insufficient_balance | The balance has run out. Add credits in Dashboard, Billing. |
404 | unknown_model | No such public model. |
413 | http_error | The body is over 128 MB. |
429 | rate_limit_exceeded, concurrency_limit, overloaded | Over the account's 100 requests a minute, too many calls in flight, or the service is full. Retry after Retry-After. |
502 | upstream_error | The model server failed; retry. |
503 | unavailable | The model is temporarily unavailable. |
504 | timeout | The call ran past its timeout. |
The Python client retries a 429, a 5xx other than 504 and a connection error up to twice (max_retries), then raises these as SGLangError with the status in .status, and a 504 as GenerationTimeout.
Pricing
$0.05 per million input tokens and $0.50 per million thinking tokens; answers are free. See Pricing.