Skip to content
Open weights · Apache-2.04B parametersCalibrated probabilities/v1/systemone compatible

Decisions, not text.
Leo-1 answers typed questions in one forward pass.

Send a state (a message, a ticket, a web page, any JSON) and a set of yes/no, choice or score questions. Leo-1 returns a calibrated probability for every answer. No generation, no parsing, no JSON repair. It speaks the same POST /v1/systemone format as TypeSafe's Jev models, so existing clients work by changing the base URL.

What it is good at

Short conditions

is(text, "angry"), "asking for a refund", "spam", even one-word inputs like "ok" or an emoji. 0.97 accuracy on a blind 211-condition suite.

Routing and classification

Pick a team, an intent or a topic from your own labels, with or without descriptions, and get a probability for each.

Moderation checks

Spam, toxicity, PII and prompt-injection rules asked together; each rule gets its own probability.

Browser agents

Next-action choices for agents: 18 of 21 runs on TypeSafe's own jev-ultrafast suite, the same as Jev, with no false "done".

Try it

Runs Leo-1 live on a GPU. The first request after a quiet period can take about a minute while the GPU starts; after that answers take about a second. Please do not paste personal or confidential data.

Answers

Results appear here.

Raw request and response

API

Base URLhttps://leo.kognare.com
EndpointPOST /v1/systemone · GET /v1/models · GET /v1/usage (your own usage today)
API keyAuthorization: Bearer free
Modelleo-1 (any model value is accepted and answered by Leo-1)
Free trial limits20 requests per minute and 200 per day per IP address; up to 16 questions and 20,000 characters of state per request. Limits are shown in X-RateLimit-* response headers.
Errors401 bad key · 413 body too large · 422 invalid request · 429 rate limit (see Retry-After) · 529 trial busy or GPU starting

Question types

typecriteriaanswer
nouloptional {"true": "...", "false": "..."}noul = probability of yes
choiceobject: option key → description (or null), 2 to 255 optionschoice, probabilities, confidence
scoreordered array of 2 to 10 level descriptionsscore (expected level), probabilities, legend, confidence

Questions never see each other, and question ids are never shown to the model. Instructions can point into JSON state with paths such as `ticket.body`. Probabilities have two decimals and sum to exactly 1.

Examples

curl https://leo.kognare.com/v1/systemone \
  -H "Authorization: Bearer free" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "leo-1",
    "state": {"ticket": {"subject": "Charged twice", "body": "I was billed twice this month. Please fix it today."}},
    "questions": {
      "team":   {"type": "choice", "instructions": "Which team should handle `ticket`?",
                 "criteria": {"billing": "charges, refunds", "technical": "bugs, outages", "other": null}},
      "urgent": {"type": "noul", "instructions": "urgent"},
      "anger":  {"type": "score", "instructions": "How upset is the customer?",
                 "criteria": ["calm", "annoyed", "very angry"]}
    }
  }'

Benchmarks

Identical requests sent to Leo-1 and to TypeSafe's Jev 1.13 (live API, September 2026). Jev was used only for scoring, never for training.

benchmarkLeo-1Jev 1.13
Browser agent tasks, jev-ultrafast (21 runs)18/2118/21
naturalcodz drop-in scenarios35/3736/37
Blind short-condition suite (211 conditions)0.9720.976
Calibration error on short conditions (lower is better)0.0290.075
Held-out classification, calibration error (lower is better)0.0690.167
Tweet topic classification0.8400.790
JevBench standard tier0.9440.986
Held-out classification, mean accuracy0.6500.689
JevBench hard tier (long policies, multi-hop, dates)0.4140.721
World knowledge, MMMLU (15 languages)0.4810.861

Where Leo-1 is behind: hard multi-step reasoning, world knowledge and low-resource languages. Put the facts it needs into the state. Full tables, protocol and raw results are on GitHub.

How it works

Prefill only

A pretrained Qwen3-4B decoder reads the state once; nothing is generated. LoRA adapters (rank 64) specialise it.

Packed questions

All questions share one encoding of the state behind a block-causal mask, so they cannot influence each other.

Pointer readout

A small head scores every option's end marker against the question's decision marker. Markers are reserved embeddings that text cannot forge.

Calibrated

Trained with proper scoring rules on about 180k requests, then temperature-calibrated per question type (dev calibration error 0.007).

Leo-1 is an exact weight merge of two training runs (leo-4b-v5 and its browser-focused continuation), trained on 2 epochs of public human-labelled datasets plus code-generated families on one H100.

GitHub

SuparvaCode/leo: code, data pipeline, training and every benchmark.

github.com/SuparvaCode

Hugging Face

huggingface.co/Suparva: open weights (leo-1.7b).

naturalcodz

SuparvaCode/naturalcodz: one-line is(), pick(), rate(), guard() helpers that run on Leo-1 or Jev.

Notice

  • Free trial for testing and evaluation only. No uptime or accuracy guarantee; limits may change and the trial may close at any time. When many people use it at once, new users are asked to come back later.
  • Privacy. Request and response contents are not stored or logged. Per-IP request counts and token usage are recorded (IP addresses are stored only as salted hashes) to enforce the limits.
  • Do not send personal, confidential or regulated data. Do not use Leo-1 alone for decisions about people's health, finances, legal status or employment, and verify agent actions independently.
  • Independent project. Leo is not affiliated with TypeSafe. It follows TypeSafe's publicly documented /v1/systemone format for compatibility, contains none of its weights and was not trained on its outputs.
  • Weights and code are Apache-2.0; training datasets keep their own licences.