Kodiak-v0.3-1B
An open 1B decision model: seven more kinds of decision than v0.2, better on real-world checks, and a "can't tell" you can trust. Kodiak is built by Cortex Agent LLC. You give it a state (text, a list of texts, or JSON) and typed questions. It returns calibrated choice, score or "can't tell" answers in one forward pass. Use it to automate routine "read this and decide" work: routing, triage, guardrails and checks. Send the cases it isn't sure about to a person or an LLM.
Built on Ettin-encoder-1B (Johns Hopkins, MIT license), with 1.04 billion parameters. Code, docs and the full public build log: https://github.com/grizzlypeaksoftware/kodiak
# pip install "kodiak-s1[infer] @ git+https://github.com/grizzlypeaksoftware/kodiak"
from kodiak_s1.hub import Kodiak
kodiak = Kodiak.from_pretrained("cortex-agent-llc/kodiak-v0.3-1b")
# Guardrail against your own written policy
kodiak.decide(
{"policy": "1. No selling outside the marketplace. 2. All prices must be listed in USD. 3. Never post a personal phone number.",
"message": "Road bike for sale, $250. Call me on 555-201-7788 if you want to see it."},
[{"type": "choice", "id": "rule", "text": "Which rule does the message break?",
"labels": ["rule 1", "rule 2", "rule 3", "no rule"]}],
)
# -> rule: "rule 3" (0.91)
# Is an agent's next step safe to run without asking?
kodiak.decide(
{"task": "Clean up old log files in /var/log/myapp to free some space.", "next_step": "rm -rf /var/log"},
[{"type": "choice", "id": "safe", "text": "Is it safe for the agent to run this next step without asking?",
"labels": ["yes, safe", "ask the user first", "no, it shouldn't run"]}],
)
# -> safe: "no, it shouldn't run" (0.81)
For the most accurate and best-calibrated answers, use accuracy mode. It averages three v0.3 models and costs about 3× the compute.
What's new in v0.3
Seven new kinds of decision, each taught with checked synthetic data (an open-weight writer, gpt-oss-120b, builds each example toward a fixed answer; a blind checker, DeepSeek-V3.2, must agree):
| Kind | Example question |
|---|---|
| Pairwise judge | Which of these two answers is better? |
| Sarcasm | Is the writer being sarcastic? |
| Policy violation | Which rule of this written policy does the message break? |
| Long-answer hallucination | Is everything in this answer supported by the sources? |
| Stance | What is the author's stance on a named target? |
| Refund eligibility | Is the customer eligible under this refund policy? |
| Agent step safety | Is it safe to run this next step without asking? (keyword-level today; see Known limits) |
Measured on real data the model never trained on (chance-corrected skill, 0 = random guessing; Kodiak figures are three-run means):
| Real-data test | Kodiak-v0.2-1B | Kodiak-v0.3-1B | v0.3 accuracy mode |
|---|---|---|---|
| RAGBench: is a long answer supported by its sources? (CC BY 4.0) | +0.27 | +0.41 | +0.44 |
| MT-Bench: which answer did expert humans prefer? (CC BY 4.0) | +0.06 | +0.29 | +0.33 |
| SemEval-2016: stance of a tweet toward a target (MIT) | +0.30 | +0.37 | +0.38 |
| Average | 0.21 | 0.36 | 0.38 |
How it compares (frozen eval set v0.2, choice questions)
| Kodiak-v0.2-1B | Kodiak-v0.3-1B | v0.3 accuracy mode | Qwen3-8B (LLM) | |
|---|---|---|---|---|
| Never-seen tasks, forced accuracy | 0.689 ± 0.008 | 0.687 ± 0.010 | 0.696 | 0.688 |
| Familiar tasks | 0.877 | 0.879 | 0.887 | 0.710 |
| Ranks its own mistakes last (never-seen) | 55.6% | 54.5% | 56.1% | 14.7% |
| Calibration error (never-seen; lower is better)¹ | 0.085 | 0.076 | 0.065 | 0.293 |
| When it says "can't tell", it's right | 0.88 | 0.93 | 0.925 | – |
| Latency (GPU, one request) | ~38 ms | ~38 ms | ~3× | ~1,500 ms |
Kodiak-v0.3-1B figures with ± are the mean of three training runs. This checkpoint is the seed-1 run, chosen by validation loss and never by the eval set. Its own scores are 0.675 never-seen forced, 0.873 familiar, 0.074 never-seen calibration error and 0.961 "can't tell" precision. "Never-seen" means tasks and label sets the model was never trained on (zero-shot). The eval set and the LLM baseline setup are in the repository. v0.3 does not improve never-seen accuracy over v0.2; its gains are the new decision kinds, real-data checks and an honest "can't tell".
Ranking is what a confidence threshold relies on: sort answers by confidence and see how much of the gap between a random order and the perfect order (all mistakes last) the model closes. It doesn't change under any recalibration. ¹ The calibration comparison is raw. Qwen3-8B is mostly overconfident by a constant amount, so a fitted recalibration (isotonic, fit on the other never-seen tasks) brings it from 0.293 to 0.178; the same treatment gives Kodiak-v0.3-1B 0.050-0.072.
Known limits (read before using)
- Option wording still matters. The same question with reworded options gets the same answer about two thirds of the time on never-seen tasks. Keep options short and distinct, and test a few wordings on your own data. On a 100-item wording-trap test (a message that repeats a wrong option's words inside a condition or negation), v0.3 is right 90% of the time (v0.2: 87%).
- Real-world checks are improved, not solved. Long-answer hallucination detection, pairwise judging and stance are well above chance on real data (table above) but far from perfect. Use "can't tell" and a confidence threshold, and review the rest.
- Ranking is slightly lower than v0.2 for the single model (54.5% vs 55.6%, within run-to-run noise); accuracy mode is 56.1%.
- Agent step safety is mostly a keyword check today. It keys on the command itself (e.g. rm -rf → never; delete or send_email → ask first) more than on the task, and can call an unsafe step safe with high confidence (a push to main while tests are failing → "safe", 0.97). Do not use it as a safety control. Refund eligibility, policy violation and sarcasm have no real-data or contrastive test yet; treat them as unverified against similar shortcuts. (Reported by a Hugging Face user; we're adding contrastive tests, where the same input appears under different answers.)
- Implied violations are harder than stated ones. A message that breaks a rule only by implication (e.g. "text me at the number in my profile, no fees that way" against "no selling outside the marketplace") can be missed or pinned on the wrong rule, with low confidence. Treat low-confidence guardrail answers as "send to a person".
- Arithmetic (e.g. does an invoice total match its lines?) is near chance and wasn't trained.
- Long inputs. The position limit is about 8,000 tokens (the state, every question and every option together). Longer requests are refused rather than truncated.
- Ratings ("how urgent is this?") are rough. Prefer choice questions.
- Speed. About 38 ms per request on a GPU. On a CPU, expect a few hundred milliseconds.
- Validate it on your own data. Don't use it for decisions about people without human review. Teach it your own decisions with the fine-tuning kit in the repository (a CSV in, a before/after report out).
License and citation
Apache-2.0 (weights and code). Base model: Ettin-encoder-1B (MIT). Training-data licenses are listed in data/LICENSES.md in the repository.
All synthetic training data was written and checked by open-weight models; no closed-model outputs and no evaluation data are in training.
Cortex Agent LLC / Grizzly Peak Software, 2026.
Model tree for cortex-agent-llc/kodiak-v0.3-1b
Base model
jhu-clsp/ettin-encoder-1b