Llama-modes: load one GGUF once, then Chat, Boolean, Choice and Scale

I wanted to explore a simple question:
can one ordinary local LLM, loaded once, be used as more than a text generator without introducing a second classifier or swapping models?
That experiment became llama-modes, a fork of llama.cpp.
The core idea is that you load one GGUF model once, keep the same weights in memory, and use that same loaded model for several different inference modes.
Normal llama-server chat remains available, but the same model can also perform:

  • BOOLEAN — directly score Yes / No
  • CHOICE — score arbitrary supplied candidates, including multi-token labels
  • SCALE — evaluate an ordinal or interval scale and return a discrete distribution over the supplied scale points

There is no second classifier, no model swapping, and no special fine-tune required.
Candidate scoring itself is obviously not new. People already use next-token logits or sequence likelihoods for classification, ranking, and multiple-choice evaluation.
What I wanted to explore was turning that idea into a reusable llama.cpp runtime primitive rather than rebuilding the logic separately in every application.
In other words: same model, same weights in memory, different inference primitive.
For single-token choices, llama-modes can work directly from the model output at the prepared evaluation state.
For multi-token choices, it uses teacher-forced sequence log-likelihood:
log P(candidate | prompt) rather than pretending that an entire multi-token candidate has a single logit.
The SCALE mode is probably the easiest part to demonstrate visually.
Instead of asking a model to generate something like:
9/10 you can supply an ordered 0–10 scale and get back the entire discrete distribution across those supplied points.
For ordinal scales, llama-modes returns things such as mode, median and quantiles.
For interval scales, where the caller explicitly asserts that numeric distances have meaning, it can also return expected value and weighted spread.
The model is not generating those summary statistics. They are derived from the returned distribution.
I originally became interested in this direction after the recent discussion around Jev and decision-first inference, but llama-modes is not a Jev reimplementation and does not claim Jev-style calibration.
It takes a different route: exposing structured scoring directly from ordinary local GGUF language models.
So far I have tested it with GPT-OSS 20B MXFP4 and Qwen3.8 Ridge on Windows with NVIDIA CUDA.
The repository now includes a Windows CUDA release, a local React demo, Python / PowerShell / curl examples, a cookbook, API documentation, and a reproducible Direct-vs-Chat benchmark harness.
Repo:

Release:

One practical note if you try the demo:
the first request after loading the model can be noticeably slower because of warm-up.
Run 2–3 requests before judging interactive latency. The Direct-vs-Chat screen in the React UI is meant as an interactive demonstration, not as the benchmark itself; the repository contains a separate harness for reproducible measurements.
Also, the returned candidate/scale weights should not be interpreted as calibrated confidence. They are relative to the supplied alternatives and their representations, and label/tokenization choices can matter.
One thing I’d especially like feedback on is whether this kind of direct structured inference is useful in real applications, and which model families behave well or badly with it.

The current roadmap item is shared-context multi-question evaluation: one context evaluation, then multiple Boolean / Choice / Scale questions over it.

I tried a small test, and the behavior changed quite a lot depending on the setup even within the same model, so I couldn’t really say much about family-level differences:


For the practical-use question: yes, I can see a fairly natural use case for this, especially when the same GGUF is already resident for normal chat.

The pattern that makes the most sense to me is not “Direct instead of Chat everywhere”, but something more like:

closed / finite decision
    -> Direct

needs reasoning, open-ended output, or Direct looks unstable
    -> normal Chat / reasoning

So things like routing, relevance decisions, small taxonomies, rubric checks, triage, or other places where the application ultimately needs one of a small set of answers seem like a good fit. If the same model is already loaded for chat, being able to use it for those decisions without loading a second classifier is operationally attractive.

There are already adjacent examples of this general pattern. vLLM’s Generative Scoring scores specified label tokens with an ordinary generative model, and models such as Qwen3-Reranker also turn a small yes/no readout into a ranking score. Older work such as monoT5 is another example of using target-token logits as a practical score.

Where I became less confident was the “which model families work well?” part. My small test suggested that before comparing families, I would probably control a few configuration variables inside each model first, because those alone were enough to change the result quite a lot.

A compact compatibility report might be more useful than just the family name:

model / GGUF:
quantization:
runtime + backend:
chat template:
answer/readout boundary:
candidate text:
candidate token IDs / token counts:
scoring rule:
task:
Direct result:
Chat result:
label-swap result:

Family-level patterns may well emerge after enough results are collected, but this would make those comparisons easier to interpret.

The cheapest smoke test I would personally do for a new model/configuration is:

  1. verify the intended assistant/final-answer boundary;
  2. inspect the candidate tokenization;
  3. try one equivalent label remapping, e.g. Yes/No → A/B;
  4. compare Direct and normal Chat on a handful of labeled examples;
  5. only then run a larger benchmark.

Those checks are cheap, and in my little test they were surprisingly informative.

Small Qwen3-4B sanity check

I tested Qwen/Qwen3-4B (1cfa9a7208912126459214e8b04321603b3df60c) on an NVIDIA L4 using Transformers.

This was not a test of the llama-modes C++ runtime. I was only trying to probe the model/template side of the question: if I take one model and perform finite-candidate scoring at different places or with different label representations, how stable is the result?

The set was intentionally small: 24 synthetic true/false propositions, split into 12 easy and 12 somewhat more reasoning-like items.

At the normal non-thinking/final-answer boundary:

  • Direct scoring: 23/24
  • non-thinking generated Chat: 23/24
  • Direct and Chat made the same binary decision on all 24 items

So on that small closed-decision set, Direct scoring did not create a different decision pattern from ordinary non-thinking generation.

The much more interesting result was the readout position.

Using the same 24 propositions but scoring Yes/No at the raw thinking-start position instead of the final-answer position:

  • final-answer boundary: 23/24
  • thinking-start boundary: 12/24

I would not interpret 12/24 as “Qwen3 is bad at Direct scoring”. That position is deliberately the wrong semantic place to ask for the final binary answer.

Rather, to me it says that the prepared evaluation state is part of the compatibility contract.

That seems particularly relevant for reasoning models. Qwen3’s own thinking-mode documentation says that in thinking mode it generates a <think>...</think> block and then the final response. In other words, “the state where reasoning is about to begin” and “the state where the model is expected to produce the final answer” are meaningfully different states.

Qwen3-Reranker is also an interesting nearby example: its official usage places a completed empty <think>...</think> block immediately before the yes/no scoring position. I do not think that exact template should be generalized to unrelated models, but it is another reason I would record the readout boundary, not just the model family.

I also reran the 12 harder items with actual Qwen3 thinking allowed to finish. The corrected run required a real </think> followed by a final Yes or No.

Results:

  • Direct at the valid final-answer boundary: 12/12
  • completed thinking: 12/12
  • Direct vs completed-thinking decisions: 12/12 agreement
  • median thinking generation: 316 tokens

Again, I would keep the interpretation narrow: these happened to be problems that Direct already solved. This does not show that Direct generally replaces reasoning.

It does show one real class of cases where the same final decision was available directly without generating hundreds of reasoning tokens.

The open question I did not answer is probably the more interesting one: which tasks are wrong under Direct but become correct after reasoning? That would define the useful fallback boundary much better.

Labels and score interpretation

I also tried equivalent output representations at the valid answer boundary.

All of the tested labels tokenized cleanly, but changing the representation still changed a few decisions:

  • Yes / No
  • A = Yes, B = No
  • A = No, B = Yes
  • changing the listing order
  • True / False

Depending on the variant, 1–2 of 24 decisions flipped.

Some of those flips happened even when the original candidate-relative margin was very large.

So I think your warning that labels/tokenization affect the scores is important. My test suggests that a tokenization check is useful but not sufficient by itself: a label can be perfectly valid at the token level and still change the semantic readout.

A very cheap robustness check would therefore be one equivalent label/mapping perturbation.

This also made me cautious about treating something like

P(Yes | Yes or No) = 0.99

as if it automatically meant

P(the answer is actually correct) = 0.99

Those are different statements.

Your current wording that the returned weights are relative candidate weights rather than calibrated confidence seems like the right distinction to preserve.

If an application wants to use a margin or entropy as a Direct → Chat fallback signal, I would probably validate that on held-out application data rather than giving the raw weight a universal interpretation. The general idea is related to selective prediction / classifier cascades: accept cheap predictions where a locally validated signal is reliable, and defer the rest.

One small wrinkle from my test is that margin alone may not be enough, because I saw representation changes flip some apparently high-margin cases. Representation stability could potentially be another useful diagnostic, but that would also need empirical validation.

Scoring primitive vs scoring policy

Another separation that seems useful to me is:

raw model scoring
        |
        +-- candidate-set normalization
        +-- length normalization
        +-- prior / MI correction
        +-- calibration
        +-- application threshold / fallback

I would treat those as separate layers rather than looking for one universally correct transformation.

For example, vLLM’s current Generative Scoring API explicitly distinguishes between:

  • softmax over only the supplied label tokens; and
  • the labels’ probabilities in the full vocabulary distribution.

Likewise, lm-evaluation-harness keeps different multiple-choice scoring policies as different metrics, including raw accuracy, length-normalized accuracy, and a mutual-information-style variant.

That seems relevant to multi-token CHOICE as well.

Using teacher-forced sequence log-likelihood SUM is a perfectly understandable raw primitive, but SUM naturally carries sequence-length and surface-form effects. A mean/length-normalized score answers a slightly different question, and prior corrections answer another one again.

I would therefore be inclined to keep the primitive semantics very explicit, expose enough raw information to reproduce the decision, and let benchmark/application layers compare alternative policies where needed.

I would not bake in a claim that one normalization is universally better. Evaluation work such as OLMES is a useful reminder that prompt formulation, probability normalization, examples, and task formulation can all materially affect measured results.

This is also why I think a family comparison needs a little metadata. Otherwise a difference caused by prompt/readout/scoring setup can easily look like a family difference.

One thought for the v0.5 shared-context direction

The shared-context multi-question roadmap also looks like a natural next step to me.

If many Boolean / Choice / Scale questions share the same large input, avoiding repeated context evaluation is exactly the kind of optimization I would want.

The main thing I would preserve while optimizing it is a deliberately simple correctness oracle:

optimized shared-context scoring
            vs
fresh independent / full-prompt scoring

and compare the actual score vectors, not only the selected winner.

This seems especially worthwhile because the project has already encountered a useful correctness lesson here. The v0.3.1 re-prefill change replaced restored candidate state with prompt recomputation after GPT-OSS CUDA/SWA tests produced candidate-order-dependent scores; re-prefill recovered agreement with the independent oracle.

I would see that less as a problem with the direction and more as a good precedent for v0.5:

optimize against a simple known-correct path, and keep parity with that path as a test invariant.

For shared-context evaluation in particular, I think the semantic requirement is slightly stronger than “these questions share a token prefix”. Each branch still needs to arrive at the same evaluation state it would have reached under an independent evaluation.

So a small matrix such as

fresh vs shared
original question order vs reversed
batch size A vs B
repeat run

could catch a lot before doing larger performance measurements.

Overall, I think the project has a useful niche.

The distinction I would keep strongest is:

Direct is a structured readout from a particular model state, not shortened Chat.

When the desired decision is already represented at that state, directly scoring the supplied alternatives can be a very convenient path. When the model still needs to transform the state through reasoning, scoring before that reasoning is simply asking a different question.

That suggests a practical default of Direct for closed decisions, Chat/reasoning as the fallback, while treating model/template/readout/scoring configuration as part of compatibility.

And if people start reporting which models work well, I think even a tiny standardized report like this would make the answers much more reusable:

model / GGUF:
quant:
backend:
chat template / answer boundary:
candidate labels + token IDs:
task:
Direct:
Chat:
label-swap:

That might eventually make the family-level picture much clearer without requiring everyone to run a large benchmark first.

Thank you !!— this is exactly the kind of feedback I was hoping for

Your distinction that Direct is a structured readout from a particular model state, not shortened Chat is especially useful. I think that captures the semantics better than any shorthand I’ve used so far.

The Qwen3 boundary experiment is also very informative. The difference between the final-answer boundary and the thinking-start boundary makes the point very clearly: the prepared evaluation state is part of the compatibility contract, especially for reasoning models.

I also agree that tokenization alone is not enough as a robustness check. Your label-remapping results are a good reminder that representation sensitivity can remain even when all candidate labels tokenize cleanly.

I’m going to make a few concrete changes based on your test:

  1. Add a small standardized compatibility report format recording model/GGUF, quantization, backend, chat template, answer/readout boundary, candidate labels and token IDs, scoring rule, task, Direct result, Chat result, and label-remapping result.

  2. Add a cheap compatibility smoke-test protocol for new models/configurations:

    • verify the intended final-answer boundary;
    • inspect candidate tokenization;
    • test one equivalent label remapping / swapped mapping;
    • compare Direct and normal Chat on a small labeled set;
    • only then move to larger benchmarks.
  3. Strengthen the documentation around semantics:

    • Direct is a readout from a specific model state;
    • candidate-relative weights are not calibrated correctness probabilities;
    • representation sensitivity is broader than tokenization alone.
  4. Keep scoring primitive and scoring policy explicitly separate. llama-modes will continue to expose reproducible raw scoring information such as sequence SUM and MEAN rather than claiming that one normalization, calibration, or fallback threshold is universally correct.

  5. Add an evaluation track specifically for Direct-vs-reasoning disagreement. Your open question is probably the most interesting one: finding cases that are wrong under Direct but become correct after actual reasoning would help define where a Direct → Chat fallback is genuinely useful.

  6. For v0.5 shared-context evaluation, I’m going to treat parity against fresh independent evaluation as a correctness invariant, not merely compare selected winners. At minimum I want tests across:

    • fresh vs shared-context scoring;
    • original vs reversed question order;
    • different batch sizes;
    • repeated runs;
    • full score-vector parity within the expected numeric tolerance.

The v0.3.1 re-prefill issue was a very useful lesson here, so I agree that the simple independent path should remain the oracle while the optimized path evolves.

Your Direct → Chat framing also feels like a much better practical story than treating them as competitors:

closed / finite decision
    -> Direct

needs reasoning, open-ended output, or Direct looks unstable
    -> Chat / reasoning

I don’t want to hard-code that as a universal policy, but it is a useful application pattern to validate.

Thanks again for actually testing the idea rather than only commenting on the concept. Several of these points will go directly into the compatibility and v0.5 work.