I tried a small test, and the behavior changed quite a lot depending on the setup even within the same model, so I couldn’t really say much about family-level differences:
For the practical-use question: yes, I can see a fairly natural use case for this, especially when the same GGUF is already resident for normal chat.
The pattern that makes the most sense to me is not “Direct instead of Chat everywhere”, but something more like:
closed / finite decision
-> Direct
needs reasoning, open-ended output, or Direct looks unstable
-> normal Chat / reasoning
So things like routing, relevance decisions, small taxonomies, rubric checks, triage, or other places where the application ultimately needs one of a small set of answers seem like a good fit. If the same model is already loaded for chat, being able to use it for those decisions without loading a second classifier is operationally attractive.
There are already adjacent examples of this general pattern. vLLM’s Generative Scoring scores specified label tokens with an ordinary generative model, and models such as Qwen3-Reranker also turn a small yes/no readout into a ranking score. Older work such as monoT5 is another example of using target-token logits as a practical score.
Where I became less confident was the “which model families work well?” part. My small test suggested that before comparing families, I would probably control a few configuration variables inside each model first, because those alone were enough to change the result quite a lot.
A compact compatibility report might be more useful than just the family name:
model / GGUF:
quantization:
runtime + backend:
chat template:
answer/readout boundary:
candidate text:
candidate token IDs / token counts:
scoring rule:
task:
Direct result:
Chat result:
label-swap result:
Family-level patterns may well emerge after enough results are collected, but this would make those comparisons easier to interpret.
The cheapest smoke test I would personally do for a new model/configuration is:
- verify the intended assistant/final-answer boundary;
- inspect the candidate tokenization;
- try one equivalent label remapping, e.g.
Yes/No → A/B;
- compare Direct and normal Chat on a handful of labeled examples;
- only then run a larger benchmark.
Those checks are cheap, and in my little test they were surprisingly informative.
Small Qwen3-4B sanity check
I tested Qwen/Qwen3-4B (1cfa9a7208912126459214e8b04321603b3df60c) on an NVIDIA L4 using Transformers.
This was not a test of the llama-modes C++ runtime. I was only trying to probe the model/template side of the question: if I take one model and perform finite-candidate scoring at different places or with different label representations, how stable is the result?
The set was intentionally small: 24 synthetic true/false propositions, split into 12 easy and 12 somewhat more reasoning-like items.
At the normal non-thinking/final-answer boundary:
- Direct scoring: 23/24
- non-thinking generated Chat: 23/24
- Direct and Chat made the same binary decision on all 24 items
So on that small closed-decision set, Direct scoring did not create a different decision pattern from ordinary non-thinking generation.
The much more interesting result was the readout position.
Using the same 24 propositions but scoring Yes/No at the raw thinking-start position instead of the final-answer position:
- final-answer boundary: 23/24
- thinking-start boundary: 12/24
I would not interpret 12/24 as “Qwen3 is bad at Direct scoring”. That position is deliberately the wrong semantic place to ask for the final binary answer.
Rather, to me it says that the prepared evaluation state is part of the compatibility contract.
That seems particularly relevant for reasoning models. Qwen3’s own thinking-mode documentation says that in thinking mode it generates a <think>...</think> block and then the final response. In other words, “the state where reasoning is about to begin” and “the state where the model is expected to produce the final answer” are meaningfully different states.
Qwen3-Reranker is also an interesting nearby example: its official usage places a completed empty <think>...</think> block immediately before the yes/no scoring position. I do not think that exact template should be generalized to unrelated models, but it is another reason I would record the readout boundary, not just the model family.
I also reran the 12 harder items with actual Qwen3 thinking allowed to finish. The corrected run required a real </think> followed by a final Yes or No.
Results:
- Direct at the valid final-answer boundary: 12/12
- completed thinking: 12/12
- Direct vs completed-thinking decisions: 12/12 agreement
- median thinking generation: 316 tokens
Again, I would keep the interpretation narrow: these happened to be problems that Direct already solved. This does not show that Direct generally replaces reasoning.
It does show one real class of cases where the same final decision was available directly without generating hundreds of reasoning tokens.
The open question I did not answer is probably the more interesting one: which tasks are wrong under Direct but become correct after reasoning? That would define the useful fallback boundary much better.
Labels and score interpretation
I also tried equivalent output representations at the valid answer boundary.
All of the tested labels tokenized cleanly, but changing the representation still changed a few decisions:
Yes / No
A = Yes, B = No
A = No, B = Yes
- changing the listing order
True / False
Depending on the variant, 1–2 of 24 decisions flipped.
Some of those flips happened even when the original candidate-relative margin was very large.
So I think your warning that labels/tokenization affect the scores is important. My test suggests that a tokenization check is useful but not sufficient by itself: a label can be perfectly valid at the token level and still change the semantic readout.
A very cheap robustness check would therefore be one equivalent label/mapping perturbation.
This also made me cautious about treating something like
P(Yes | Yes or No) = 0.99
as if it automatically meant
P(the answer is actually correct) = 0.99
Those are different statements.
Your current wording that the returned weights are relative candidate weights rather than calibrated confidence seems like the right distinction to preserve.
If an application wants to use a margin or entropy as a Direct → Chat fallback signal, I would probably validate that on held-out application data rather than giving the raw weight a universal interpretation. The general idea is related to selective prediction / classifier cascades: accept cheap predictions where a locally validated signal is reliable, and defer the rest.
One small wrinkle from my test is that margin alone may not be enough, because I saw representation changes flip some apparently high-margin cases. Representation stability could potentially be another useful diagnostic, but that would also need empirical validation.
Scoring primitive vs scoring policy
Another separation that seems useful to me is:
raw model scoring
|
+-- candidate-set normalization
+-- length normalization
+-- prior / MI correction
+-- calibration
+-- application threshold / fallback
I would treat those as separate layers rather than looking for one universally correct transformation.
For example, vLLM’s current Generative Scoring API explicitly distinguishes between:
- softmax over only the supplied label tokens; and
- the labels’ probabilities in the full vocabulary distribution.
Likewise, lm-evaluation-harness keeps different multiple-choice scoring policies as different metrics, including raw accuracy, length-normalized accuracy, and a mutual-information-style variant.
That seems relevant to multi-token CHOICE as well.
Using teacher-forced sequence log-likelihood SUM is a perfectly understandable raw primitive, but SUM naturally carries sequence-length and surface-form effects. A mean/length-normalized score answers a slightly different question, and prior corrections answer another one again.
I would therefore be inclined to keep the primitive semantics very explicit, expose enough raw information to reproduce the decision, and let benchmark/application layers compare alternative policies where needed.
I would not bake in a claim that one normalization is universally better. Evaluation work such as OLMES is a useful reminder that prompt formulation, probability normalization, examples, and task formulation can all materially affect measured results.
This is also why I think a family comparison needs a little metadata. Otherwise a difference caused by prompt/readout/scoring setup can easily look like a family difference.
One thought for the v0.5 shared-context direction
The shared-context multi-question roadmap also looks like a natural next step to me.
If many Boolean / Choice / Scale questions share the same large input, avoiding repeated context evaluation is exactly the kind of optimization I would want.
The main thing I would preserve while optimizing it is a deliberately simple correctness oracle:
optimized shared-context scoring
vs
fresh independent / full-prompt scoring
and compare the actual score vectors, not only the selected winner.
This seems especially worthwhile because the project has already encountered a useful correctness lesson here. The v0.3.1 re-prefill change replaced restored candidate state with prompt recomputation after GPT-OSS CUDA/SWA tests produced candidate-order-dependent scores; re-prefill recovered agreement with the independent oracle.
I would see that less as a problem with the direction and more as a good precedent for v0.5:
optimize against a simple known-correct path, and keep parity with that path as a test invariant.
For shared-context evaluation in particular, I think the semantic requirement is slightly stronger than “these questions share a token prefix”. Each branch still needs to arrive at the same evaluation state it would have reached under an independent evaluation.
So a small matrix such as
fresh vs shared
original question order vs reversed
batch size A vs B
repeat run
could catch a lot before doing larger performance measurements.
Overall, I think the project has a useful niche.
The distinction I would keep strongest is:
Direct is a structured readout from a particular model state, not shortened Chat.
When the desired decision is already represented at that state, directly scoring the supplied alternatives can be a very convenient path. When the model still needs to transform the state through reasoning, scoring before that reasoning is simply asking a different question.
That suggests a practical default of Direct for closed decisions, Chat/reasoning as the fallback, while treating model/template/readout/scoring configuration as part of compatibility.
And if people start reporting which models work well, I think even a tiny standardized report like this would make the answers much more reusable:
model / GGUF:
quant:
backend:
chat template / answer boundary:
candidate labels + token IDs:
task:
Direct:
Chat:
label-swap:
That might eventually make the family-level picture much clearer without requiring everyone to run a large benchmark first.