Scaling Beatrix V3: a 376M byte-level model and the training

Scaling Beatrix: a 376M byte-level model with splat memory in every block, detachable arms, and the first slider experiments

Disclosure: this post and the article it summarizes were written by Claude (Anthropic) working with AbstractPhil, who directs the research and made every decision in it; the numbers come from the training record, and the record prints its own retractions. Feedback from people and from AI tools is equally welcome. If you run the article through a model and it finds a hole, a missing control or a better explanation, please post what it said.

What it is. The third installment on Beatrix, a byte-level language model (vocabulary of 256, context 4,096 bytes) whose every attention block is splat memory: a signed, softmax-free addressed memory read by a closed-form rule over a chunked causal scan. Article: https://huggingface.co/blog/AbstractPhil/beatrix-ft3 (about 15,000 words; the technical companions ship with the model). Model: AbstractPhil/mini-beatrix-3 · Hugging Face . Chat with it: Beatrix — AlephLLM Chat - a Hugging Face Space by AbstractPhil .

The trunk. 376 million parameters, 32 blocks, trained on 64.4 billion bytes on two RTX 5090 cards in 295 hours: a web pretraining phase, nine curriculum stages of synthetic text that each teach one skill, and two anneals at a flat learning rate. It finished at 0.9861 bits per byte on held-out web text with no loss spike and the gradient clip never reached. A depth ladder priced the decision: held-out loss fell at every rung from 16 to 32 blocks (1.561 to 1.476 bits per byte at a matched step).

The arms. Detachable 13.7M-parameter adapters after every block, trained with the trunk frozen and penalized for changing predictions outside their job. The arms trained beside the trunk turned out empty: the trunk took each stage’s text first. Refitted on the finished weights, the eight stage arms read, for example, .5538 to .0492 bits per byte on their own stage at a web-text cost of +.0019, two seeds, and detach bit for bit. Solo arms switched on together fail together; a staged group with a roster rule gives whole, separable members; routed dispatch over them reads worse than the plain stack.

What the size did not buy. Recall at a distance. A planted code comes back at .753 of its digits at once and .031 about 3,500 bytes later, where a softmax control model held .588. The additive memory never erases; the next prototype is a hybrid with standard attention in the model’s own weakest-rank blocks.

As a text encoder, with nothing trained. Its middle blocks agree with T5 at .62 to .65 (an untrained copy: .23 to .25) and bind attributes at .89. One fixed direction at the last block carries 94 to 99 percent of each byte’s energy and acts as a per-byte temperature (removing it costs +43.5 bits per byte); its size fluctuates with one self-similar exponent, .62 to .67, against .48 to .51 for shuffled controls.

Sliders. Mood is a dial in the conditioning of two image models. On Sana a learned direction moves a mood judge +0.795 per unit against 14% of that for a random direction; on Anima the same push is dead before the text adapter and alive after it, and the adapter reads every caption twice, token ids as questions and Qwen3 states as answers, with the answers carrying the mood in pictures. Beatrix’s own sliders missed their registered bars five times; the picture test of her relay through the nine arms is built, public ( GitHub - AbstractEyes/alephllm-diffusion-experiments · GitHub ) and ahead.

Where feedback would help most

  1. The hybrid: which blocks should go to standard attention, and what guard (query-key normalization, decay) keeps softmax stable under a recipe tuned for splat memory.
  2. Memory that can forget: overwrite (delta-rule) writes, whitened writes or a direct lookback path. Experience from linear-attention and state-space work is very welcome.
  3. The one large direction at the last block: has anyone measured the same per-byte temperature in other architectures, and how do you bound it without flattening it?
  4. The slider and relay design on Anima’s adapter, and better ways to score a conditioning signal in pictures than a CLIP mood judge.
  5. The method itself: pre-registered bars, two seeds, controls that can fail, retractions printed. Tell us where it is not enough.
  6. The writing: what is unclear, what is too long, what you would cut.

Everything, including what failed, is in the article and the companions; corrections and disagreements are the point of posting it.

For now, here’s what I found from a quick test of the final model:


Small final-checkpoint probe: later values are often preferred, but earlier records still interfere

I ran a small inference-only probe on the published mini-beatrix-3 final checkpoint (step 245,674; arms off), using six synthetic English key–five-digit-code fixtures. The aim was to add a few observations to your forget/rewrite question, and perhaps help interpret the planned hybrid comparisons. This is not a reproduction of your 64-trial battery, a retraining experiment, or an inspection of Splat’s internal writes.

Alongside distance-related degradation, which you already discuss in the Beatrix V3 report, one pattern stood out: an earlier competing record could hurt retrieval even when the later value was preferred.

  • Adding an earlier, conflicting same-key record increased the later correct value’s summed five-byte teacher-forced negative log-likelihood (NLL) by 4.56 bits on average (6/6 positive). Swapping which code appeared earlier produced a 4.30-bit penalty (also 6/6 positive).
  • Yet the later value received a better candidate score in 6/6 and 5/6 fixtures in those two orders. Exact five-byte greedy generation succeeded less often: 4/6 and 2/6.
  • A different key’s numerical code record increased target NLL by 7.47 bits relative to neutral prose. With the distractor’s key and digits unchanged, replacing its field label note with code added another 2.00 bits.

These are paired changes in teacher-forced NLL for a five-byte answer, not held-out bits per byte, percentage-point changes, or estimates of information stored inside Splat. They come from six fixed synthetic fixtures under one checkpoint—not six training seeds or a representative benchmark.

For the forget/rewrite question, the distinction seems useful: a preference for the later value does not show that the earlier value was erased; interference does not show that updating is impossible. Accumulated writes, address formation, readout competition and prompt framing are still different candidate explanations. Output likelihoods alone cannot tell us which part of Splat is responsible.

A practical default for interpreting the next design

Three questions seem worth keeping separate:

  1. Retention: does an isolated fact become harder to recover as its byte gap grows? Persistence, normalized accumulation, position and readout could each matter.
  2. Editing/interference: does a second structured record impose an extra penalty beyond matched filler? Selective erase/write, address separation and better readout are distinct responses.
  3. Selection versus generation: can the model favor the right candidate but still fail to generate all five digits? Candidate scores and unconstrained continuation measure different things.

Global decay could ease stale interference while harming useful distant facts. Gated DeltaNet and Gated DeltaNet-2 provide useful design precedents for retention and selective editing, not evidence that their update rules are drop-in fixes for Splat.

For the planned hybrid, the default comparison I would favor is keyed recall (including interference), held-out bpb, stability, and inference cost, kept as separate outcomes. With the same number of softmax blocks, rank-guided placement versus even spacing seems especially informative; a boundary-focused variant is optional. Low representation rank may suggest where to look, but does not establish that a layer is disposable or that replacing it improves recall. The systematic hybrid linear-attention study likewise reports that language modeling and associative recall can respond differently to full-attention mixing. Its preferred ratios are not Beatrix-specific prescriptions.

If recall improves without unacceptable bpb or cost, that supports the hybrid for the memory objective. A bpb-only gain is still useful, but answers another question. If softmax training becomes unstable, check Q/K and logit scales, norm growth, clipping and precision before trying Query-Key Normalization or a broader optimizer change.

One methodological caution: the first factor probe treated a literal ----- substitution as its comparison; the later control probe actually removed absent records. The first reports an NLL difference; the latter reports absolute NLL and within-probe contrasts. Those figures should not be pooled. The rest of this post gives the fixtures, limits, and conditional design options for anyone who wants to examine them—not an additional benchmark request.

1. What I ran, what the scores mean, and the distance result

Artifact and scoring contract

The completed tests used AbstractPhil/mini-beatrix-3, checkpoint revision:

0ae0e924965b657525bd49199cd3664a5cf37f14

The run recorded 32 blocks, a 256-byte vocabulary, a 4,096-byte window and arms off. It used a Colab L4, FP32 construction and BF16 autocast (PyTorch 2.11.0+cu130, Transformers 4.57.6). A fixed-input re-score matched; training-runtime parity was not tested.

The two main completed experiments were:

Probe Purpose Conditions
v3 Distance, absolute placement, initial distractor/update conditions 72
v4 Better record-format controls; true absence and reversed update order 102

The two probes reuse the same six synthetic fixture families; these are not held-out documents or fresh independent samples across runs.

For candidate c1...c5, I scored the correct byte sequence with teacher forcing:

NLL_bits(code | prefix)
  = -sum(i=1..5) log2 P(c_i | prefix, c_1, ..., c_(i-1))

Lower NLL means a better score for the supplied candidate. Each earlier correct byte is provided before predicting the next one. This can look better than free five-byte generation, which was also measured. Scoring both candidates against one prefix reveals a preference in the output distribution, not whether either association was physically retained.

In v3, saving_bits means:

NLL(correct code | dashed-source version)
  - NLL(correct code | planted-code version)

This is the within-pair benefit of the planted code; positive means the planted source improved its score. v4 instead reports absolute NLL_bits or differences between matched absolute scores. Neither is directly comparable to the article’s recall fractions or held-out bpb, and the v3/v4 numbers must not be subtracted from one another.

Distance with query position fixed

The query was near byte position 3,800, while the planted fact moved:

Source-to-query gap Planted-vs-dash saving, mean Exact greedy 5 digits
48 bytes +31.325 bits 3/6
1,200 bytes +24.093 bits 2/6
3,200 bytes +8.454 bits 0/6

All six fixtures lost planted-source advantage between 48 and 3,200 bytes. This supports distance-associated degradation under these prompts, not the magnitude of the published recall deficit and not a matched softmax comparison.

There is a confound: fixing the query while increasing the gap also moves the source’s absolute position. To probe that, I held the gap at 1,200 bytes and moved the layout:

Query byte position Saving, mean Exact greedy 5 digits
1,700 +19.051 bits 1/6
2,700 +25.140 bits 2/6
3,800 +24.093 bits 2/6

At a fixed gap, the 3,800-position saving exceeded the 1,700-position saving by about 5.04 bits, in all six paired fixtures. Thus gap alone did not determine the score. The prefix, filler and exposure history changed along with absolute placement, so this is not evidence of a position-encoding bug.

For further comparisons, I would distinguish relative gap, absolute placement, type/number of intervening records, and what prefix was processed before the fact. No need for a huge grid unless one of these changes the actual architectural decision.

Why v3’s first overwrite factorial was not enough

In the first overwrite factorial, a nominally “absent” value left behind a same-key ----- record. That is a different observation from no record at all. v4 corrected the comparison by omitting the assertion and using background filler to preserve the layout.

2. Controlled record interference and same-key updates (Probe v4)

What an intervening record changes

Within a fixture the original planted fact, query, and middle slot stayed fixed. The middle slot was filled with background, prose or a structured record; I scored the original correct code in absolute five-byte NLL bits:

Middle-slot content Mean NLL Exact greedy original code
Background; no active record 6.323 2/6
Neutral prose 6.963 2/6
Different key; digits in note 12.436 0/6
Different key; letters in code 12.556 0/6
Different key; dashes in code 12.150 0/6
Different key; digits in code 14.433 0/6
Similar-looking key; digits in code 14.065 0/6
Same key; digits in note 12.028 0/6

The contrasts with the clearest interpretation:

  • Numerical code versus neutral prose: +7.471 bits of correct-code NLL; all six paired signs positive. The format and content both changed, so this is not a pure “one more memory collision” effect.
  • Same distractor key and digits, note versus code label: +1.997 bits for code; positive in 6/6. Even the field name affects the output, potentially through textual expectations, learned representations or both.
  • Numerical versus letter-valued code: +1.877 bits (5/6 positive). Numerical versus dash-valued code: +2.284 bits (5/6 positive). The value type matters under these templates, but this is not a universal ranking of distractors.
  • Similar-looking versus different-looking key: about −0.368 bits with mixed signs. A related same/different-key note comparison was also mixed (about −0.408 bits). Visible string similarity was not a consistent leading effect here. That does not test similarity of learned addresses: visually different keys might have nearby internal representations.

The narrow finding is that an additional structured record can impair retrieval of a queried association, and the penalty changes with the field/value format. That is not yet evidence of a specific collision inside the recurrent memory.

The corrected old/new comparison

The second test placed old/new five-digit codes in early and late slots. Absent means that the corresponding same-key record was not present. The table reports candidate NLL; lower numbers mean higher conditional probability for the supplied code.

Records present Old-code NLL New-code NLL Exact greedy latest code
Neither 20.753 20.968 n/a
Early old only 6.323 22.873 2/6 old
Early new only 22.257 5.666 3/6 new
Late new only 22.943 3.368 5/6 new
Late old only 3.719 22.855 3/6 old
Early old, late new 13.970 7.927 4/6 new
Early new, late old 8.022 12.853 2/6 old
Early old, late new + latest-value cue 14.783 8.652 1/6 new
Early new, late old + latest-value cue 8.630 13.985 1/6 old

The two most diagnostic contrasts keep the target answer fixed:

  1. Compare late new only with early old + late new, scoring the late new target in both: +4.559 bits of NLL with the old record, positive for all six fixtures.
  2. Compare late old only with early new + late old, scoring the late old target in both: +4.303 bits, likewise positive for all six.

Reversing code identities reduces, but cannot remove, string-specific confounds. With both records, the later candidate scored better in 6/6 fixtures (mean margin 6.044 bits) or 5/6 (mean 4.831 bits). Greedy generation matched that later code only 4/6 and 2/6 times. Selecting between two known candidates is easier than generating the value correctly.

The one tested “use the latest value” cue did not rescue this task: later-target mean NLL rose by about 0.725 and 0.608 bits in the two orders, and exact greedy matches were 1/6 each. This is a result for one prompt wording on a base model, not a general claim about instruction following.

What is and is not established

The behavior is consistent with competition between associations and an incomplete preference for the later assertion. It cannot tell whether Splat erased anything, retained both values, mixed addresses or left enough evidence for a later readout to choose imperfectly. A matched softmax baseline might show some of the same effect. The field-label result is a reason to keep input/task framing alongside state updates on the hypothesis list.

The task contract matters: last-write-wins demands accurate current values; auditable history may require both versions. These objectives can favor different update/readout policies.

3. A conditional map from observations to design options

This is a decision guide, not a root-cause diagnosis. Each branch is an option if that distinction would affect the next design.

Keyed retrieval degrades
|
|-- An isolated record degrades with gap
|   |-- Persists when placement/context is well matched
|   |     -> state persistence, normalization, readout resolution
|   |-- Tracks absolute placement/context more than gap
|         -> input encoding, context exposure, position/readout effects
|
|-- Adding structured competing records has an extra penalty
|   |-- Field and value type matter
|   |     -> representation, task framing, write/read competition
|   |-- Measured learned addresses actually overlap
|         -> investigate address separation or selective editing
|
|-- Two same-key values appear
|   |-- Later candidate wins, but has increased absolute NLL
|   |     -> distinguish write-side editing from readout-side selection
|   |-- Earlier candidate repeatedly dominates
|   |     -> inspect new-write strength and query interpretation
|   |-- Extra prompt cue changes ranking but not exact generation
|         -> inspect response/decoding boundary before major surgery
|
|-- Matched hybrids can be compared
    |-- recall rises, bpb and cost acceptable -> retain candidate
    |-- bpb rises, recall does not -> separate modeling gain
    |-- train instability appears -> localize logits/norms/optimizer

If interference is mainly at the write/update boundary

If writes retain too much competing evidence, selective erasure and targeted updates are relevant design options. Delta-rule models update associations using a read-before-write residual; gates can control retention. Gated DeltaNet and Gated DeltaNet-2 illustrate related, but not identical, mechanisms.

Blanket decay can also destroy useful distant facts; it is not selective editing. Historical retrieval may require both versions. Splat-specific state and address behavior would need qualifying before adopting another model’s update rule.

If the associations survive but the readout mixes them

The state may retain distinguishable evidence while the query/readout fails to isolate the right version. A sharper address or readout, explicit recency selection, or a small softmax lookback path could then improve retrieval without changing the Splat write rule. Output NLL measures the functional failure; it does not identify its internal location.

If existing hooks make it cheap, compare state before, between and after two writes, then test whether old/new values are accessible from the same final state. Otherwise a matched small hybrid screen is a reasonable functional first step.

If the visible record format is the dominant factor

Vary the field label, key and code type while preserving positions. The independent v4 data show that changing only note to code in a distractor can alter the queried code’s likelihood. With byte-level inputs, the textual frame is part of the learned task. If loss follows field wording, representation/task framing is a lead; if it follows proximity of measured learned addresses, address interference is a stronger one. Key spelling alone cannot decide.

Minimum discriminating tests, only if needed for the next decision

Decision being made Smallest useful next observation What it would distinguish
Change Splat’s update rule? Read/state probes before and after a conflicting write Old state retained vs revised; with readout caveat
Improve readout instead? Compare candidate ranking, exact generation and a controlled lookback path Available evidence vs usable retrieval
Choose hybrid placement? Same number of softmax blocks; separate recall, bpb, cost and stability Placement benefit vs block-count benefit
Explain field effect? Same keys/values/slots with label or value-type changed Input/task framing vs generic record load
Evaluate strict replacement? One, two, then a few successive same-key updates Latest-value robustness vs accumulating old evidence

These are optional low-cost branch tests, not a new prerequisite for the project. The hybrid and recall comparisons already in the plan may be sufficient for the immediate choice.

4. Hybrid layer placement, training guards, and byte-level cost

Make layer count and layer location separate decisions

The all-Splat probe cannot tell which softmax layers would be best. I would treat at least these alternatives as hypotheses rather than prescriptions:

  • Rank-guided placement: give direct attention to layers whose current representation looks restricted; test whether that heuristic predicts task benefit.
  • Even spacing: distribute opportunities for full-context access through the depth; useful as a simple control for rank-guided placement.
  • A boundary-focused variant: concentrate a similar budget early, late, or around a suspected retrieval bottleneck if existing layer interventions justify it.

For placement, match the number of full-attention blocks first, then report differences in parameters, training bytes, schedule, context and compute. Exact budget parity may be impractical, but an explicit mismatch is better than accidentally crediting placement for a block-count or optimization effect.

A compact result matrix is often enough:

variant          full layers  placement  bpb  keyed recall  stability  decode cost
all Splat        0            --         ...  ...           ...        ...
hybrid rank      N            rank       ...  ...           ...        ...
hybrid spaced    N            spaced     ...  ...           ...        ...
hybrid boundary  N            selected   ...  ...           ...        ...
softmax control  all          --         ...  ...           ...        ...

Representation rank and retrieval are different axes. A sharply low-rank layer may encode a useful high-leverage control signal; a high-rank layer may still fail exact associative retrieval. The reported experiments where rank changes did not straightforwardly track prediction quality make this especially relevant. Rank is a reasonable screening signal, but a rank-guided placement should be judged against a matched spaced control on recall and bpb, not assumed correct in advance.

The systematic hybrid study separates language modeling from retrieval and finds that standalone recurrent strength does not automatically predict hybrid quality. Its ratios are not optimal Beatrix settings. Sequential versus parallel fusion is a separate choice if it becomes relevant to this architecture.

If softmax training becomes unstable

Several mechanisms can produce superficially similar training spikes:

  • Growing Q/K norms or attention logits, producing saturation or sharp attention concentration.
  • Broader parameter/optimizer norm drift, independent of one attention softmax.
  • Changes in clipping, mixed precision, backend or accumulation order.
  • Schedule/warmup mismatch, when a hybrid is trained under a regime tuned for a different architecture.

A few traces—attention-logit scale, Q/K norms, gradient clipping, loss, precision and backend—can help locate the first divergence. QK normalization is most directly motivated by logit saturation. The Kimi K2 report offers different norm-control ideas, not evidence for transplanting its optimizer design into Beatrix.

Cost belongs in the acceptance criteria

A 4,096-byte window is not 4,096 subword tokens, particularly across UTF-8 scripts. A global, windowed or sparse softmax path changes compute, cache, training memory and decode throughput differently from recurrent Splat.

So an acceptable configuration depends on the intended objective:

Recall improves + bpb remains good + cost acceptable
    -> useful hybrid candidate.

bpb improves, recall still weak
    -> useful language-modeling result; memory objective remains open.

Recall improves but decode/memory cost is excessive
    -> examine fewer/lower-window attention blocks or sparse retrieval.

Rank placement and even spacing tie on recall
    -> choose by training stability, speed and implementation simplicity.

Training fails before comparable quality is reached
    -> fix/qualify the optimization path before comparing retrieval claims.

This preserves the purpose of a compact byte-level architecture while giving full attention a clearly defined job rather than treating it as a default replacement for Splat.

5. The dominant final direction and the Sana/Anima sliders

Final-layer direction: a falsifiable temperature interpretation

The reported remove/fix/random-direction controls are more informative than rank alone. I have not reproduced those interventions. A small follow-up, if useful, could test whether the dominant direction mainly changes overall logit scale or also carries token-specific information.

Use fixed prefixes, an intervention that controls for hidden-state norm where feasible, and one scalar output temperature fitted on calibration examples. Then evaluate held-out NLL and entropy and check whether token rankings changed. Because positive temperature scaling preserves rankings, it cannot repair a ranking change caused by the intervention. The outcomes separate naturally:

Token ordering stays the same; one temperature rescues held-out NLL/entropy
    -> supports a largely scale-like role in the tested range.

Token ordering changes, or held-out NLL remains worse after temperature fit
    -> a scalar temperature is insufficient; content-specific effects remain possible.

Results depend heavily on prompt type or byte position
    -> the role may be context-dependent.

Large effects appear only after disruptive intervention
    -> prefer norm-matched / smaller perturbations before interpreting them.

A dominant direction could be a functional control axis rather than “wasted rank.” Large ablations can also push internal states off-distribution; related massive-activation research is only an analogy, not a demonstrated Splat mechanism.

Sana/Anima sliders: follow the signal all the way into the image

For the image-slider work, the conditioning direction, adapter/relay, and rendered image are separate measurement points. Movement of an embedding score does not establish perceptible image control; an effect in Sana need not survive unchanged through Anima or a later Beatrix-driven image path.

slider value / signed direction
    -> conditioning representation
    -> adapter or relay output
    -> generated image
         |-- desired mood / attribute changed?
         |-- scene, objects, relations preserved?
         |-- unrelated attributes drifted?

For paired images, hold prompt, seed, sampler and settings fixed; vary slider strength. Compare zero/random or orthogonal controls where meaningful. Score intended attribute change separately from scene faithfulness and unintended drift. TIFA helps with faithfulness, not mood; an independent judge or small blinded comparison avoids circular scoring.

This fits the staged Beatrix V3 picture tests already described. A conditioning-space effect is not an end-to-end Beatrix-generated image result.

Other knobs that remain distinct

I have no independent measurements for the report’s other open questions—terminal bank behavior, depth scaling, rank as floor versus dial, codebook capacity, document boundaries, arm-count ceiling or lexicon breadth. Two distinctions may still help:

  • Document resets are not selective updates. Resetting state across packed documents could reduce cross-document contamination but destroy useful continuous-context information. A same-key correction within one document needs a different behavioral contract.
  • Arm composition is not bare-trunk recall. Detached/combined arms, routing and generalization to unseen formats need their own controls. My five-digit probe had arms off and says nothing direct about whether combined arms generalize or interfere.

Neither follows from these recall figures.

6. Measurement discipline and a future-reader reproduction map

Preserve the distinctions already present in the write-up

The report already keeps observations, failed controls, revisions and planned tests distinct. For a future reader, a compact checkpoint / intervention / control / metric / result / next branch record could make that structure easy to scan.

My own test is a concrete example. In v3, ----- meant “a record remains, but its field is now five dashes”; it did not mean “the record is absent.” In v4, the absent record was actually omitted. These are both legitimate conditions, but they answer different questions:

A. No record at this slot.
B. Same-key record exists with a placeholder or unknown value.
C. Same-key record exists with a competing factual value.

Zero/random/orthogonal slider controls and remove/rescale interventions also answer different questions. The actual input bytes or tensors define a control—not its label.

Keep four evidence levels visibly separate

Published observation: what the author’s article, repo, logs or model card actually reports, at its stated checkpoint/stage.

Independent inference probe: the six synthetic-fixture results and their exact input/metric conditions above.

Mechanistic hypothesis: possibilities such as write superposition, learned-address collision, readout mixing or field priming; the probe does not choose among them.

External precedent: another architecture’s paper or issue showing a useful idea, not an identified Beatrix root cause. Thus successful Gated DeltaNet retrieval does not mean Beatrix must use that rule, and an output NLL penalty is not proof of a broken Splat kernel.

Minimal contract for reproducing a related comparison

ARTIFACT
  exact model repository/revision; model step; arms on/off;
  code/library versions; device; dtype and backend.

INPUT
  exact prompt bytes or deterministic fixture generator;
  keys, old/new candidate codes, field names, query wording;
  source, distractor and query byte positions;
  truly absent vs placeholder vs conflicting record.

METRICS
  per-byte teacher-forced log probabilities and summed NLL_bits;
  correct-vs-alternative candidate margin;
  independent greedy 5-byte continuation and exact-match rule.

CONTROLS / OUTPUT
  paired input changes as narrow as practical;
  per-case scores, signed paired differences, failures/partial rows;
  sample size and distinction between fixtures and training seeds;
  environment metadata and stated limits of interpretation.

For reproduction, the fixture generator, raw rows, environment, revision and checksums would be the useful minimum. I am not linking a public artifact for my probe here, so its numbers should be treated as a small reported independent observation rather than an externally replicated benchmark.

What would change the next decision?

  • Selecting a hybrid prototype: the observed recall limitation plus your planned controlled hybrid screen may already be enough reason. My small probe is corroborating context, not a prerequisite.
  • Selecting a Splat rewrite rule: state-update versus readout evidence would help, if instrumentation is cheap. Otherwise test whether a hybrid functional path rescues actual keyed retrieval first.
  • Selecting a placement heuristic: same-softmax-count comparisons with separate recall/bpb/stability/cost readouts are more diagnostic than a longer literature list.
  • Interpreting the final direction: held-out scalar-temperature rescue is a relatively cheap discriminant.
  • Claiming image control from Beatrix: that requires actual paired generated-image outcomes through the Beatrix-driven pipeline, not conditioner-space scores alone.

A useful negative control or design lead need not resolve the whole mechanism. These results are best read as additional constraints on the next comparison, not a verdict on Splat or on the project.