For now, here’s what I found from a quick test of the final model:
Small final-checkpoint probe: later values are often preferred, but earlier records still interfere
I ran a small inference-only probe on the published mini-beatrix-3 final checkpoint (step 245,674; arms off), using six synthetic English key–five-digit-code fixtures. The aim was to add a few observations to your forget/rewrite question, and perhaps help interpret the planned hybrid comparisons. This is not a reproduction of your 64-trial battery, a retraining experiment, or an inspection of Splat’s internal writes.
Alongside distance-related degradation, which you already discuss in the Beatrix V3 report, one pattern stood out: an earlier competing record could hurt retrieval even when the later value was preferred.
- Adding an earlier, conflicting same-key record increased the later correct value’s summed five-byte teacher-forced negative log-likelihood (NLL) by 4.56 bits on average (6/6 positive). Swapping which code appeared earlier produced a 4.30-bit penalty (also 6/6 positive).
- Yet the later value received a better candidate score in 6/6 and 5/6 fixtures in those two orders. Exact five-byte greedy generation succeeded less often: 4/6 and 2/6.
- A different key’s numerical
code record increased target NLL by 7.47 bits relative to neutral prose. With the distractor’s key and digits unchanged, replacing its field label note with code added another 2.00 bits.
These are paired changes in teacher-forced NLL for a five-byte answer, not held-out bits per byte, percentage-point changes, or estimates of information stored inside Splat. They come from six fixed synthetic fixtures under one checkpoint—not six training seeds or a representative benchmark.
For the forget/rewrite question, the distinction seems useful: a preference for the later value does not show that the earlier value was erased; interference does not show that updating is impossible. Accumulated writes, address formation, readout competition and prompt framing are still different candidate explanations. Output likelihoods alone cannot tell us which part of Splat is responsible.
A practical default for interpreting the next design
Three questions seem worth keeping separate:
- Retention: does an isolated fact become harder to recover as its byte gap grows? Persistence, normalized accumulation, position and readout could each matter.
- Editing/interference: does a second structured record impose an extra penalty beyond matched filler? Selective erase/write, address separation and better readout are distinct responses.
- Selection versus generation: can the model favor the right candidate but still fail to generate all five digits? Candidate scores and unconstrained continuation measure different things.
Global decay could ease stale interference while harming useful distant facts. Gated DeltaNet and Gated DeltaNet-2 provide useful design precedents for retention and selective editing, not evidence that their update rules are drop-in fixes for Splat.
For the planned hybrid, the default comparison I would favor is keyed recall (including interference), held-out bpb, stability, and inference cost, kept as separate outcomes. With the same number of softmax blocks, rank-guided placement versus even spacing seems especially informative; a boundary-focused variant is optional. Low representation rank may suggest where to look, but does not establish that a layer is disposable or that replacing it improves recall. The systematic hybrid linear-attention study likewise reports that language modeling and associative recall can respond differently to full-attention mixing. Its preferred ratios are not Beatrix-specific prescriptions.
If recall improves without unacceptable bpb or cost, that supports the hybrid for the memory objective. A bpb-only gain is still useful, but answers another question. If softmax training becomes unstable, check Q/K and logit scales, norm growth, clipping and precision before trying Query-Key Normalization or a broader optimizer change.
One methodological caution: the first factor probe treated a literal ----- substitution as its comparison; the later control probe actually removed absent records. The first reports an NLL difference; the latter reports absolute NLL and within-probe contrasts. Those figures should not be pooled. The rest of this post gives the fixtures, limits, and conditional design options for anyone who wants to examine them—not an additional benchmark request.
1. What I ran, what the scores mean, and the distance result
Artifact and scoring contract
The completed tests used AbstractPhil/mini-beatrix-3, checkpoint revision:
0ae0e924965b657525bd49199cd3664a5cf37f14
The run recorded 32 blocks, a 256-byte vocabulary, a 4,096-byte window and arms off. It used a Colab L4, FP32 construction and BF16 autocast (PyTorch 2.11.0+cu130, Transformers 4.57.6). A fixed-input re-score matched; training-runtime parity was not tested.
The two main completed experiments were:
| Probe |
Purpose |
Conditions |
| v3 |
Distance, absolute placement, initial distractor/update conditions |
72 |
| v4 |
Better record-format controls; true absence and reversed update order |
102 |
The two probes reuse the same six synthetic fixture families; these are not held-out documents or fresh independent samples across runs.
For candidate c1...c5, I scored the correct byte sequence with teacher forcing:
NLL_bits(code | prefix)
= -sum(i=1..5) log2 P(c_i | prefix, c_1, ..., c_(i-1))
Lower NLL means a better score for the supplied candidate. Each earlier correct byte is provided before predicting the next one. This can look better than free five-byte generation, which was also measured. Scoring both candidates against one prefix reveals a preference in the output distribution, not whether either association was physically retained.
In v3, saving_bits means:
NLL(correct code | dashed-source version)
- NLL(correct code | planted-code version)
This is the within-pair benefit of the planted code; positive means the planted source improved its score. v4 instead reports absolute NLL_bits or differences between matched absolute scores. Neither is directly comparable to the article’s recall fractions or held-out bpb, and the v3/v4 numbers must not be subtracted from one another.
Distance with query position fixed
The query was near byte position 3,800, while the planted fact moved:
| Source-to-query gap |
Planted-vs-dash saving, mean |
Exact greedy 5 digits |
| 48 bytes |
+31.325 bits |
3/6 |
| 1,200 bytes |
+24.093 bits |
2/6 |
| 3,200 bytes |
+8.454 bits |
0/6 |
All six fixtures lost planted-source advantage between 48 and 3,200 bytes. This supports distance-associated degradation under these prompts, not the magnitude of the published recall deficit and not a matched softmax comparison.
There is a confound: fixing the query while increasing the gap also moves the source’s absolute position. To probe that, I held the gap at 1,200 bytes and moved the layout:
| Query byte position |
Saving, mean |
Exact greedy 5 digits |
| 1,700 |
+19.051 bits |
1/6 |
| 2,700 |
+25.140 bits |
2/6 |
| 3,800 |
+24.093 bits |
2/6 |
At a fixed gap, the 3,800-position saving exceeded the 1,700-position saving by about 5.04 bits, in all six paired fixtures. Thus gap alone did not determine the score. The prefix, filler and exposure history changed along with absolute placement, so this is not evidence of a position-encoding bug.
For further comparisons, I would distinguish relative gap, absolute placement, type/number of intervening records, and what prefix was processed before the fact. No need for a huge grid unless one of these changes the actual architectural decision.
Why v3’s first overwrite factorial was not enough
In the first overwrite factorial, a nominally “absent” value left behind a same-key ----- record. That is a different observation from no record at all. v4 corrected the comparison by omitting the assertion and using background filler to preserve the layout.
2. Controlled record interference and same-key updates (Probe v4)
What an intervening record changes
Within a fixture the original planted fact, query, and middle slot stayed fixed. The middle slot was filled with background, prose or a structured record; I scored the original correct code in absolute five-byte NLL bits:
| Middle-slot content |
Mean NLL |
Exact greedy original code |
| Background; no active record |
6.323 |
2/6 |
| Neutral prose |
6.963 |
2/6 |
Different key; digits in note |
12.436 |
0/6 |
Different key; letters in code |
12.556 |
0/6 |
Different key; dashes in code |
12.150 |
0/6 |
Different key; digits in code |
14.433 |
0/6 |
Similar-looking key; digits in code |
14.065 |
0/6 |
Same key; digits in note |
12.028 |
0/6 |
The contrasts with the clearest interpretation:
- Numerical
code versus neutral prose: +7.471 bits of correct-code NLL; all six paired signs positive. The format and content both changed, so this is not a pure “one more memory collision” effect.
- Same distractor key and digits,
note versus code label: +1.997 bits for code; positive in 6/6. Even the field name affects the output, potentially through textual expectations, learned representations or both.
- Numerical versus letter-valued
code: +1.877 bits (5/6 positive). Numerical versus dash-valued code: +2.284 bits (5/6 positive). The value type matters under these templates, but this is not a universal ranking of distractors.
- Similar-looking versus different-looking key: about −0.368 bits with mixed signs. A related same/different-key
note comparison was also mixed (about −0.408 bits). Visible string similarity was not a consistent leading effect here. That does not test similarity of learned addresses: visually different keys might have nearby internal representations.
The narrow finding is that an additional structured record can impair retrieval of a queried association, and the penalty changes with the field/value format. That is not yet evidence of a specific collision inside the recurrent memory.
The corrected old/new comparison
The second test placed old/new five-digit codes in early and late slots. Absent means that the corresponding same-key record was not present. The table reports candidate NLL; lower numbers mean higher conditional probability for the supplied code.
| Records present |
Old-code NLL |
New-code NLL |
Exact greedy latest code |
| Neither |
20.753 |
20.968 |
n/a |
| Early old only |
6.323 |
22.873 |
2/6 old |
| Early new only |
22.257 |
5.666 |
3/6 new |
| Late new only |
22.943 |
3.368 |
5/6 new |
| Late old only |
3.719 |
22.855 |
3/6 old |
| Early old, late new |
13.970 |
7.927 |
4/6 new |
| Early new, late old |
8.022 |
12.853 |
2/6 old |
| Early old, late new + latest-value cue |
14.783 |
8.652 |
1/6 new |
| Early new, late old + latest-value cue |
8.630 |
13.985 |
1/6 old |
The two most diagnostic contrasts keep the target answer fixed:
- Compare late new only with early old + late new, scoring the late new target in both: +4.559 bits of NLL with the old record, positive for all six fixtures.
- Compare late old only with early new + late old, scoring the late old target in both: +4.303 bits, likewise positive for all six.
Reversing code identities reduces, but cannot remove, string-specific confounds. With both records, the later candidate scored better in 6/6 fixtures (mean margin 6.044 bits) or 5/6 (mean 4.831 bits). Greedy generation matched that later code only 4/6 and 2/6 times. Selecting between two known candidates is easier than generating the value correctly.
The one tested “use the latest value” cue did not rescue this task: later-target mean NLL rose by about 0.725 and 0.608 bits in the two orders, and exact greedy matches were 1/6 each. This is a result for one prompt wording on a base model, not a general claim about instruction following.
What is and is not established
The behavior is consistent with competition between associations and an incomplete preference for the later assertion. It cannot tell whether Splat erased anything, retained both values, mixed addresses or left enough evidence for a later readout to choose imperfectly. A matched softmax baseline might show some of the same effect. The field-label result is a reason to keep input/task framing alongside state updates on the hypothesis list.
The task contract matters: last-write-wins demands accurate current values; auditable history may require both versions. These objectives can favor different update/readout policies.
3. A conditional map from observations to design options
This is a decision guide, not a root-cause diagnosis. Each branch is an option if that distinction would affect the next design.
Keyed retrieval degrades
|
|-- An isolated record degrades with gap
| |-- Persists when placement/context is well matched
| | -> state persistence, normalization, readout resolution
| |-- Tracks absolute placement/context more than gap
| -> input encoding, context exposure, position/readout effects
|
|-- Adding structured competing records has an extra penalty
| |-- Field and value type matter
| | -> representation, task framing, write/read competition
| |-- Measured learned addresses actually overlap
| -> investigate address separation or selective editing
|
|-- Two same-key values appear
| |-- Later candidate wins, but has increased absolute NLL
| | -> distinguish write-side editing from readout-side selection
| |-- Earlier candidate repeatedly dominates
| | -> inspect new-write strength and query interpretation
| |-- Extra prompt cue changes ranking but not exact generation
| -> inspect response/decoding boundary before major surgery
|
|-- Matched hybrids can be compared
|-- recall rises, bpb and cost acceptable -> retain candidate
|-- bpb rises, recall does not -> separate modeling gain
|-- train instability appears -> localize logits/norms/optimizer
If interference is mainly at the write/update boundary
If writes retain too much competing evidence, selective erasure and targeted updates are relevant design options. Delta-rule models update associations using a read-before-write residual; gates can control retention. Gated DeltaNet and Gated DeltaNet-2 illustrate related, but not identical, mechanisms.
Blanket decay can also destroy useful distant facts; it is not selective editing. Historical retrieval may require both versions. Splat-specific state and address behavior would need qualifying before adopting another model’s update rule.
If the associations survive but the readout mixes them
The state may retain distinguishable evidence while the query/readout fails to isolate the right version. A sharper address or readout, explicit recency selection, or a small softmax lookback path could then improve retrieval without changing the Splat write rule. Output NLL measures the functional failure; it does not identify its internal location.
If existing hooks make it cheap, compare state before, between and after two writes, then test whether old/new values are accessible from the same final state. Otherwise a matched small hybrid screen is a reasonable functional first step.
If the visible record format is the dominant factor
Vary the field label, key and code type while preserving positions. The independent v4 data show that changing only note to code in a distractor can alter the queried code’s likelihood. With byte-level inputs, the textual frame is part of the learned task. If loss follows field wording, representation/task framing is a lead; if it follows proximity of measured learned addresses, address interference is a stronger one. Key spelling alone cannot decide.
Minimum discriminating tests, only if needed for the next decision
| Decision being made |
Smallest useful next observation |
What it would distinguish |
| Change Splat’s update rule? |
Read/state probes before and after a conflicting write |
Old state retained vs revised; with readout caveat |
| Improve readout instead? |
Compare candidate ranking, exact generation and a controlled lookback path |
Available evidence vs usable retrieval |
| Choose hybrid placement? |
Same number of softmax blocks; separate recall, bpb, cost and stability |
Placement benefit vs block-count benefit |
| Explain field effect? |
Same keys/values/slots with label or value-type changed |
Input/task framing vs generic record load |
| Evaluate strict replacement? |
One, two, then a few successive same-key updates |
Latest-value robustness vs accumulating old evidence |
These are optional low-cost branch tests, not a new prerequisite for the project. The hybrid and recall comparisons already in the plan may be sufficient for the immediate choice.
4. Hybrid layer placement, training guards, and byte-level cost
Make layer count and layer location separate decisions
The all-Splat probe cannot tell which softmax layers would be best. I would treat at least these alternatives as hypotheses rather than prescriptions:
- Rank-guided placement: give direct attention to layers whose current representation looks restricted; test whether that heuristic predicts task benefit.
- Even spacing: distribute opportunities for full-context access through the depth; useful as a simple control for rank-guided placement.
- A boundary-focused variant: concentrate a similar budget early, late, or around a suspected retrieval bottleneck if existing layer interventions justify it.
For placement, match the number of full-attention blocks first, then report differences in parameters, training bytes, schedule, context and compute. Exact budget parity may be impractical, but an explicit mismatch is better than accidentally crediting placement for a block-count or optimization effect.
A compact result matrix is often enough:
variant full layers placement bpb keyed recall stability decode cost
all Splat 0 -- ... ... ... ...
hybrid rank N rank ... ... ... ...
hybrid spaced N spaced ... ... ... ...
hybrid boundary N selected ... ... ... ...
softmax control all -- ... ... ... ...
Representation rank and retrieval are different axes. A sharply low-rank layer may encode a useful high-leverage control signal; a high-rank layer may still fail exact associative retrieval. The reported experiments where rank changes did not straightforwardly track prediction quality make this especially relevant. Rank is a reasonable screening signal, but a rank-guided placement should be judged against a matched spaced control on recall and bpb, not assumed correct in advance.
The systematic hybrid study separates language modeling from retrieval and finds that standalone recurrent strength does not automatically predict hybrid quality. Its ratios are not optimal Beatrix settings. Sequential versus parallel fusion is a separate choice if it becomes relevant to this architecture.
If softmax training becomes unstable
Several mechanisms can produce superficially similar training spikes:
- Growing Q/K norms or attention logits, producing saturation or sharp attention concentration.
- Broader parameter/optimizer norm drift, independent of one attention softmax.
- Changes in clipping, mixed precision, backend or accumulation order.
- Schedule/warmup mismatch, when a hybrid is trained under a regime tuned for a different architecture.
A few traces—attention-logit scale, Q/K norms, gradient clipping, loss, precision and backend—can help locate the first divergence. QK normalization is most directly motivated by logit saturation. The Kimi K2 report offers different norm-control ideas, not evidence for transplanting its optimizer design into Beatrix.
Cost belongs in the acceptance criteria
A 4,096-byte window is not 4,096 subword tokens, particularly across UTF-8 scripts. A global, windowed or sparse softmax path changes compute, cache, training memory and decode throughput differently from recurrent Splat.
So an acceptable configuration depends on the intended objective:
Recall improves + bpb remains good + cost acceptable
-> useful hybrid candidate.
bpb improves, recall still weak
-> useful language-modeling result; memory objective remains open.
Recall improves but decode/memory cost is excessive
-> examine fewer/lower-window attention blocks or sparse retrieval.
Rank placement and even spacing tie on recall
-> choose by training stability, speed and implementation simplicity.
Training fails before comparable quality is reached
-> fix/qualify the optimization path before comparing retrieval claims.
This preserves the purpose of a compact byte-level architecture while giving full attention a clearly defined job rather than treating it as a default replacement for Splat.
5. The dominant final direction and the Sana/Anima sliders
Final-layer direction: a falsifiable temperature interpretation
The reported remove/fix/random-direction controls are more informative than rank alone. I have not reproduced those interventions. A small follow-up, if useful, could test whether the dominant direction mainly changes overall logit scale or also carries token-specific information.
Use fixed prefixes, an intervention that controls for hidden-state norm where feasible, and one scalar output temperature fitted on calibration examples. Then evaluate held-out NLL and entropy and check whether token rankings changed. Because positive temperature scaling preserves rankings, it cannot repair a ranking change caused by the intervention. The outcomes separate naturally:
Token ordering stays the same; one temperature rescues held-out NLL/entropy
-> supports a largely scale-like role in the tested range.
Token ordering changes, or held-out NLL remains worse after temperature fit
-> a scalar temperature is insufficient; content-specific effects remain possible.
Results depend heavily on prompt type or byte position
-> the role may be context-dependent.
Large effects appear only after disruptive intervention
-> prefer norm-matched / smaller perturbations before interpreting them.
A dominant direction could be a functional control axis rather than “wasted rank.” Large ablations can also push internal states off-distribution; related massive-activation research is only an analogy, not a demonstrated Splat mechanism.
Sana/Anima sliders: follow the signal all the way into the image
For the image-slider work, the conditioning direction, adapter/relay, and rendered image are separate measurement points. Movement of an embedding score does not establish perceptible image control; an effect in Sana need not survive unchanged through Anima or a later Beatrix-driven image path.
slider value / signed direction
-> conditioning representation
-> adapter or relay output
-> generated image
|-- desired mood / attribute changed?
|-- scene, objects, relations preserved?
|-- unrelated attributes drifted?
For paired images, hold prompt, seed, sampler and settings fixed; vary slider strength. Compare zero/random or orthogonal controls where meaningful. Score intended attribute change separately from scene faithfulness and unintended drift. TIFA helps with faithfulness, not mood; an independent judge or small blinded comparison avoids circular scoring.
This fits the staged Beatrix V3 picture tests already described. A conditioning-space effect is not an end-to-end Beatrix-generated image result.
Other knobs that remain distinct
I have no independent measurements for the report’s other open questions—terminal bank behavior, depth scaling, rank as floor versus dial, codebook capacity, document boundaries, arm-count ceiling or lexicon breadth. Two distinctions may still help:
- Document resets are not selective updates. Resetting state across packed documents could reduce cross-document contamination but destroy useful continuous-context information. A same-key correction within one document needs a different behavioral contract.
- Arm composition is not bare-trunk recall. Detached/combined arms, routing and generalization to unseen formats need their own controls. My five-digit probe had arms off and says nothing direct about whether combined arms generalize or interfere.
Neither follows from these recall figures.
6. Measurement discipline and a future-reader reproduction map
Preserve the distinctions already present in the write-up
The report already keeps observations, failed controls, revisions and planned tests distinct. For a future reader, a compact checkpoint / intervention / control / metric / result / next branch record could make that structure easy to scan.
My own test is a concrete example. In v3, ----- meant “a record remains, but its field is now five dashes”; it did not mean “the record is absent.” In v4, the absent record was actually omitted. These are both legitimate conditions, but they answer different questions:
A. No record at this slot.
B. Same-key record exists with a placeholder or unknown value.
C. Same-key record exists with a competing factual value.
Zero/random/orthogonal slider controls and remove/rescale interventions also answer different questions. The actual input bytes or tensors define a control—not its label.
Keep four evidence levels visibly separate
Published observation: what the author’s article, repo, logs or model card actually reports, at its stated checkpoint/stage.
Independent inference probe: the six synthetic-fixture results and their exact input/metric conditions above.
Mechanistic hypothesis: possibilities such as write superposition, learned-address collision, readout mixing or field priming; the probe does not choose among them.
External precedent: another architecture’s paper or issue showing a useful idea, not an identified Beatrix root cause. Thus successful Gated DeltaNet retrieval does not mean Beatrix must use that rule, and an output NLL penalty is not proof of a broken Splat kernel.
Minimal contract for reproducing a related comparison
ARTIFACT
exact model repository/revision; model step; arms on/off;
code/library versions; device; dtype and backend.
INPUT
exact prompt bytes or deterministic fixture generator;
keys, old/new candidate codes, field names, query wording;
source, distractor and query byte positions;
truly absent vs placeholder vs conflicting record.
METRICS
per-byte teacher-forced log probabilities and summed NLL_bits;
correct-vs-alternative candidate margin;
independent greedy 5-byte continuation and exact-match rule.
CONTROLS / OUTPUT
paired input changes as narrow as practical;
per-case scores, signed paired differences, failures/partial rows;
sample size and distinction between fixtures and training seeds;
environment metadata and stated limits of interpretation.
For reproduction, the fixture generator, raw rows, environment, revision and checksums would be the useful minimum. I am not linking a public artifact for my probe here, so its numbers should be treated as a small reported independent observation rather than an externally replicated benchmark.
What would change the next decision?
- Selecting a hybrid prototype: the observed recall limitation plus your planned controlled hybrid screen may already be enough reason. My small probe is corroborating context, not a prerequisite.
- Selecting a Splat rewrite rule: state-update versus readout evidence would help, if instrumentation is cheap. Otherwise test whether a hybrid functional path rescues actual keyed retrieval first.
- Selecting a placement heuristic: same-softmax-count comparisons with separate recall/bpb/stability/cost readouts are more diagnostic than a longer literature list.
- Interpreting the final direction: held-out scalar-temperature rescue is a relatively cheap discriminant.
- Claiming image control from Beatrix: that requires actual paired generated-image outcomes through the Beatrix-driven pipeline, not conditioner-space scores alone.
A useful negative control or design lead need not resolve the whole mechanism. These results are best read as additional constraints on the next comparison, not a verdict on Splat or on the project.