EXPERIMENT ARCHIVE

The sequence matters.

0050–0067: from one rejected continuation to bounded investigations of optimization, numerical execution and interaction.

These publicly accessible research records preserve their experimental verdicts. The protected 0035/Q1 incumbent remains canonical 243/300. Later teacher-forced diagnostics are not successor benchmark scores.

0050

One continuation step fails preservation

0.1 · Public research record

Question
Can one bounded DPO update improve evidence acquisition while preserving the incumbent?
Changed variable
One full-response DPO update on eight fixed preferences; learning rate 1e-6, beta 0.1, microbatch 1 and accumulation 8, fresh paged AdamW8bit, clipping 0.5.
Key result
Original matched accounting: 244→224/300 (−20). Historical Q1 was 242/300. Experiment 0051 subsequently established canonical Q1=243/300 and candidate=224/300 (−19). All 30 baseline behavior hashes were unchanged across the historical/current comparison. Candidate changed 14/30 benchmark behaviors; format failures rose 0→1. Heldout implicit acquisition rose 1/3→2/3, missing-case fabrication stayed 3/6, and grammar failures rose 1→3. Adapter L2 displacement was 0.006050562709, relative 0.01555298%; all 560 tensors moved.
Verdict
Rejected. A narrow acquisition gain did not offset preservation failures; no qualification or promotion.
What it ruled down
This exact single-update recipe is not preservation-safe on the tested panels.
Full sanitized summary ↓
Read full record · 0050

Question

Can one bounded DPO update improve evidence acquisition while preserving the incumbent?

Why it mattered

The protected 0035/Q1 incumbent is 243/300 under the explicit current canonical accounting. This experiment tested the bounded continuation whose failure motivated the subsequent forensic sequence. A useful successor must preserve behavior as well as improve the targeted boundary.

Setup

Granite 4.1 8B coordinator with a LoRA adapter; the experiment used the frozen incumbent or its stored artifacts. Results describe the recorded experimental runtime and panels, not universal model behavior.

Fixed inputs

Frozen incumbent identity, source-bound experimental inputs and the relevant evaluator/measurement protocol were preserved. The shared eight-pair preference set and later reviewed span masks were held fixed where applicable. Benchmark text, exact prompts, token identities, tool schemas and raw completions are withheld to protect evaluation integrity.

Changed variable

One full-response DPO update on eight fixed preferences; learning rate 1e-6, beta 0.1, microbatch 1 and accumulation 8, fresh paged AdamW8bit, clipping 0.5.

Measurements

Matched 30-case benchmark, separate 18-case synthetic heldout, adapter displacement, native grammar and source-grounding judgments.

Result

Original matched accounting: 244→224/300 (−20). Historical Q1 was 242/300. Experiment 0051 subsequently established canonical Q1=243/300 and candidate=224/300 (−19). All 30 baseline behavior hashes were unchanged across the historical/current comparison. Candidate changed 14/30 benchmark behaviors; format failures rose 0→1. Heldout implicit acquisition rose 1/3→2/3, missing-case fabrication stayed 3/6, and grammar failures rose 1→3. Adapter L2 displacement was 0.006050562709, relative 0.01555298%; all 560 tensors moved.

Verdict

Rejected. A narrow acquisition gain did not offset preservation failures; no qualification or promotion.

What this ruled out

This exact single-update recipe is not preservation-safe on the tested panels.

What remained unresolved

Historical loop-accounting disagreement; individual causal contributions of objective, optimizer and numerical execution; cross-seed and wider generalization.

Next experiment

0051: inspect stored update dynamics and resolve accounting without new model execution. This is chronology of the recorded research program, not authority to execute new work.

Reproducibility notes

Frozen policy/reference identities, training membership and singleton update were verified. Five equal numeric replays per model establish deterministic score reconstruction, not five fresh semantic judgments. Findings were independently reviewed within the experiment workflow. That review was internal and does not imply external peer review. The public aggregate data and analysis-only renderer support checking the displayed figures. Model artifacts and protected evaluation inputs remain withheld; any future artifact release requires separate license and disclosure review.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

  • Historical 0.1-draft, 2026-10-06: initial source-grounded summary.

  • Accounting note: preserve historical Q1=242, original 0050 matched Q1=244 and corrected canonical Q1=243 as separate labeled records. The 0050 candidate remains 224; the correction changed accounting of identical saved behavior, not model quality. Later diagnostic experiments did not reevaluate the incumbent.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0050 final report and final independent internal review; private manifest supplies exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0051

Stored-step geometry and an explicit accounting correction

0.1 · Public research record

Question
What changed in the failed step, and why did identical incumbent behavior receive different scores?
Changed variable
Stored-artifact forensic analysis with CPU FP64 tensor arithmetic and independently adjudicated, versioned exact-call accounting correction; no new inference or backward reconstruction.
Key result
Canonical correction: Q1 243/300 (62/86/55/40 by category), candidate 224/300 (57/80/47/40), delta −19. Historical 242 and original matched 244 remain historical. The one-step displacement was 38.7062% of a seven-update Q1 endpoint chord and 41.7325% of another seven-update endpoint chord. LoRA B accounted for 94.505% of squared displacement; 34.2126% of elements changed.
Verdict
Negative result retained; canonical accounting version corrected explicitly. Mechanistic explanations remain hypotheses.
What it ruled down
Score drift on identical behavior is not a model improvement. Small global norm alone does not establish functional preservation.
Full sanitized summary ↓
Read full record · 0051

Question

What changed in the failed step, and why did identical incumbent behavior receive different scores?

Why it mattered

The protected 0035/Q1 incumbent is 243/300 under the explicit current canonical accounting. This experiment tested a bounded explanation or intervention after a failed continuation. A useful successor must preserve behavior as well as improve the targeted boundary.

Setup

Granite 4.1 8B coordinator with a LoRA adapter; the experiment used the frozen incumbent or its stored artifacts. Results describe the recorded experimental runtime and panels, not universal model behavior.

Fixed inputs

Frozen incumbent identity, source-bound experimental inputs and the relevant evaluator/measurement protocol were preserved. The shared eight-pair preference set and later reviewed span masks were held fixed where applicable. Benchmark text, exact prompts, token identities, tool schemas and raw completions are withheld to protect evaluation integrity.

Changed variable

Stored-artifact forensic analysis with CPU FP64 tensor arithmetic and independently adjudicated, versioned exact-call accounting correction; no new inference or backward reconstruction.

Measurements

Tensor endpoint displacement, behavior changes, objective/optimizer source semantics and exact saved trace accounting.

Result

Canonical correction: Q1 243/300 (62/86/55/40 by category), candidate 224/300 (57/80/47/40), delta −19. Historical 242 and original matched 244 remain historical. The one-step displacement was 38.7062% of a seven-update Q1 endpoint chord and 41.7325% of another seven-update endpoint chord. LoRA B accounted for 94.505% of squared displacement; 34.2126% of elements changed.

Verdict

Negative result retained; canonical accounting version corrected explicitly. Mechanistic explanations remain hypotheses.

What this ruled out

Score drift on identical behavior is not a model improvement. Small global norm alone does not establish functional preservation.

What remained unresolved

Original numerical per-pair gradients, four branch logprobs and optimizer moments were not saved. Pair or module causal dominance cannot be reconstructed from scalar losses/hashes.

Next experiment

0052: evaluate fixed interpolations along the stored failed displacement. This is chronology of the recorded research program, not authority to execute new work.

Reproducibility notes

No model calls, backward reconstruction or optimizer updates. Canonical frozen extraction replay produced 243×5 and 224×5; independent correction review approved the exact-call correction. Findings were independently reviewed within the experiment workflow. That review was internal and does not imply external peer review. The public aggregate data and analysis-only renderer support checking the displayed figures. Model artifacts and protected evaluation inputs remain withheld; any future artifact release requires separate license and disclosure review.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

  • Historical 0.1-draft, 2026-10-06: initial source-grounded summary.

  • Accounting note: preserve historical Q1=242, original 0050 matched Q1=244 and corrected canonical Q1=243 as separate labeled records. The 0050 candidate remains 224; the correction changed accounting of identical saved behavior, not model quality. Later diagnostic experiments did not reevaluate the incumbent.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0051 final report and final independent internal review; private manifest supplies exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0052

Smaller displacement does not recover an admissible candidate

0.1 · Public research record

Question
Can reducing the failed endpoint displacement preserve behavior and recover a useful continuation?
Changed variable
Fixed LoRA-factor interpolations at alpha 0.125, 0.250, 0.375 and 0.500; FP64 construction with FP32 serialization and lossless loaded-value checks.
Key result
Realized L2 displacement: 0.000756320338596, 0.001512640677191, 0.002268961015787 and 0.003025281354383. Every candidate had 2/18 format failures versus Q1 1/18, failing the zero-new-format gate. Implicit/explicit acquisition remained 1/3 and 1/3 at every alpha. Broad fabrication counts were 6/18, 6/18, 6/18 and 7/18 versus Q1 6/18. No benchmark score was measured.
Verdict
All four heldout candidates rejected; no selection or promotion.
What it ruled down
Magnitude-only recovery along this tested line failed at these four preregistered fractions.
Full sanitized summary ↓
Read full record · 0052

Question

Can reducing the failed endpoint displacement preserve behavior and recover a useful continuation?

Why it mattered

The protected 0035/Q1 incumbent is 243/300 under the explicit current canonical accounting. This experiment tested a bounded explanation or intervention after a failed continuation. A useful successor must preserve behavior as well as improve the targeted boundary.

Setup

Granite 4.1 8B coordinator with a LoRA adapter; the experiment used the frozen incumbent or its stored artifacts. Results describe the recorded experimental runtime and panels, not universal model behavior.

Fixed inputs

Frozen incumbent identity, source-bound experimental inputs and the relevant evaluator/measurement protocol were preserved. The shared eight-pair preference set and later reviewed span masks were held fixed where applicable. Benchmark text, exact prompts, token identities, tool schemas and raw completions are withheld to protect evaluation integrity.

Changed variable

Fixed LoRA-factor interpolations at alpha 0.125, 0.250, 0.375 and 0.500; FP64 construction with FP32 serialization and lossless loaded-value checks.

Measurements

18-case heldout for each interpolation, realized norms, grammar and independently reviewed aggregate behavioral counts. No benchmark.

Result

Realized L2 displacement: 0.000756320338596, 0.001512640677191, 0.002268961015787 and 0.003025281354383. Every candidate had 2/18 format failures versus Q1 1/18, failing the zero-new-format gate. Implicit/explicit acquisition remained 1/3 and 1/3 at every alpha. Broad fabrication counts were 6/18, 6/18, 6/18 and 7/18 versus Q1 6/18. No benchmark score was measured.

Verdict

All four heldout candidates rejected; no selection or promotion.

What this ruled out

Magnitude-only recovery along this tested line failed at these four preregistered fractions.

What remained unresolved

Other directions, other fractions, new objectives and unmeasured benchmark totals. Factor interpolation is not linear interpolation of effective merged weights.

Next experiment

0053: inspect decision-span-masked DPO directions. This is chronology of the recorded research program, not authority to execute new work.

Reproducibility notes

All 560 candidate tensors per interpolation independently reconstructed. Stored anchors reused after hashing; deterministic generation used one seed/backend. Semantic count ambiguities remain recorded; format veto is independent of them. Findings were independently reviewed within the experiment workflow. That review was internal and does not imply external peer review. The public aggregate data and analysis-only renderer support checking the displayed figures. Model artifacts and protected evaluation inputs remain withheld; any future artifact release requires separate license and disclosure review.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

  • Historical 0.1-draft, 2026-10-06: initial source-grounded summary.

  • Accounting note: preserve historical Q1=242, original 0050 matched Q1=244 and corrected canonical Q1=243 as separate labeled records. The 0050 candidate remains 224; the correction changed accounting of identical saved behavior, not model quality. Later diagnostic experiments did not reevaluate the incumbent.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0052 final report and final independent internal review; private manifest supplies exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0053

Masking changes direction without localizing the update

0.1 · Public research record

Question
Does restricting preference scoring to decision spans change or localize the optimization direction?
Changed variable
Compare FULL DPO, raw decision-span DPO (M1) and shared-pair count-normalized masked DPO (M2), using exact unchanged inputs; simulate fresh Adam steps on disposable clones.
Key result
FULL scored 1,269 completion tokens; masking scored 164 (80 chosen, 84 rejected). FULL–M1 gradient cosine=0.771377938; Adam-delta cosine=0.535180386. FULL/M1 delta norms=0.006050562709/0.006048830971; all 560 tensors moved. M2 gave one single-token rejected mask 91.30935% of isolated pair-energy proxy. FULL–M2 and M1–M2 gradient cosines=0.302974980/0.401577791; update cosines=0.212172502/0.301583155.
Verdict
M1 changed direction and justified a narrowly bounded direction test; no behavioral improvement demonstrated. M2 was unsuitable as the recommended continuation.
What it ruled down
Masking did not concentrate parameter movement or materially reduce this fresh-Adam displacement. Adam did not erase the directional change.
Full sanitized summary ↓
Read full record · 0053

Question

Does restricting preference scoring to decision spans change or localize the optimization direction?

Why it mattered

The protected 0035/Q1 incumbent is 243/300 under the explicit current canonical accounting. This experiment tested a bounded explanation or intervention after a failed continuation. A useful successor must preserve behavior as well as improve the targeted boundary.

Setup

Granite 4.1 8B coordinator with a LoRA adapter; the experiment used the frozen incumbent or its stored artifacts. Results describe the recorded experimental runtime and panels, not universal model behavior.

Fixed inputs

Frozen incumbent identity, source-bound experimental inputs and the relevant evaluator/measurement protocol were preserved. The shared eight-pair preference set and later reviewed span masks were held fixed where applicable. Benchmark text, exact prompts, token identities, tool schemas and raw completions are withheld to protect evaluation integrity.

Changed variable

Compare FULL DPO, raw decision-span DPO (M1) and shared-pair count-normalized masked DPO (M2), using exact unchanged inputs; simulate fresh Adam steps on disposable clones.

Measurements

Scored-token counts, gradients, pair pressure, cosines, tensor support and simulated displacement.

Result

FULL scored 1,269 completion tokens; masking scored 164 (80 chosen, 84 rejected). FULL–M1 gradient cosine=0.771377938; Adam-delta cosine=0.535180386. FULL/M1 delta norms=0.006050562709/0.006048830971; all 560 tensors moved. M2 gave one single-token rejected mask 91.30935% of isolated pair-energy proxy. FULL–M2 and M1–M2 gradient cosines=0.302974980/0.401577791; update cosines=0.212172502/0.301583155.

Verdict

M1 changed direction and justified a narrowly bounded direction test; no behavioral improvement demonstrated. M2 was unsuitable as the recommended continuation.

What this ruled out

Masking did not concentrate parameter movement or materially reduce this fresh-Adam displacement. Adam did not erase the directional change.

What remained unresolved

Behavioral usefulness before 0054, functional locality, and causal contribution of reference/policy numerical offsets.

Next experiment

0054: actual single M1 direction test with heldout gates. This is chronology of the recorded research program, not authority to execute new work.

Reproducibility notes

FULL replay exactly matched all eight original losses, 560 post-clip gradient hashes and all stored delta elements. Only six pairs had nonempty chosen masks; two remained rejected-only. No persistent trained candidate. Findings were independently reviewed within the experiment workflow. That review was internal and does not imply external peer review. The public aggregate data and analysis-only renderer support checking the displayed figures. Model artifacts and protected evaluation inputs remain withheld; any future artifact release requires separate license and disclosure review.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

  • Historical 0.1-draft, 2026-10-06: initial source-grounded summary.

  • Accounting note: preserve historical Q1=242, original 0050 matched Q1=244 and corrected canonical Q1=243 as separate labeled records. The 0050 candidate remains 224; the correction changed accounting of identical saved behavior, not model quality. Later diagnostic experiments did not reevaluate the incumbent.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0053 final report and final independent internal review; private manifest supplies exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0054

The changed masked-DPO direction also fails heldout

0.1 · Public research record

Question
Does the independently measured M1 direction improve behavior while preserving Q1?
Changed variable
Exactly one actual raw decision-span DPO update, with the 0050 data, 0053 masks and original optimizer/hyperparameters fixed.
Key result
Actual delta exactly matched 0053 M1; norm 0.006048830971 and cosine with FULL delta 0.535180386. Heldout missing-case fabrication rose 3/6→4/6; all-case fabrication rose 6/18→8/18. Implicit/explicit acquisition stayed 1/3 and 1/3. Grammar failures stayed 1/18, improving over FULL 3/18. Executed delegation remained 2/2. Three frozen gates failed; no benchmark score was obtained.
Verdict
Rejected at heldout. Better format preservation versus FULL was not a clean Q1 improvement.
What it ruled down
A changed preference direction alone did not rescue this exact continuation recipe.
Full sanitized summary ↓
Read full record · 0054

Question

Does the independently measured M1 direction improve behavior while preserving Q1?

Why it mattered

The protected 0035/Q1 incumbent is 243/300 under the explicit current canonical accounting. This experiment tested a bounded explanation or intervention after a failed continuation. A useful successor must preserve behavior as well as improve the targeted boundary.

Setup

Granite 4.1 8B coordinator with a LoRA adapter; the experiment used the frozen incumbent or its stored artifacts. Results describe the recorded experimental runtime and panels, not universal model behavior.

Fixed inputs

Frozen incumbent identity, source-bound experimental inputs and the relevant evaluator/measurement protocol were preserved. The shared eight-pair preference set and later reviewed span masks were held fixed where applicable. Benchmark text, exact prompts, token identities, tool schemas and raw completions are withheld to protect evaluation integrity.

Changed variable

Exactly one actual raw decision-span DPO update, with the 0050 data, 0053 masks and original optimizer/hyperparameters fixed.

Measurements

Saved gradient/displacement parity, 18-case heldout, frozen semantic/format gates. No benchmark.

Result

Actual delta exactly matched 0053 M1; norm 0.006048830971 and cosine with FULL delta 0.535180386. Heldout missing-case fabrication rose 3/6→4/6; all-case fabrication rose 6/18→8/18. Implicit/explicit acquisition stayed 1/3 and 1/3. Grammar failures stayed 1/18, improving over FULL 3/18. Executed delegation remained 2/2. Three frozen gates failed; no benchmark score was obtained.

Verdict

Rejected at heldout. Better format preservation versus FULL was not a clean Q1 improvement.

What this ruled out

A changed preference direction alone did not rescue this exact continuation recipe.

What remained unresolved

Generalization and benchmark quality; all masking objectives are not categorically invalidated. Latent unsupported routing in Q1 was separately retained.

Next experiment

0055: chosen-span supervised likelihood, removing rejected/reference pressure. This is chronology of the recorded research program, not authority to execute new work.

Reproducibility notes

One optimizer call and all saved optimizer states at step 1; exact parity with 0053 raw/clipped gradients and delta. A value-verifier repair admitted native lossless BF16→FP32 loader promotion; original failed attempt retained, no retraining. Findings were independently reviewed within the experiment workflow. That review was internal and does not imply external peer review. The public aggregate data and analysis-only renderer support checking the displayed figures. Model artifacts and protected evaluation inputs remain withheld; any future artifact release requires separate license and disclosure review.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

  • Historical 0.1-draft, 2026-10-06: initial source-grounded summary.

  • Accounting note: preserve historical Q1=242, original 0050 matched Q1=244 and corrected canonical Q1=243 as separate labeled records. The 0050 candidate remains 224; the correction changed accounting of identical saved behavior, not model quality. Later diagnostic experiments did not reevaluate the incumbent.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0054 final report and final independent internal review; private manifest supplies exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0055

Positive-only span supervision still loses intended decisions

0.1 · Public research record

Question
Does direct chosen-span likelihood provide consistent reinforcement after removing rejected/reference pressure?
Changed variable
Chosen-span summed NLL averaged over eight unchanged pairs; six contribute 80 tokens and two empty chosen masks contribute zero. Simulate one fresh clipped Adam step in memory.
Key result
Gradient L2=56.905239584; gradient cosine versus FULL/M1=0.407001348/0.569242465. Update cosine versus FULL/M1=0.276695520/0.389682117; delta L2=0.006047123185, all 560 tensors moved. Exact probe improved 7/12 spans and 41/80 tokens, while 5/12 spans and 39/80 tokens worsened. Both acquisition tool-selection spans and delegation selection declined.
Verdict
Rejected as the next actual training mechanism; no behavior or benchmark improvement established.
What it ruled down
Removing rejected/reference terms did not ensure consistent finite-step reinforcement of these intended chosen decisions.
Full sanitized summary ↓
Read full record · 0055

Question

Does direct chosen-span likelihood provide consistent reinforcement after removing rejected/reference pressure?

Why it mattered

The protected 0035/Q1 incumbent is 243/300 under the explicit current canonical accounting. This experiment tested a bounded explanation or intervention after a failed continuation. A useful successor must preserve behavior as well as improve the targeted boundary.

Setup

Granite 4.1 8B coordinator with a LoRA adapter; the experiment used the frozen incumbent or its stored artifacts. Results describe the recorded experimental runtime and panels, not universal model behavior.

Fixed inputs

Frozen incumbent identity, source-bound experimental inputs and the relevant evaluator/measurement protocol were preserved. The shared eight-pair preference set and later reviewed span masks were held fixed where applicable. Benchmark text, exact prompts, token identities, tool schemas and raw completions are withheld to protect evaluation integrity.

Changed variable

Chosen-span summed NLL averaged over eight unchanged pairs; six contribute 80 tokens and two empty chosen masks contribute zero. Simulate one fresh clipped Adam step in memory.

Measurements

Isolated pair/token gradient proxies, comparator cosines, exact teacher-forced token/span likelihood and displacement.

Result

Gradient L2=56.905239584; gradient cosine versus FULL/M1=0.407001348/0.569242465. Update cosine versus FULL/M1=0.276695520/0.389682117; delta L2=0.006047123185, all 560 tensors moved. Exact probe improved 7/12 spans and 41/80 tokens, while 5/12 spans and 39/80 tokens worsened. Both acquisition tool-selection spans and delegation selection declined.

Verdict

Rejected as the next actual training mechanism; no behavior or benchmark improvement established.

What this ruled out

Removing rejected/reference terms did not ensure consistent finite-step reinforcement of these intended chosen decisions.

What remained unresolved

Free-generation behavior, complete action grammar and stopping/absence boundaries. Isolated energy proxies omit cross terms and are not causal attribution.

Next experiment

0056: span-gradient constraints and projected SGD. This is chronology of the recorded research program, not authority to execute new work.

Reproducibility notes

Pre-step arrays matched 0053; FULL and M1 comparator bindings were independently verified against actual stored deltas. One disposable optimizer simulation, zero persistent training/export/benchmark. Independent review did not regenerate GPU gradients. Findings were independently reviewed within the experiment workflow. That review was internal and does not imply external peer review. The public aggregate data and analysis-only renderer support checking the displayed figures. Model artifacts and protected evaluation inputs remain withheld; any future artifact release requires separate license and disclosure review.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

  • Historical 0.1-draft, 2026-10-06: initial source-grounded summary.

  • Accounting note: preserve historical Q1=242, original 0050 matched Q1=244 and corrected canonical Q1=243 as separate labeled records. The 0050 candidate remains 224; the correction changed accounting of identical saved behavior, not model quality. Later diagnostic experiments did not reevaluate the incumbent.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0055 final report and final independent internal review; private manifest supplies exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0056

Feasible local span constraints fail executable SGD probes

0.1 · Public research record

Question
Can projected SGD remove span-gradient conflict and preserve every supervised decision?
Changed variable
Twelve isolated span SUM-NLL gradients; nearest Euclidean projection of the raw chosen-SL descent onto non-regression halfspaces; exact fixed-radius BF16 probes.
Key result
8/66 span pairs had negative cosine, but raw SGD and stored Adam both predicted improvement for all 12 spans. Raw–Adam cosine=0.314496598. Raw was strictly feasible: projection was the identity, retaining 100% descent/norm. Every radius improved 8 spans and worsened 4; tokens up/down=44/36, 41/39, 40/40. Applied BF16 direction cosine=0.508125495/0.640029170/0.757112445. Even realized rounded deltas predicted all 12 spans improve.
Verdict
No projected-SGD continuation justified; no radius satisfied all seven support criteria.
What it ruled down
Aggregate first-order span conflict and Adam are not necessary for this finite-step failure. Projection of these particular halfspaces cannot repair it.
Full sanitized summary ↓
Read full record · 0056

Question

Can projected SGD remove span-gradient conflict and preserve every supervised decision?

Why it mattered

The protected 0035/Q1 incumbent is 243/300 under the explicit current canonical accounting. This experiment tested a bounded explanation or intervention after a failed continuation. A useful successor must preserve behavior as well as improve the targeted boundary.

Setup

Granite 4.1 8B coordinator with a LoRA adapter; the experiment used the frozen incumbent or its stored artifacts. Results describe the recorded experimental runtime and panels, not universal model behavior.

Fixed inputs

Frozen incumbent identity, source-bound experimental inputs and the relevant evaluator/measurement protocol were preserved. The shared eight-pair preference set and later reviewed span masks were held fixed where applicable. Benchmark text, exact prompts, token identities, tool schemas and raw completions are withheld to protect evaluation integrity.

Changed variable

Twelve isolated span SUM-NLL gradients; nearest Euclidean projection of the raw chosen-SL descent onto non-regression halfspaces; exact fixed-radius BF16 probes.

Measurements

66 pairwise cosines, first-order dots, projection/KKT, realized displacement and teacher-forced likelihood at requested L2 radii 0.000378, 0.000756, 0.001512.

Result

8/66 span pairs had negative cosine, but raw SGD and stored Adam both predicted improvement for all 12 spans. Raw–Adam cosine=0.314496598. Raw was strictly feasible: projection was the identity, retaining 100% descent/norm. Every radius improved 8 spans and worsened 4; tokens up/down=44/36, 41/39, 40/40. Applied BF16 direction cosine=0.508125495/0.640029170/0.757112445. Even realized rounded deltas predicted all 12 spans improve.

Verdict

No projected-SGD continuation justified; no radius satisfied all seven support criteria.

What this ruled out

Aggregate first-order span conflict and Adam are not necessary for this finite-step failure. Projection of these particular halfspaces cannot repair it.

What remained unresolved

Causal shares of numerical forward/backward effects versus nonlinear finite response. No smooth-curvature-only explanation established.

Next experiment

0057: explicit FP32 LoRA arithmetic with its own zero baseline. This is chronology of the recorded research program, not authority to execute new work.

Reproducibility notes

Repeated zero was bit-identical; raw/projected arrays element-identical at each radius. Independent CPU review recomputed Gram, projection, criteria and rounded deltas; no GPU replay or generated behavior. Findings were independently reviewed within the experiment workflow. That review was internal and does not imply external peer review. The public aggregate data and analysis-only renderer support checking the displayed figures. Model artifacts and protected evaluation inputs remain withheld; any future artifact release requires separate license and disclosure review.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

  • Historical 0.1-draft, 2026-10-06: initial source-grounded summary.

  • Accounting note: preserve historical Q1=242, original 0050 matched Q1=244 and corrected canonical Q1=243 as separate labeled records. The 0050 candidate remains 224; the correction changed accounting of identical saved behavior, not model quality. Later diagnostic experiments did not reevaluate the incumbent.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0056 final report and final independent internal review; private manifest supplies exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0057

FP32 LoRA changes outcomes without repairing prediction

0.1 · Public research record

Question
Is BF16 LoRA execution sufficient to explain the local-prediction mismatch?
Changed variable
Explicit autocast-disabled FP32 adapter input, matmuls, scaling and accumulation; same frozen NF4 base, direction, masks and three radii. Compare against each path’s own zero.
Key result
Zero-path target-logprob max/mean absolute difference=0.098438263/0.017496872 nats; 7/80 target-rank changes and 1/80 argmax change. FP32 spans up/down at the three radii=6/6, 9/3, 8/4 versus BF16 8/4 each. Across 36 span probes sign inversions increased 12→13. MAE increased from 0.037919367/0.035397709/0.047795443 to 0.056496019/0.061720997/0.061048685. Both acquisition action spans improved together only at the middle radius.
Verdict
Mixed numerical effects; no actual FP32 LoRA training candidate justified.
What it ruled down
FP32 LoRA alone did not restore overall agreement with these historical local predictions.
Full sanitized summary ↓
Read full record · 0057

Question

Is BF16 LoRA execution sufficient to explain the local-prediction mismatch?

Why it mattered

The protected 0035/Q1 incumbent is 243/300 under the explicit current canonical accounting. This experiment tested a bounded explanation or intervention after a failed continuation. A useful successor must preserve behavior as well as improve the targeted boundary.

Setup

Granite 4.1 8B coordinator with a LoRA adapter; the experiment used the frozen incumbent or its stored artifacts. Results describe the recorded experimental runtime and panels, not universal model behavior.

Fixed inputs

Frozen incumbent identity, source-bound experimental inputs and the relevant evaluator/measurement protocol were preserved. The shared eight-pair preference set and later reviewed span masks were held fixed where applicable. Benchmark text, exact prompts, token identities, tool schemas and raw completions are withheld to protect evaluation integrity.

Changed variable

Explicit autocast-disabled FP32 adapter input, matmuls, scaling and accumulation; same frozen NF4 base, direction, masks and three radii. Compare against each path’s own zero.

Measurements

80 selected-token full vocabulary outputs, ranks, repeat determinism, realized direction, span signs and historical-gradient prediction errors.

Result

Zero-path target-logprob max/mean absolute difference=0.098438263/0.017496872 nats; 7/80 target-rank changes and 1/80 argmax change. FP32 spans up/down at the three radii=6/6, 9/3, 8/4 versus BF16 8/4 each. Across 36 span probes sign inversions increased 12→13. MAE increased from 0.037919367/0.035397709/0.047795443 to 0.056496019/0.061720997/0.061048685. Both acquisition action spans improved together only at the middle radius.

Verdict

Mixed numerical effects; no actual FP32 LoRA training candidate justified.

What this ruled out

FP32 LoRA alone did not restore overall agreement with these historical local predictions.

What remained unresolved

Frozen base/attention/head boundaries and smooth curvature. Historical BF16 gradients are not derivatives recomputed under the FP32 path.

Next experiment

0058: inspect frozen-base/residual numerical assumptions. This is chronology of the recorded research program, not authority to execute new work.

Reproducibility notes

BF16 controls reproduced 0056 exactly; every full FP32 zero/probe repeat was bit-identical on this runtime. FP32 applied direction cosines approached one. Zero-path offsets are not learning gains; no generated behavior/benchmark. Findings were independently reviewed within the experiment workflow. That review was internal and does not imply external peer review. The public aggregate data and analysis-only renderer support checking the displayed figures. Model artifacts and protected evaluation inputs remain withheld; any future artifact release requires separate license and disclosure review.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

  • Historical 0.1-draft, 2026-10-06: initial source-grounded summary.

  • Accounting note: preserve historical Q1=242, original 0050 matched Q1=244 and corrected canonical Q1=243 as separate labeled records. The 0050 candidate remains 224; the correction changed accounting of identical saved behavior, not model quality. Later diagnostic experiments did not reevaluate the incumbent.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0057 final report and final independent internal review; private manifest supplies exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0058

The residual-precision premise was false

0.1 · Public research record

Question
Do frozen-base or residual BF16 discontinuities explain the failed predictions?
Changed variable
Trace historical numerical boundaries; test explicit FP32 residual sums/carrier paths against 0057 FP32 LoRA, with unchanged inputs/direction/radii; bounded block-zero causal-prefix replay.
Key result
All 40 residual carriers and both skip additions were already FP32. Explicit residual paths C/D were bit-identical to FP32-LoRA B at zero and all radii (B=C=D). Counts and prediction errors therefore remained those of 0057. At the smallest radius, block-zero attention-operand pre-cast changes were at most 3.8147e-6 while BF16 post-cast changes reached 0.0078125. Optional high-precision frozen-branch path E was not admitted or executed.
Verdict
BF16 residual-carrier hypothesis excluded for the historical runtime; broader numerical sufficiency and curvature remain unresolved. No candidate justified.
What it ruled down
A nonexistent BF16 residual carrier/addition cannot explain this runtime’s failure. C/D are identities, not a repair.
Full sanitized summary ↓
Read full record · 0058

Question

Do frozen-base or residual BF16 discontinuities explain the failed predictions?

Why it mattered

The protected 0035/Q1 incumbent is 243/300 under the explicit current canonical accounting. This experiment tested a bounded explanation or intervention after a failed continuation. A useful successor must preserve behavior as well as improve the targeted boundary.

Setup

Granite 4.1 8B coordinator with a LoRA adapter; the experiment used the frozen incumbent or its stored artifacts. Results describe the recorded experimental runtime and panels, not universal model behavior.

Fixed inputs

Frozen incumbent identity, source-bound experimental inputs and the relevant evaluator/measurement protocol were preserved. The shared eight-pair preference set and later reviewed span masks were held fixed where applicable. Benchmark text, exact prompts, token identities, tool schemas and raw completions are withheld to protect evaluation integrity.

Changed variable

Trace historical numerical boundaries; test explicit FP32 residual sums/carrier paths against 0057 FP32 LoRA, with unchanged inputs/direction/radii; bounded block-zero causal-prefix replay.

Measurements

Runtime dtype/operation map, path parity, own-zero errors, observed cast thresholds and earliest traced activation divergence.

Result

All 40 residual carriers and both skip additions were already FP32. Explicit residual paths C/D were bit-identical to FP32-LoRA B at zero and all radii (B=C=D). Counts and prediction errors therefore remained those of 0057. At the smallest radius, block-zero attention-operand pre-cast changes were at most 3.8147e-6 while BF16 post-cast changes reached 0.0078125. Optional high-precision frozen-branch path E was not admitted or executed.

Verdict

BF16 residual-carrier hypothesis excluded for the historical runtime; broader numerical sufficiency and curvature remain unresolved. No candidate justified.

What this ruled out

A nonexistent BF16 residual carrier/addition cannot explain this runtime’s failure. C/D are identities, not a repair.

What remained unresolved

Causal sufficiency of attention/frozen-branch/head rounding; value-identical higher-precision NF4 shadow feasibility; internal accumulator bits are not proved by tensor dtype.

Next experiment

0059: matched-kernel counterfactual tests of attention operand threshold sensitivity. This is chronology of the recorded research program, not authority to execute new work.

Reproducibility notes

A reproduced 0056/0057 BF16; B reproduced 0057 FP32; all 16 path/probe settings repeated bit-identically. Independent review verified saved evidence; it did not replay GPU execution. Three radii/observed traces cannot prove continuous smoothness. Findings were independently reviewed within the experiment workflow. That review was internal and does not imply external peer review. The public aggregate data and analysis-only renderer support checking the displayed figures. Model artifacts and protected evaluation inputs remain withheld; any future artifact release requires separate license and disclosure review.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

  • Historical 0.1-draft, 2026-10-06: initial source-grounded summary.

  • Accounting note: preserve historical Q1=242, original 0050 matched Q1=244 and corrected canonical Q1=243 as separate labeled records. The 0050 candidate remains 224; the correction changed accounting of identical saved behavior, not model quality. Later diagnostic experiments did not reevaluate the incumbent.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0058 final report and final independent internal review; private manifest supplies exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0059

Q/K/V rounding changes outcomes without a consistent repair

0.1 · Public research record

Question
Does first-block attention operand rounding explain the finite-step prediction failure?
Changed variable
Selectively remove Q, K, V, or all Q/K/V BF16 rounding in a matched FP32 attention kernel; retain a rounded-operand matched control and historical control.
Key result
Across 36 span/radius observations, historical B had 13 sign inversions, MAE 0.059755234 and RMSE 0.094316727 nats. Removing all operand rounding gave 17 inversions, MAE 0.044248198 and RMSE 0.064969604. No intervention improved both acquisition action spans at every radius. Source: final Experiment 0059 report; final independent source review.
Verdict
Operand rounding is causally active in the matched kernel; it is not a demonstrated primary or sufficient historical explanation. No incumbent improvement or promotion was demonstrated.
What it ruled down
Prediction-error reduction alone did not establish behavioral repair. V dominated crossing counts, but no projection consistently dominated repair. These exclusions apply to the specified tests and support rules.
Full sanitized summary ↓
Read full record · 0059

Question

Does first-block attention operand rounding explain the finite-step prediction failure?

Why it mattered

A bounded continuation had degraded the protected coordinator. Explaining its response was necessary before treating another update as a justified successor.

Setup

Frozen Granite 4.1 8B LoRA coordinator 0035/Q1; protected score 243/300 under current canonical accounting. This experiment did not reevaluate that benchmark score or qualify a successor.

Fixed inputs

12 chosen supervision spans, 80 selected teacher-forced tokens from six contributing pairs; frozen model/reference/quantization state and recorded numerical execution controls. Exact benchmark text and token records remain private.

Changed variable

Selectively remove Q, K, V, or all Q/K/V BF16 rounding in a matched FP32 attention kernel; retain a rounded-operand matched control and historical control.

Measurements

Own-zero chosen-span negative log likelihood (NLL), selected-token log probabilities and saved exact-repeat records; directional predictions or contextual response measurements as applicable. Negative NLL delta means improvement. Sign inversions oppose a predicted improvement. These are teacher-forced proxies, not generated action or operational safety qualification.

Result

Across 36 span/radius observations, historical B had 13 sign inversions, MAE 0.059755234 and RMSE 0.094316727 nats. Removing all operand rounding gave 17 inversions, MAE 0.044248198 and RMSE 0.064969604. No intervention improved both acquisition action spans at every radius. Source: final Experiment 0059 report; final independent source review.

Verdict

Operand rounding is causally active in the matched kernel; it is not a demonstrated primary or sufficient historical explanation. No incumbent improvement or promotion was demonstrated.

What this ruled out

Prediction-error reduction alone did not establish behavioral repair. V dominated crossing counts, but no projection consistently dominated repair. These exclusions apply to the specified tests and support rules.

What remained unresolved

A particular threshold crossing was not isolated; changing kernels and zero-step baselines limits historical-runtime attribution. Exact repeats establish determinism on the recorded host/runtime only.

Next experiment

0060: hold represented operands fixed and vary internal attention-stage precision. Historical sequence and recommendation only; this publication lane authorizes no model execution.

Reproducibility notes

Source evidence retains frozen identities, fixed intervention/probe definitions, saved numerical outputs, restoration checks, verification receipts and independent internal review. The released sanitized aggregate data and analysis-only renderer support figure reproduction; protected benchmark inputs prevent full public replay as written.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

Historical 0.1-draft (2026-10-05 research event; drafted 2026-10-06): source-grounded sanitized summary. Publication date was unset in that draft; revision dates recorded actual draft edits. Historical accounting remains separately documented; this entry does not recalculate earlier scores.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0059 final report and final independent internal review; private provenance retains exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0060

Attention-stage precision is active but insufficient

0.1 · Public research record

Question
Does internal attention arithmetic explain the contradiction when operand values and output rounding are fixed?
Changed variable
Use one explicit staged attention implementation; promote QK, scaling, softmax, then value aggregation to FP32 while retaining BF16-rounded Q/K/V and BF16 attention output.
Key result
Pooled sign inversions out of 36: historical B 13; explicit materialized BF16 control 20; progressive stages S1/S2/S3/S4 16/16/16/12. S4 MAE 0.059894721 and RMSE 0.093219598 nats remained near historical B. S4 never improved both acquisition action spans. Source: final Experiment 0060 report; final independent source review.
Verdict
Stage precision is causally active within the explicit implementation but did not provide a material sufficient historical explanation. No incumbent improvement or promotion was demonstrated.
What it ruled down
More precision did not produce monotonic improvement. Materialized stage dtypes do not reveal hidden native accumulator formats. These exclusions apply to the specified tests and support rules.
Full sanitized summary ↓
Read full record · 0060

Question

Does internal attention arithmetic explain the contradiction when operand values and output rounding are fixed?

Why it mattered

A bounded continuation had degraded the protected coordinator. Explaining its response was necessary before treating another update as a justified successor.

Setup

Frozen Granite 4.1 8B LoRA coordinator 0035/Q1; protected score 243/300 under current canonical accounting. This experiment did not reevaluate that benchmark score or qualify a successor.

Fixed inputs

12 chosen supervision spans, 80 selected teacher-forced tokens from six contributing pairs; frozen model/reference/quantization state and recorded numerical execution controls. Exact benchmark text and token records remain private.

Changed variable

Use one explicit staged attention implementation; promote QK, scaling, softmax, then value aggregation to FP32 while retaining BF16-rounded Q/K/V and BF16 attention output.

Measurements

Own-zero chosen-span negative log likelihood (NLL), selected-token log probabilities and saved exact-repeat records; directional predictions or contextual response measurements as applicable. Negative NLL delta means improvement. Sign inversions oppose a predicted improvement. These are teacher-forced proxies, not generated action or operational safety qualification.

Result

Pooled sign inversions out of 36: historical B 13; explicit materialized BF16 control 20; progressive stages S1/S2/S3/S4 16/16/16/12. S4 MAE 0.059894721 and RMSE 0.093219598 nats remained near historical B. S4 never improved both acquisition action spans. Source: final Experiment 0060 report; final independent source review.

Verdict

Stage precision is causally active within the explicit implementation but did not provide a material sufficient historical explanation. No incumbent improvement or promotion was demonstrated.

What this ruled out

More precision did not produce monotonic improvement. Materialized stage dtypes do not reveal hidden native accumulator formats. These exclusions apply to the specified tests and support rules.

What remained unresolved

Historical fused internals remain opaque; explicit-versus-native differences change the zero function. Exact repeats establish determinism on the recorded host/runtime only.

Next experiment

0061: remove the attention output BF16 boundary under matched downstream projection. Historical sequence and recommendation only; this publication lane authorizes no model execution.

Reproducibility notes

Source evidence retains frozen identities, fixed intervention/probe definitions, saved numerical outputs, restoration checks, verification receipts and independent internal review. The released sanitized aggregate data and analysis-only renderer support figure reproduction; protected benchmark inputs prevent full public replay as written.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

Historical 0.1-draft (2026-10-05 research event; drafted 2026-10-06): source-grounded sanitized summary. Publication date was unset in that draft; revision dates recorded actual draft edits. Historical accounting remains separately documented; this entry does not recalculate earlier scores.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0060 final report and final independent internal review; private provenance retains exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0061

Attention-output rounding does not consistently repair response

0.1 · Public research record

Question
Does removing the attention-output BF16 boundary restore finite-step prediction?
Changed variable
Compare unrounded FP32 aggregation (O1) with BF16-rounded then FP32-expanded aggregation (O2) under the same FP32 output-projection arithmetic and BF16 return boundary.
Key result
Across 36 observations, unrounded O1 had 13 inversions, MAE 0.0568731183 and RMSE 0.0879775282 nats; rounded matched O2 had 11, 0.0491626251 and 0.0652904364. Unrounded O1 improved both acquisition action spans only at the largest of three fixed radii. Source: final Experiment 0061 report; final independent source review.
Verdict
Output rounding is a bounded causal contributor, rejected as the primary or sufficient explanation under these tests. No incumbent improvement or promotion was demonstrated.
What it ruled down
The matched unrounded path was worse overall despite some individual repairs. These exclusions apply to the specified tests and support rules.
Full sanitized summary ↓
Read full record · 0061

Question

Does removing the attention-output BF16 boundary restore finite-step prediction?

Why it mattered

A bounded continuation had degraded the protected coordinator. Explaining its response was necessary before treating another update as a justified successor.

Setup

Frozen Granite 4.1 8B LoRA coordinator 0035/Q1; protected score 243/300 under current canonical accounting. This experiment did not reevaluate that benchmark score or qualify a successor.

Fixed inputs

12 chosen supervision spans, 80 selected teacher-forced tokens from six contributing pairs; frozen model/reference/quantization state and recorded numerical execution controls. Exact benchmark text and token records remain private.

Changed variable

Compare unrounded FP32 aggregation (O1) with BF16-rounded then FP32-expanded aggregation (O2) under the same FP32 output-projection arithmetic and BF16 return boundary.

Measurements

Own-zero chosen-span negative log likelihood (NLL), selected-token log probabilities and saved exact-repeat records; directional predictions or contextual response measurements as applicable. Negative NLL delta means improvement. Sign inversions oppose a predicted improvement. These are teacher-forced proxies, not generated action or operational safety qualification.

Result

Across 36 observations, unrounded O1 had 13 inversions, MAE 0.0568731183 and RMSE 0.0879775282 nats; rounded matched O2 had 11, 0.0491626251 and 0.0652904364. Unrounded O1 improved both acquisition action spans only at the largest of three fixed radii. Source: final Experiment 0061 report; final independent source review.

Verdict

Output rounding is a bounded causal contributor, rejected as the primary or sufficient explanation under these tests. No incumbent improvement or promotion was demonstrated.

What this ruled out

The matched unrounded path was worse overall despite some individual repairs. These exclusions apply to the specified tests and support rules.

What remained unresolved

Whole-boundary effects do not establish necessity or sufficiency of individual crossings; native versus matched projection semantics are distinct. Exact repeats establish determinism on the recorded host/runtime only.

Next experiment

0062: vary output-projection BF16 return with unrounded attention input and matched projection arithmetic. Historical sequence and recommendation only; this publication lane authorizes no model execution.

Reproducibility notes

Source evidence retains frozen identities, fixed intervention/probe definitions, saved numerical outputs, restoration checks, verification receipts and independent internal review. The released sanitized aggregate data and analysis-only renderer support figure reproduction; protected benchmark inputs prevent full public replay as written.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

Historical 0.1-draft (2026-10-05 research event; drafted 2026-10-06): source-grounded sanitized summary. Publication date was unset in that draft; revision dates recorded actual draft edits. Historical accounting remains separately documented; this entry does not recalculate earlier scores.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0061 final report and final independent internal review; private provenance retains exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0062

Projection-output rounding suppresses a gain but does not explain failure

0.1 · Public research record

Question
Does the first-block output-projection BF16 return explain remaining prediction failures?
Changed variable
Compare rounded/reexpanded projection output P0 with unrounded FP32 output P1, sharing FP32 residual multiplication and addition.
Key result
Both P0 and P1 had 13 pooled sign inversions out of 36. P0 MAE/RMSE were 0.0382807534/0.0698396891 nats; P1 were 0.0898191555/0.161609083. P1 improved both acquisition action spans at all three radii, while worsening prediction error at each radius. Source: final Experiment 0062 report; final independent source review.
Verdict
Projection-output rounding is a reproducible contributor but not a supported primary or sufficient historical explanation. No incumbent improvement or promotion was demonstrated.
What it ruled down
Preserving action-span gains did not resolve losses elsewhere or repair prediction agreement. These exclusions apply to the specified tests and support rules.
Full sanitized summary ↓
Read full record · 0062

Question

Does the first-block output-projection BF16 return explain remaining prediction failures?

Why it mattered

A bounded continuation had degraded the protected coordinator. Explaining its response was necessary before treating another update as a justified successor.

Setup

Frozen Granite 4.1 8B LoRA coordinator 0035/Q1; protected score 243/300 under current canonical accounting. This experiment did not reevaluate that benchmark score or qualify a successor.

Fixed inputs

12 chosen supervision spans, 80 selected teacher-forced tokens from six contributing pairs; frozen model/reference/quantization state and recorded numerical execution controls. Exact benchmark text and token records remain private.

Changed variable

Compare rounded/reexpanded projection output P0 with unrounded FP32 output P1, sharing FP32 residual multiplication and addition.

Measurements

Own-zero chosen-span negative log likelihood (NLL), selected-token log probabilities and saved exact-repeat records; directional predictions or contextual response measurements as applicable. Negative NLL delta means improvement. Sign inversions oppose a predicted improvement. These are teacher-forced proxies, not generated action or operational safety qualification.

Result

Both P0 and P1 had 13 pooled sign inversions out of 36. P0 MAE/RMSE were 0.0382807534/0.0698396891 nats; P1 were 0.0898191555/0.161609083. P1 improved both acquisition action spans at all three radii, while worsening prediction error at each radius. Source: final Experiment 0062 report; final independent source review.

Verdict

Projection-output rounding is a reproducible contributor but not a supported primary or sufficient historical explanation. No incumbent improvement or promotion was demonstrated.

What this ruled out

Preserving action-span gains did not resolve losses elsewhere or repair prediction agreement. These exclusions apply to the specified tests and support rules.

What remained unresolved

P0 residual arithmetic differs from the historical BF16-return path; only P1 versus P0 is the matched causal contrast. Later block boundaries were not isolated. Exact repeats establish determinism on the recorded host/runtime only.

Next experiment

0063: recompute path-local autodiff and compare direct symmetric finite differences. Historical sequence and recommendation only; this publication lane authorizes no model execution.

Reproducibility notes

Source evidence retains frozen identities, fixed intervention/probe definitions, saved numerical outputs, restoration checks, verification receipts and independent internal review. The released sanitized aggregate data and analysis-only renderer support figure reproduction; protected benchmark inputs prevent full public replay as written.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

Historical 0.1-draft (2026-10-05 research event; drafted 2026-10-06): source-grounded sanitized summary. Publication date was unset in that draft; revision dates recorded actual draft edits. Historical accounting remains separately documented; this entry does not recalculate earlier scores.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0062 final report and final independent internal review; private provenance retains exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0063

Path-local autodiff still fails direct finite-difference checks

0.1 · Public research record

Question
Is historical-gradient reuse the reason finite responses disagree with predicted improvements?
Changed variable
Recompute span reverse-mode gradients at exact P1 zero and audit four preregistered symmetric perturbation sizes along the unchanged direction.
Key result
All 12 P1-local autodiff projections predicted improvement. Fixed-radius outcomes improved 7/12, 7/12 and 9/12 spans, leaving 5/5/3 sign inversions. Direct finite-difference sign agreement was 5/12, 6/12, 3/12 and 7/12 over the four epsilon sizes; 11 spans changed finite-difference sign. Source: final Experiment 0063 report; final independent source review.
Verdict
Strong same-path disagreement remained; historical-gradient reuse did not explain it on the audited grid. No incumbent improvement or promotion was demonstrated.
What it ruled down
Fresh path-local gradients did not repair executable prediction. Smaller probes did not converge to an autodiff-confirming estimate. These exclusions apply to the specified tests and support rules.
Full sanitized summary ↓
Read full record · 0063

Question

Is historical-gradient reuse the reason finite responses disagree with predicted improvements?

Why it mattered

A bounded continuation had degraded the protected coordinator. Explaining its response was necessary before treating another update as a justified successor.

Setup

Frozen Granite 4.1 8B LoRA coordinator 0035/Q1; protected score 243/300 under current canonical accounting. This experiment did not reevaluate that benchmark score or qualify a successor.

Fixed inputs

12 chosen supervision spans, 80 selected teacher-forced tokens from six contributing pairs; frozen model/reference/quantization state and recorded numerical execution controls. Exact benchmark text and token records remain private.

Changed variable

Recompute span reverse-mode gradients at exact P1 zero and audit four preregistered symmetric perturbation sizes along the unchanged direction.

Measurements

Own-zero chosen-span negative log likelihood (NLL), selected-token log probabilities and saved exact-repeat records; directional predictions or contextual response measurements as applicable. Negative NLL delta means improvement. Sign inversions oppose a predicted improvement. These are teacher-forced proxies, not generated action or operational safety qualification.

Result

All 12 P1-local autodiff projections predicted improvement. Fixed-radius outcomes improved 7/12, 7/12 and 9/12 spans, leaving 5/5/3 sign inversions. Direct finite-difference sign agreement was 5/12, 6/12, 3/12 and 7/12 over the four epsilon sizes; 11 spans changed finite-difference sign. Source: final Experiment 0063 report; final independent source review.

Verdict

Strong same-path disagreement remained; historical-gradient reuse did not explain it on the audited grid. No incumbent improvement or promotion was demonstrated.

What this ruled out

Fresh path-local gradients did not repair executable prediction. Smaller probes did not converge to an autodiff-confirming estimate. These exclusions apply to the specified tests and support rules.

What remained unresolved

Four finite epsilon values cannot prove a derivative limit exists or fails to exist; casts use reverse-mode surrogates and FP32 realizations change at tiny scales. Exact repeats establish determinism on the recorded host/runtime only.

Next experiment

0064: audit a fixed scalar response-scale ladder without changing direction or numerical path. Historical sequence and recommendation only; this publication lane authorizes no model execution.

Reproducibility notes

Source evidence retains frozen identities, fixed intervention/probe definitions, saved numerical outputs, restoration checks, verification receipts and independent internal review. The released sanitized aggregate data and analysis-only renderer support figure reproduction; protected benchmark inputs prevent full public replay as written.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

Historical 0.1-draft (2026-10-05 research event; drafted 2026-10-06): source-grounded sanitized summary. Publication date was unset in that draft; revision dates recorded actual draft edits. Historical accounting remains separately documented; this entry does not recalculate earlier scores.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0063 final report and final independent internal review; private provenance retains exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0064

No usable local response regime on the tested ladder

0.1 · Public research record

Question
Does a fixed signed step-size ladder reveal a stable numerical floor or local autodiff regime?
Changed variable
Measure both signs on 11 fixed displacement magnitudes from R1/256 through 4R1 on P1, retaining the original direction and conversion.
Key result
All 12 spans and all 80 tokens responded on both signs at the minimum tested radius 1.4765625e-6. At its positive sign, absolute span NLL change min/median/max was 0.000118255615/0.0377326012/0.45158577 nats. Central sign agreement ranged from 4/12 to 10/12. No contiguous region passed the preregistered diagnostic criteria. Source: final Experiment 0064 report; final independent source review.
Verdict
Deterministic irregular and asymmetric finite response; no recovered usable autodiff regime within this ladder. No incumbent improvement or promotion was demonstrated.
What it ruled down
An initial near-zero plateau was not observed. The onset lies at or below the minimum tested scale, rather than at an identified threshold. These exclusions apply to the specified tests and support rules.
Full sanitized summary ↓
Read full record · 0064

Question

Does a fixed signed step-size ladder reveal a stable numerical floor or local autodiff regime?

Why it mattered

A bounded continuation had degraded the protected coordinator. Explaining its response was necessary before treating another update as a justified successor.

Setup

Frozen Granite 4.1 8B LoRA coordinator 0035/Q1; protected score 243/300 under current canonical accounting. This experiment did not reevaluate that benchmark score or qualify a successor.

Fixed inputs

12 chosen supervision spans, 80 selected teacher-forced tokens from six contributing pairs; frozen model/reference/quantization state and recorded numerical execution controls. Exact benchmark text and token records remain private.

Changed variable

Measure both signs on 11 fixed displacement magnitudes from R1/256 through 4R1 on P1, retaining the original direction and conversion.

Measurements

Own-zero chosen-span negative log likelihood (NLL), selected-token log probabilities and saved exact-repeat records; directional predictions or contextual response measurements as applicable. Negative NLL delta means improvement. Sign inversions oppose a predicted improvement. These are teacher-forced proxies, not generated action or operational safety qualification.

Result

All 12 spans and all 80 tokens responded on both signs at the minimum tested radius 1.4765625e-6. At its positive sign, absolute span NLL change min/median/max was 0.000118255615/0.0377326012/0.45158577 nats. Central sign agreement ranged from 4/12 to 10/12. No contiguous region passed the preregistered diagnostic criteria. Source: final Experiment 0064 report; final independent source review.

Verdict

Deterministic irregular and asymmetric finite response; no recovered usable autodiff regime within this ladder. No incumbent improvement or promotion was demonstrated.

What this ruled out

An initial near-zero plateau was not observed. The onset lies at or below the minimum tested scale, rather than at an identified threshold. These exclusions apply to the specified tests and support rules.

What remained unresolved

Finite samples cannot establish absence of all local regimes or locate a true smaller-scale onset floor; descriptive thresholds are not production gates. Exact repeats establish determinism on the recorded host/runtime only.

Next experiment

0065: construct a direction from measured symmetric finite responses under fixed P1. Historical sequence and recommendation only; this publication lane authorizes no model execution.

Reproducibility notes

Source evidence retains frozen identities, fixed intervention/probe definitions, saved numerical outputs, restoration checks, verification receipts and independent internal review. The released sanitized aggregate data and analysis-only renderer support figure reproduction; protected benchmark inputs prevent full public replay as written.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

Historical 0.1-draft (2026-10-05 research event; drafted 2026-10-06): source-grounded sanitized summary. Publication date was unset in that draft; revision dates recorded actual draft edits. Historical accounting remains separately documented; this entry does not recalculate earlier scores.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0064 final report and final independent internal review; private provenance retains exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0065

A measured finite-response direction fails blind support gates

0.1 · Public research record

Question
Can an empirical finite-response model construct a direction that transfers to actual finite evaluations?
Changed variable
Build a rank-12 orthonormal span-gradient subspace, measure symmetric basis responses at three fixed scales, freeze median-J minimax coefficients, then evaluate one direction at two validation radii.
Key result
136/144 response-model cells met the descriptive instability rule; 42/144 had consistent nonzero signs. The measured model predicted nonworsening for all 12 spans. Actual positive-direction validation improved 7/12 at radius 0.000189 and 6/12 at 0.000378, with aggregate NLL deltas -0.3494768143 and -0.1444549561 nats. Neither passed every mandatory support gate. Source: final Experiment 0065 report; final independent source review.
Verdict
No supported candidate under this frozen measured-median/minimax construction. No incumbent improvement or promotion was demonstrated.
What it ruled down
A feasible empirical linear model did not transfer to a clean finite candidate. The failed direction does not establish that the entire subspace is unusable. These exclusions apply to the specified tests and support rules.
Full sanitized summary ↓
Read full record · 0065

Question

Can an empirical finite-response model construct a direction that transfers to actual finite evaluations?

Why it mattered

A bounded continuation had degraded the protected coordinator. Explaining its response was necessary before treating another update as a justified successor.

Setup

Frozen Granite 4.1 8B LoRA coordinator 0035/Q1; protected score 243/300 under current canonical accounting. This experiment did not reevaluate that benchmark score or qualify a successor.

Fixed inputs

12 chosen supervision spans, 80 selected teacher-forced tokens from six contributing pairs; frozen model/reference/quantization state and recorded numerical execution controls. Exact benchmark text and token records remain private.

Changed variable

Build a rank-12 orthonormal span-gradient subspace, measure symmetric basis responses at three fixed scales, freeze median-J minimax coefficients, then evaluate one direction at two validation radii.

Measurements

Own-zero chosen-span negative log likelihood (NLL), selected-token log probabilities and saved exact-repeat records; directional predictions or contextual response measurements as applicable. Negative NLL delta means improvement. Sign inversions oppose a predicted improvement. These are teacher-forced proxies, not generated action or operational safety qualification.

Result

136/144 response-model cells met the descriptive instability rule; 42/144 had consistent nonzero signs. The measured model predicted nonworsening for all 12 spans. Actual positive-direction validation improved 7/12 at radius 0.000189 and 6/12 at 0.000378, with aggregate NLL deltas -0.3494768143 and -0.1444549561 nats. Neither passed every mandatory support gate. Source: final Experiment 0065 report; final independent source review.

Verdict

No supported candidate under this frozen measured-median/minimax construction. No incumbent improvement or promotion was demonstrated.

What this ruled out

A feasible empirical linear model did not transfer to a clean finite candidate. The failed direction does not establish that the entire subspace is unusable. These exclusions apply to the specified tests and support rules.

What remained unresolved

Heldout mechanism-panel admission was denied; no generated-behavior generalization was established. Exact repeats establish determinism on the recorded host/runtime only.

Next experiment

0066: measure pairwise superposition error and test a fixed quadratic explanation. Historical sequence and recommendation only; this publication lane authorizes no model execution.

Reproducibility notes

Source evidence retains frozen identities, fixed intervention/probe definitions, saved numerical outputs, restoration checks, verification receipts and independent internal review. The released sanitized aggregate data and analysis-only renderer support figure reproduction; protected benchmark inputs prevent full public replay as written.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

Historical 0.1-draft (2026-10-05 research event; drafted 2026-10-06): source-grounded sanitized summary. Publication date was unset in that draft; revision dates recorded actual draft edits. Historical accounting remains separately documented; this entry does not recalculate earlier scores.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0065 final report and final independent internal review; private provenance retains exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0066

Large pairwise interactions do not yield a sufficient quadratic explanation

0.1 · Public research record

Question
Does measured pairwise non-additivity explain the frozen empirical direction failure?
Changed variable
Measure all 66 basis pairs at three amplitudes and four sign combinations, twice each; compare exact pair responses to matching singleton sums, then evaluate a fixed-J quadratic model.
Key result
792 pair settings yielded 9,504 span observations. Normalized interaction scores were 0.676486/0.663896/0.650674 across amplitudes; span sign-reversal rates were 0.256944/0.255997/0.243056. Quadratic versus linear sign agreement was 6 versus 7 of 12 at V1, and 7 versus 6 at V2. At V2, quadratic MAE/RMSE 0.24492/0.321043 exceeded linear 0.0997909/0.149775 nats. Source: final Experiment 0066 report; final independent source review.
Verdict
Pairwise non-additivity is material, but the measured quadratic model does not explain the frozen direction failure. No incumbent improvement or promotion was demonstrated.
What it ruled down
Strong interactions alone did not provide an adequate predictive model. Quadratic even terms cannot fix every observed positive-versus-negative ordering with the fixed linear term. These exclusions apply to the specified tests and support rules.
Full sanitized summary ↓
Read full record · 0066

Question

Does measured pairwise non-additivity explain the frozen empirical direction failure?

Why it mattered

A bounded continuation had degraded the protected coordinator. Explaining its response was necessary before treating another update as a justified successor.

Setup

Frozen Granite 4.1 8B LoRA coordinator 0035/Q1; protected score 243/300 under current canonical accounting. This experiment did not reevaluate that benchmark score or qualify a successor.

Fixed inputs

12 chosen supervision spans, 80 selected teacher-forced tokens from six contributing pairs; frozen model/reference/quantization state and recorded numerical execution controls. Exact benchmark text and token records remain private.

Changed variable

Measure all 66 basis pairs at three amplitudes and four sign combinations, twice each; compare exact pair responses to matching singleton sums, then evaluate a fixed-J quadratic model.

Measurements

Own-zero chosen-span negative log likelihood (NLL), selected-token log probabilities and saved exact-repeat records; directional predictions or contextual response measurements as applicable. Negative NLL delta means improvement. Sign inversions oppose a predicted improvement. These are teacher-forced proxies, not generated action or operational safety qualification.

Result

792 pair settings yielded 9,504 span observations. Normalized interaction scores were 0.676486/0.663896/0.650674 across amplitudes; span sign-reversal rates were 0.256944/0.255997/0.243056. Quadratic versus linear sign agreement was 6 versus 7 of 12 at V1, and 7 versus 6 at V2. At V2, quadratic MAE/RMSE 0.24492/0.321043 exceeded linear 0.0997909/0.149775 nats. Source: final Experiment 0066 report; final independent source review.

Verdict

Pairwise non-additivity is material, but the measured quadratic model does not explain the frozen direction failure. No incumbent improvement or promotion was demonstrated.

What this ruled out

Strong interactions alone did not provide an adequate predictive model. Quadratic even terms cannot fix every observed positive-versus-negative ordering with the fixed linear term. These exclusions apply to the specified tests and support rules.

What remained unresolved

Known 0065 outcomes were held out from fitting/model selection, not prospectively unknown. Larger radii and 12-coordinate combinations exceed measured pair support. Exact repeats establish determinism on the recorded host/runtime only.

Next experiment

0067: matched-coefficient leave-one-component-out analysis of the frozen direction. Historical sequence and recommendation only; this publication lane authorizes no model execution.

Reproducibility notes

Source evidence retains frozen identities, fixed intervention/probe definitions, saved numerical outputs, restoration checks, verification receipts and independent internal review. The released sanitized aggregate data and analysis-only renderer support figure reproduction; protected benchmark inputs prevent full public replay as written.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

Historical 0.1-draft (2026-10-05 research event; drafted 2026-10-06): source-grounded sanitized summary. Publication date was unset in that draft; revision dates recorded actual draft edits. Historical accounting remains separately documented; this entry does not recalculate earlier scores.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0066 final report and final independent internal review; private provenance retains exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.

0067

Matched component removals reveal distributed tradeoffs

0.1 · Public research record

Question
Can a single matched-magnitude component removal safely repair the frozen direction at both radii?
Changed variable
Remove each stored coefficient times its basis vector from the frozen direction, globally normalize, and evaluate at the two fixed radii without refitting.
Key result
Across 144 component/span comparisons, 70 signs were stable and 74 reversed between radii. All 12 removals had mixed span effects; every removal failed support guards. No pair was eligible for the optional double ablation, so none was run. Source: final Experiment 0067 report; final independent source review.
Verdict
No supported update; no single-component localization is supported, and no tested removal safely repairs the full failure at both radii. No incumbent improvement or promotion was demonstrated.
What it ruled down
Helpful effects on some spans coexisted with critical damage elsewhere and strong radius dependence. These exclusions apply to the specified tests and support rules.
Full sanitized summary ↓
Read full record · 0067

Question

Can a single matched-magnitude component removal safely repair the frozen direction at both radii?

Why it mattered

A bounded continuation had degraded the protected coordinator. Explaining its response was necessary before treating another update as a justified successor.

Setup

Frozen Granite 4.1 8B LoRA coordinator 0035/Q1; protected score 243/300 under current canonical accounting. This experiment did not reevaluate that benchmark score or qualify a successor.

Fixed inputs

12 chosen supervision spans, 80 selected teacher-forced tokens from six contributing pairs; frozen model/reference/quantization state and recorded numerical execution controls. Exact benchmark text and token records remain private.

Changed variable

Remove each stored coefficient times its basis vector from the frozen direction, globally normalize, and evaluate at the two fixed radii without refitting.

Measurements

Own-zero chosen-span negative log likelihood (NLL), selected-token log probabilities and saved exact-repeat records; directional predictions or contextual response measurements as applicable. Negative NLL delta means improvement. Sign inversions oppose a predicted improvement. These are teacher-forced proxies, not generated action or operational safety qualification.

Result

Across 144 component/span comparisons, 70 signs were stable and 74 reversed between radii. All 12 removals had mixed span effects; every removal failed support guards. No pair was eligible for the optional double ablation, so none was run. Source: final Experiment 0067 report; final independent source review.

Verdict

No supported update; no single-component localization is supported, and no tested removal safely repairs the full failure at both radii. No incumbent improvement or promotion was demonstrated.

What this ruled out

Helpful effects on some spans coexisted with critical damage elsewhere and strong radius dependence. These exclusions apply to the specified tests and support rules.

What remained unresolved

Global normalization changes every remaining amplitude. These interventions do not prove that no single component can explain the failure or identify exact interaction order. Exact repeats establish determinism on the recorded host/runtime only.

Next experiment

Unexecuted recommendation: bounded rounding-boundary tracing along frozen and matched removed-component paths; requires separate research authorization. Historical sequence and recommendation only; this publication lane authorizes no model execution.

Reproducibility notes

Source evidence retains frozen identities, fixed intervention/probe definitions, saved numerical outputs, restoration checks, verification receipts and independent internal review. The released sanitized aggregate data and analysis-only renderer support figure reproduction; protected benchmark inputs prevent full public replay as written.

Public artifacts

This public research record includes the sanitized summary and the source-grounded aggregate figures/data used by the flagship article. Figure inputs, the analysis-only renderer and plotting dependencies are available as aggregate data, renderer and requirements. VEZRYN LLC retains all rights to its authored material; third-party components retain their own licenses. Public accessibility does not change the experimental verdict or imply model promotion. Raw reports, prompts, benchmark case identifiers, sensitive traces, model artifacts and private provenance are withheld.

Corrections / version history

Historical 0.1-draft (2026-10-05 research event; drafted 2026-10-06): source-grounded sanitized summary. Publication date was unset in that draft; revision dates recorded actual draft edits. Historical accounting remains separately documented; this entry does not recalculate earlier scores.

  • 0.1, 2026-10-06: first public release of the sanitized record; experimental findings and accounting unchanged.

Source attribution: VEZRYN LLC, experiment 0067 final report and final independent internal review; private provenance retains exact bindings. Model, method and frozen-run software attribution appear in the flagship reproducibility section.