A stable Granite 4.1 8B coordinator, adapted with LoRA, scored 243/300 under current canonical accounting. One deliberately bounded continuation step produced a candidate scoring 224/300. The protected incumbent remained intact. No successor in the research record covered here cleanly beat it. 0050, 0051

The ensuing experiments tested smaller displacement, different preference objectives, positive-only supervision, local gradient constraints, numerical execution boundaries, local derivatives, measured finite-response models and pair interactions. They narrowed several plausible explanations. They did not yield a sufficient mechanism or a successful replacement.

This article reports that negative result. “Stable” refers to repeatability within the recorded evidence, not flawless behavior or production qualification. “Broke” refers to degradation of the tested continuation, not destruction of the retained adapter or a general defect in Granite. The observations below concern one lineage, specific panels and the recorded runtime.

1. Why we were modifying the coordinator

The coordinator is a bounded model for tool and delegation decisions. Improving such a model means more than increasing the likelihood of a preferred response. It must acquire evidence when needed, remain grounded when evidence is missing, preserve valid action structure and retain existing behavior. A successor that repairs one decision boundary while damaging another is not an improvement by those requirements.

The initial question was narrow: could a small preference-based continuation improve evidence acquisition without sacrificing the incumbent? Direct Preference Optimization (DPO) offered a way to contrast chosen and rejected responses against a fixed reference policy.1 LoRA provided the adapter parameterization.2 The underlying model was IBM Granite 4.1 8B.3 The experiments used a quantized numerical runtime; NF4 context comes from QLoRA, without implying adoption of every recipe in that paper.4

A useful reading distinction is score, generated behavior and teacher-forced response. The benchmark aggregates evaluated behavior. Later diagnostics often measured selected-token or selected-span likelihood with the target response supplied. Those diagnostics can investigate an update without qualifying its generated actions. This distinction becomes essential as the chronology moves from rejected candidates to mechanism tests.

Figure 1A sequence of bounded negative results
Eighteen experiments progress from continuation and objective tests to numerical-path, response-model and interaction diagnostics; no successor qualified.

0050–0067: final-record chronology and bounded verdicts.

Dates identify final experiment records, not execution duration. Diagnostic results are not benchmark scores; own-zero conventions apply only to response diagnostics.

Sources: 0050, 0051, 0052, 0053, 0054, 0055, 0056, 0057, 0058, 0059, 0060, 0061, 0062, 0063, 0064, 0065, 0066, 0067 · Reproducible aggregate data (JSON)

2. The stable incumbent

The protected adapter was 0035/Q1. Its current canonical score is 243/300. Existing failures remained; stability did not mean perfection. The accounting history must remain visible because identical saved behavior had previously received different totals. 0050, 0051

Accounting recordIncumbent Q10050 candidateMeaning
Historical incumbent record242/300—Preserved historical total
Original matched 0050 comparison244/300224/300Original reported comparison
Current canonical correction in 0051243/300224/300Corrected comparison: −19 points

All thirty baseline behavior hashes were unchanged across the historical/current comparison. The correction was an accounting correction on saved behavior, not a gain from training. Five equal numeric replays per model support deterministic score reconstruction; they are not five fresh semantic judgments. 0050, 0051

This boundary matters throughout the article. Later experiments did not repeatedly rescore Q1. Many did not run the benchmark at all. A chronology of diagnostics cannot honestly become a curve of candidate benchmark scores.

Figure 2Accounting changes are not model gains
Two comparisons show the same failed candidate at 224 of 300; incumbent 244 under original matched accounting and 243 under canonical accounting. Earlier historical accounting was 242.

0050 original comparison: 244→224. 0051 canonical correction: 243→224.

The incumbent behavior is unchanged by accounting. Historical 242 is preserved separately, not plotted as a learning trajectory. Own-zero does not apply to these benchmark totals.

Sources: 0050, 0051 · Reproducible aggregate data (JSON)

3. The one-step failure

Experiment 0050 applied one full-response DPO update to eight fixed preferences. Its settings included learning rate 10−610^{-6}, DPO coefficient β=0.1\beta=0.1, microbatch size one, accumulation eight, fresh paged AdamW8bit and clipping at 0.5. The candidate changed fourteen of thirty benchmark behaviors. Benchmark format failures increased from zero to one. Under canonical accounting, the score fell from 243 to 224. 0050, 0051

The separate eighteen-case synthetic heldout showed a narrow benefit and consequential regressions:

Heldout observationQ1Full-response continuation
Implicit evidence acquisition1/32/3
Fabrication in missing-evidence cases3/63/6
Grammar failures1/183/18

Source: 0050. These are distinct heldout counts, not benchmark category scores.

The acquisition improvement did not compensate for failed preservation. The continuation was rejected. This was the first durable result: a deliberately bounded update could produce substantial functional damage in this setup, and a targeted gain was insufficient to justify promotion.

The adapter’s L2 displacement was 0.006050562709, about 0.015553% relative to its parameter norm. All 560 tensors moved. Those measurements describe geometry, not behavioral locality. A small norm is an observation; “therefore safe” would have been an unsupported inference. 0050, 0051

4. Step magnitude investigation

The intuitive repair was to take less of the failed step. Experiment 0051 first analyzed stored artifacts and corrected accounting without new model execution. It also found that the one-step displacement was a substantial fraction of two earlier seven-update endpoint chords: 38.7062% and 41.7325%. A low learning-rate label did not by itself characterize the effective movement. 0051

Experiment 0052 tested four fixed LoRA-factor interpolation fractions: 0.125, 0.250, 0.375 and 0.500. Every candidate had two format failures out of eighteen heldout cases, against one for Q1. Every candidate failed the zero-new-format gate. No benchmark score was measured. 0052

Figure 3Smaller factor-interpolated updates still failed preservation
All four sampled fractions have two format failures; the incumbent has one. Fabrication is six at the first three fractions and seven at the fourth.

0052: four LoRA-factor interpolation fractions retain 2/18 format failures versus incumbent 1/18; fabrication is 6/18, 6/18, 6/18, 7/18 versus 6/18. Endpoint norm: 0050.

Only sampled fractions are measured. These are factor interpolations, not linear merged-weight interpolation. Own-zero does not apply to failure counts.

Sources: 0050, 0052 · Reproducible aggregate data (JSON)

The result ruled down magnitude-only recovery on this tested line at these fractions. It did not show that every smaller step fails, or that another direction cannot work. Interpolating LoRA factors also differs from linear interpolation of the effective merged weights. Treating the plotted fractions as a smooth functional response curve would add assumptions the experiment did not establish.

5. Full-sequence versus masked DPO

Full-response scoring might apply pressure to tokens unrelated to the intended decision. Experiment 0053 therefore compared FULL DPO with raw decision-span masking, M1, and shared-pair count-normalized masking, M2. Inputs remained unchanged while the scoring scope changed. 0053

FULL scored 1,269 completion tokens; masking scored 164, comprising eighty chosen and eighty-four rejected tokens. M1 changed direction: its gradient cosine with FULL was 0.771377938, while the simulated Adam-update cosine was 0.535180386. Yet the FULL and M1 update norms were almost equal, 0.006050562709 and 0.006048830971, and all 560 tensors still moved. Masking scoring tokens did not establish parameter or functional locality. 0053

M2 changed the direction further, but concentrated 91.30935% of an isolated pair-energy proxy in one single-token rejected mask. That proxy is not causal attribution. It was nevertheless a reason not to recommend M2 as the continuation. M1 received the actual behavioral test. 0053

In 0054, one actual M1 update exactly matched the previously simulated displacement. Grammar failures stayed at one of eighteen, better than FULL’s three. But missing-evidence fabrication increased from three of six to four of six, and all-case fabrication from six of eighteen to eight of eighteen. Three frozen gates failed. No benchmark was run. 0054

Changing direction mattered numerically. It did not rescue this candidate behaviorally. The negative finding applies to the tested masking recipe, not to all masked preference objectives.

Figure 4Masking and chosen supervision change direction
Two sparse matrices report gradient and fresh Adam update cosines. FULL–M1 is highest; the M2–SL measurement is unavailable.

0053 and 0055: measured cosine similarities for full DPO, two masked DPO objectives (M1/M2), and chosen-span supervised loss (SL).

Direction similarity is not behavioral quality. M2–SL is unmeasured and blank; diagonals are omitted. Own-zero is not applicable to cosines.

Sources: 0053, 0055 · Reproducible aggregate data (JSON)

6. Chosen-span supervision

DPO combines chosen, rejected and reference-policy terms. Perhaps rejected/reference pressure was responsible. Experiment 0055 removed those terms and used chosen-span supervised likelihood: summed negative log likelihood, averaged over the same eight pairs. Six pairs contributed eighty selected tokens; two empty chosen masks contributed zero. 0055

A disposable fresh-Adam simulation again produced a different direction and a similar update norm: 0.006047123185. Exact teacher-forced probing improved seven of twelve spans and forty-one of eighty tokens. Five spans and thirty-nine tokens worsened. Both acquisition tool-selection spans and delegation selection declined. 0055

This was more specific than a failed benchmark total: even positive-only pressure did not consistently improve the selected targets at the finite executable step. It did not establish what free generation would do, because that test was not run. It rejected this mechanism as the next justified actual continuation.

7. Gradient conflict and projected SGD

The next question separated the objective from the optimizer. Did conflicting span gradients or Adam’s transformation explain the regressions? Experiment 0056 computed twelve isolated span gradients and projected raw descent onto local non-regression halfspaces. 0056

For span loss ℓi\ell_i and direction dd, the first-order preservation condition is:

∇ℓi(θ)Td≤0.\nabla\ell_i(\theta)^\mathsf{T}d \leq 0.

Eight of sixty-six gradient pairs had negative cosine. But the aggregate raw descent was already strictly feasible for all twelve constraints; the projection was the identity. Both raw SGD and stored Adam predicted improvement for every span. The constrained construction therefore did not change the direction. 0056

Actual probes at requested L2 radii 0.000378, 0.000756 and 0.001512 improved eight spans and worsened four at each radius. Even the realized rounded deltas predicted improvement for all twelve under the local calculation. No radius passed all seven support criteria. 0056

The mismatch survived removal of Adam and satisfaction of the tested local constraints. That weakens these explanations as necessary causes of this failure. It does not imply gradient conflict never matters, or establish that smooth curvature alone accounts for the response.

8. Numerical execution path

A parameter-space direction becomes behavior only through executable arithmetic. Experiment 0057 made adapter input, matrix products, scaling and accumulation explicitly FP32, retaining the frozen NF4 base and the same direction, masks and radii. 0057

The zero-step function changed, so each path needed its own-zero comparison:

Δℓi,p(r)=ℓi,p(θ+rd)−ℓi,p(θ).\Delta\ell_{i,p}(r)=\ell_{i,p}(\theta+r d)-\ell_{i,p}(\theta).

Here pp identifies an execution path. Negative delta improves that path’s measured target. Comparing two zero baselines is a numerical-path contrast, not a learning gain.

FP32 LoRA improved six, nine and eight of twelve spans at the three radii, versus eight at every radius under BF16. Pooled prediction-sign inversions increased from twelve to thirteen out of thirty-six observations. Precision changed outcomes without restoring overall prediction agreement. Historical BF16 gradients were not freshly computed derivatives of the substituted FP32 path. 0057

The next step was to inspect rather than assume the residual path. Experiment 0058 found all forty residual carriers and both skip additions already FP32. Explicit FP32 residual alternatives were bit-identical to the existing FP32-LoRA path. The BF16 residual-carrier premise was false for this runtime. 0058

Figure 6Numerical-path contrasts through transformer block 0
Five separate investigation boxes cover residual carriers, QKV operands, explicit attention stages, attention output and projection output with FP32 rejoin; native accumulator precision remains unknown.

0058–0062: conceptual map of matched numerical-path interventions and residual controls.

Interventions are contrasted against their own zero. This is an investigation map, not a proven causal chain; explicit attention is not evidence of native internal precision. P1 is not an all-FP32 model.

Sources: 0058, 0059, 0060, 0061, 0062 · Reproducible aggregate data (JSON)

9. BF16/FP32 attention and rejoin boundaries

Experiments 0059–0062 then isolated matched first-block boundaries. They found contributors, but no single tested boundary sufficiently explained or repaired the failure.

Experiment and contrastPrediction-sign inversionsBounded observation
0059: historical versus all Q/K/V operands unrounded13/36 versus 17/36Mean absolute error fell despite more wrong signs
0060: materialized control, progressive FP32 stages20/36; 16, 16, 16, 12/36Explicit arithmetic changed outcomes; historical control was 13/36
0061: unrounded versus rounded attention output13/36 versus 11/36Unrounded matched path was worse overall
0062: rounded versus unrounded projection output13/36 in bothAction-span gains coexisted with worse prediction error

Sources: 0059, 0060, 0061, 0062. These observations use teacher-forced spans across three radii; they are not generated-behavior failure rates.

The contrasts resist a simple “more precision fixes it” story. In 0059, removing operand rounding reduced mean absolute error from 0.059755234 to 0.044248198 nats while increasing sign inversions. In 0062, the unrounded projection output, P1, improved both acquisition action spans at all three radii while worsening prediction error at each radius. Selective improvement and global explanatory adequacy were different questions. 0059, 0062

P1 retained an unrounded FP32 first-block projection output and FP32 residual rejoin. It was not an all-FP32 model. Materialized attention stages did not reveal the hidden accumulator formats of native fused attention. Different implementations could change the zero-step function, limiting attribution back to the historical runtime. These distinctions set the conditions for the next tests.

10. Local derivative instability

Experiment 0063 addressed the historical-gradient caveat directly. It recomputed path-local gradients at exact P1 zero. Every local projection predicted improvement for all twelve spans. Actual outcomes improved seven, seven and nine spans at the fixed radii. Direct finite-difference sign agreement was five, six, three and seven of twelve over four tested epsilon sizes; eleven spans changed finite-difference sign. Fresh same-path gradients did not resolve the disagreement. 0063

For a symmetric executable probe, the directional estimate is:

D^i(h)=ℓi(θ+hd)−ℓi(θ−hd)2h.\widehat D_i(h)=\frac{\ell_i(\theta+h d)-\ell_i(\theta-h d)}{2h}.

This is a finite response estimate. It cannot establish a derivative limit from a few samples, particularly when casts and finite realizations enter the executable path.

Experiment 0064 expanded the test to an eleven-point signed ladder. All twelve spans and all eighty selected tokens responded at the minimum tested radius, 1.4765625×10−61.4765625\times10^{-6}. Central sign agreement ranged from four to ten of twelve. No contiguous region passed the preregistered diagnostic criteria. 0064

The evidence therefore supports no usable local prediction regime on that ladder. It does not show that derivatives mathematically do not exist, or that every smaller scale lacks a useful regime. Deterministic response can still be irregular and strongly scale-dependent; repeatability alone does not make the tested local approximation reliable.

11. Finite-response modeling

If analytic local gradients were unreliable predictors, perhaps measurements of actual finite response could provide the direction. Experiment 0065 built a rank-twelve orthonormal span-gradient subspace, measured signed basis responses at three scales, and froze a median-response minimax construction before validating its direction. 0065

The response matrix was itself unstable: 136 of 144 cells met the descriptive instability rule, while forty-two had consistent nonzero signs. The resulting model nevertheless predicted nonworsening for all twelve spans. Actual positive-direction validation improved seven spans at radius 0.000189 and six at 0.000378. Neither result passed every mandatory support gate. 0065

Aggregate NLL decreased at both radii, by 0.3494768143 and 0.1444549561 nats respectively. Those favorable totals did not establish preservation of every required span. The construction failed its support claim; no generated-behavior generalization was established. 0065

Figure 5Local forecasts do not transfer reliably to finite execution
Projected SGD forecasts 12 improving spans but observes 8 at each radius; same-path derivatives observe 7, 7 and 9; the empirical model forecasts nonworsening for 12 but observes positive improvement in 7 and 6.

0056/0063 predict improvement for all 12 spans. 0065 predicts nonworsening for all 12; observed positive improvement counts are 7 and 6.

Teacher-forced chosen-span NLL responses relative to each path’s own zero, not benchmark scores or generated safety. 0065 predicted nonworsening and observed positive improvement are distinct criteria; panels use different paths/models.

Sources: 0056, 0063, 0065 · Reproducible aggregate data (JSON)

This failure did not prove the subspace unusable. It showed that this frozen empirical linear model and its selected direction did not transfer cleanly to the tested finite evaluations.

12. Pairwise interaction failure

Experiment 0066 asked whether pair interactions could explain that transfer failure. It measured all sixty-six basis pairs at three amplitudes and four sign combinations: 792 settings, each measured twice. Pair interaction meant combined response minus the exact matching signed singleton responses, not a fitted residual against an assumed derivative. 0066

With own-zero responses RiR_i, the measured interaction has the form:

Ii(a,b)=Ri(a+b)−Ri(a)−Ri(b).I_i(a,b)=R_i(a+b)-R_i(a)-R_i(b).

Median absolute interaction was 0.0444965363, 0.0449018478 and 0.0417175293 nats across the amplitudes. Corresponding span-entry sign-reversal rates were 0.256944444, 0.255997475 and 0.243055556, with 3,168 entries per amplitude. These rates describe combined-versus-singleton response comparisons, not benchmark error rates. 0066

Figure 7Pair responses are materially non-additive
Three amplitude panels show median interactions about 0.042–0.045 nats, normalized scores about 0.65–0.68, and sign-reversal rates about 24–26 percent. Quadratic prediction improves one sign at V2 while worsening V2 errors.

0066: exact combined response minus exact signed-singleton response, over 3,168 span entries per amplitude. Linear/quadratic sign agreement: V1 7/12 vs 6/12; V2 6/12 vs 7/12.

Chosen-span own-zero teacher-forced response, not benchmark failure rates. Normalized score and reversal rate are different aggregates. No pair identities or causal rankings are disclosed.

Sources: 0066 · Reproducible aggregate data (JSON)

Strong non-additivity did not yield a sufficient quadratic explanation. At the larger validation radius, the quadratic model achieved seven of twelve sign agreements versus six for the linear model, but its mean absolute error was worse: 0.244920146 versus 0.0997909317 nats. A small sign-count gain did not rescue quantitative prediction. 0066

There was also a structural constraint: a homogeneous quadratic term is even under whole-direction reversal. Keeping the linear term fixed therefore prevents a quadratic correction from changing every positive-versus-negative ordering. The tested fixed-linear-term quadratic model could not fully match the observed ordering. This is a bounded model failure, not exclusion of every nonlinear model.

The validation directions came from known historical 0065 outcomes excluded from fitting and model selection. They were not prospectively unknown outcomes. Larger-radius, multi-coordinate combinations also exceeded the support of the measured pairs. Those limits matter when interpreting the failed extrapolation.

13. Leave-one-component-out findings

Could one component be spoiling an otherwise useful direction? Experiment 0067 removed each stored coefficient-times-basis term from the frozen direction, globally normalized the remainder and evaluated at two radii without refitting. 0067

Across 144 component/span comparisons, seventy effect signs were stable and seventy-four reversed between radii. All twelve removals produced mixed span effects; every removal failed support guards. No pair met eligibility for the optional double ablation, so no double removal was run. 0067

The result supports distributed, context-dependent tradeoffs in these interventions. It does not support a safe complete repair by any tested single removal, or a supported single-component localization. Because global normalization changes every retained amplitude, the effects are contextual, not isolated intrinsic labels of “good” or “bad” components. The experiment does not categorically exclude every possible single-component explanation.

14. What we currently know

The one-step continuation degraded the incumbent under corrected accounting. Reducing its tested factor-interpolation magnitude did not recover an admissible heldout candidate. Masking changed direction without providing a clean behavioral gain. Chosen-only supervision and already-feasible raw descent still produced selected-response regressions. 0050–0056

Numerical-path changes affected outcomes. The residual-carrier hypothesis rested on a false premise, and no single tested precision boundary restored sufficient prediction. Fresh local gradients, a measured finite-response direction and a fixed quadratic pair model each failed their specific support claims. Single normalized removals exposed tradeoffs rather than a clean repair. 0057–0067

Figure 8What the tested explanations did—and did not—resolve
Eight branches record failed tested repairs, the false BF16 residual premise and unresolved precision-boundary sufficiency; no branch claims a universal impossibility.

0052–0067: bounded findings for tested explanations and repairs.

Conceptual verdict tree, not cause shares or global impossibility. Numerical-response tests use own-zero paths; generation-preservation tests do not. Untested scales, paths and models remain open.

Sources: 0052, 0054, 0055, 0056, 0057, 0058, 0059, 0060, 0061, 0062, 0063, 0064, 0065, 0066, 0067 · Reproducible aggregate data (JSON)

These are narrower claims than a complete explanation. Negative results establish which tested constructions failed and which assumptions deserve less confidence. They also preserve information a successful-score narrative would omit: a direction can improve an aggregate and still damage mandatory decisions, and a repeatable executable response can resist the tested local prediction models.

15. What we do not know

We do not have a sufficient causal mechanism for the original degradation. We do not know a direction that improves generated behavior while preserving the full incumbent requirements. This record does not establish cross-hardware, cross-backend or cross-seed generality.

Nor does it establish that DPO, LoRA or Granite generally fail; that mathematical derivatives do not exist; that every local regime is unusable; that the measured subspace contains no safe direction; or that higher-order interactions are the unique cause. Failed pair prediction does not identify exact interaction order. The benchmark and proxy measurements answer different questions, and neither should stand in for an unperformed test.

16. Limitations

This is one incumbent lineage and a specific quantized runtime, with small frozen panels. Exact repeats establish repeatability only where demonstrated on that host/runtime. Twelve selected teacher-forced spans and eighty tokens are investigative proxies, not comprehensive behavioral qualification. Likelihood gains do not certify complete actions, grammar, grounding, query validity or operational safety.

Precision interventions were bounded matched contrasts. Native fused internals remained partly opaque; observed tensor dtype did not establish hidden accumulator precision. Finite ladders did not locate a true onset floor or resolve derivative limits. Linear, minimax and quadratic failures exclude their tested support claims, not whole model families.

The experiments were independently reviewed within the experiment workflow. Review included saved-data and source checks with disclosed limits; it was not external peer review or full independent GPU replication. Restricted benchmark and artifact access also limits public end-to-end reproduction.

17. Reproducibility and corrections

Each immutable experiment ID links to a sanitized record with its fixed inputs, intervention, measurements, verdict and uncertainty. Private final reports and independent reviews remain the authoritative source bindings. Public material omits prompts, completions, case identifiers, token mappings, proprietary schemas, sensitive traces, private infrastructure and model artifacts.

The figure renderer consumes only the sanitized aggregate export. Download the aggregate data, analysis-only renderer and plotting requirements into one directory, install those plotting dependencies, and run:

python3 render_figures.py --data figures.json --out figures

This command redraws figures; it does not run a model or reevaluate the benchmark. The data include units, source IDs and reproducible rendering inputs. They support checking the displayed arithmetic and graphics; they do not enable exact replay of withheld benchmarks. No weights or adapters are distributed with this record.

Frozen 0050 attribution and configuration

The following inventory is bound to Experiment 0050’s frozen run manifest, resolved DPO configuration and incumbent adapter configuration. It identifies the original one-step continuation, not every later intervention. Later numerical-path experiments deliberately changed execution boundaries; their individual records describe those changes. The model revision is an upstream snapshot identifier, not a SHA-256 weight digest. 0050

ItemFrozen 0050 value
Model developer / identifierIBM Granite Team / ibm-granite/granite-4.1-8b
Upstream revision1504002f650e656a0a3789d99574df12e3e94ed0
Base configuration SHA-256dea9d856cb57018117fe2fe3366f37cb4aa39424890061db2c0045a6a4efbda0
Python3.12.13
PyTorch / CUDA build2.14.0 / 13.0
Transformers / TRL5.17.0 / 1.14.0
PEFT / Accelerate0.21.0 / 1.15.0
bitsandbytes / NumPy / safetensors0.50.2 / 2.5.3 / 0.8.0
AdapterLoRA rank 16, alpha 32, bias none
Adapter targetsAttention q/k/v/o projections; MLP gate/up/down projections
DropoutAdapter configuration 0.05; resolved DPO execution disabled dropout
Base quantizationNF4 4-bit, double quantization, BF16 compute
UpdateEight preferences, one epoch, one optimizer update
Learning rate / DPO beta1e-6 / 0.1
Microbatch / accumulation1 / 8
Optimizerpaged_adamw_8bit
Adam betas / epsilon0.9, 0.999 / 1e-8
Weight decay / maximum gradient norm0 / 0.5
Scheduler / warmupConstant / zero
Seed26100231

IBM identifies this model revision as Apache-2.0.3 The model license does not grant rights to Vezryn’s private benchmark, datasets or adapters. VEZRYN LLC retains all rights to its authored text, aggregate data, figures and analysis code; third-party components retain their own licenses. The site preserves the KaTeX MIT notice for its distributed math-rendering assets. Model and method references do not imply endorsement or partnership.

This record’s first public edition is 0.1, dated 2026-10-06, last revised 2026-10-06. Public release describes accessibility of this engineering record, not model promotion or external peer review. Historical 0.1-draft was prepared on 2026-10-06 before release. The correction log remains empty because this is the first public edition; publication-state and attribution completion did not change experimental findings. The historical 242, original matched 244 and corrected canonical 243 scores remain labeled rather than silently replaced. Future corrections must state what changed, why, the evidence and the effect on conclusions.

18. Current research direction

The recorded next hypothesis was bounded boundary tracing along the frozen and matched removed-component paths, aimed at understanding preservation tradeoffs. It is a proposal for mechanism investigation, not a result. Experiment 0068 remains paused; nothing in this article presents it as completed or authorizes its execution.

The present conclusion remains unchanged: the protected 0035/Q1 incumbent, canonical 243/300, has not been cleanly beaten. This sequence produced a more constrained account of what the tested explanations failed to predict. A future successor would need fresh evidence of both improvement and preservation, rather than a favorable local approximation alone.

Footnotes

  1. Rafailov et al., Direct Preference Optimization: Your Language Model is Secretly a Reward Model. The experiments distinguish full-response, masked and normalized variants rather than treating them as identical objectives. ↩

  2. Hu et al., LoRA: Low-Rank Adaptation of Large Language Models. The frozen 0050 adapter used rank 16 and alpha 32; the reproducibility inventory distinguishes configured dropout from resolved DPO execution. ↩

  3. IBM, Granite 4.1 8B model card at the experimental revision and Granite 4.1 repository. IBM identifies revision 1504002f650e656a0a3789d99574df12e3e94ed0 as Apache-2.0. ↩ ↩2

  4. Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs. Cited for quantization/NF4 context, not as a claim that every experimental choice follows that recipe. ↩

Correction log

No corrections to this public edition.