ABOUT VEZRYN RESEARCH
A record that can
change its mind.
Vezryn Research is an applied AI research publication of VEZRYN LLC. We document interventions and evidence, including experiments that did not improve the model.
Observation and interpretation
What happened and why it happened are different claims. Each record identifies its question, fixed inputs, changed variable, measurements, result and unresolved questions. A plausible explanation remains a hypothesis until the evidence supports it.
Negative results belong in the record
A failed continuation, an inadequate approximation or an intervention that does not preserve behavior is worth documenting. Aggregate gains cannot waive safety or preservation requirements. No clean successor in experiments 0050โ0067 beat the protected incumbent.
Review and evidence
The experiments were independently reviewed within the experiment workflow, using saved evidence and source checks with disclosed limitations. This is independent internal review, not external peer review or a claim of fresh independent GPU replication.
Reproducibility with boundaries
This public edition includes sanitized aggregate figure data and an analysis-only rendering script. Exact prompts, benchmark text, case identifiers, token excerpts, traces, schemas and model artifacts remain withheld. Protecting future evaluation limits full public end-to-end replication. Exact repeated measurements on one host/runtime do not establish portability across seeds, hardware or execution backends.
Disclosure
Aggregate methodology and measurements can be made visible without copying raw internal reports. Private user/customer information, credentials, infrastructure addresses and sensitive implementation details are excluded. No model weights, adapters or datasets are redistributed. Public access does not grant redistribution rights to withheld material.
Corrections preserve history
Each publication carries a stable slug, version, original publication date, revision date and correction log. Historical results retain their original labels. A later accounting correction explains what changed without pretending the model learned something new.
In this sequence, the historical incumbent score of 242, the original matched score of 244 and the corrected canonical score of 243 describe accounting of unchanged saved behavior. The failed candidate remains 224. The article explains the distinction.
Attribution
The studied base model is IBM Granite 4.1 8B. The methods draw on LoRA and DPO, with quantized execution discussed in the QLoRA/NF4 context. These references do not imply affiliation or endorsement.