RESEARCH RECORD 001REV.
POST-TRAINING / GRANITE 4.1 8B / COORDINATION
When One Fine-Tuning Step Broke a Stable 8B Coordinator
What a sequence of failed post-training experiments taught us about DPO, LoRA, numerical execution, finite-response instability, and behavioral interaction.
No clean successor has beaten the protected incumbent.
Read research record- RECORD
- RR-001
- SYSTEM
- Granite 4.1 8B
- METHOD
- LoRA / DPO
- OUTCOME
- Rejected continuation
- EXPERIMENTS
- 0050–0067
- EDITION
- 0.1