RELEASE · MODELS · #723
TinyCeNN-LM: quality-gated conversion framework replacing pretrained attention with CeNN-like recurrent layers
arXiv:2609.21139v1 introduces TinyCeNN-LM, a quality-gated post-training conversion framework that replaces attention layers with CeNN-inspired cellular-recurrent modules containing bounded local processing, compact recurrent memory, routing/fusion, and accept-or-rollback validation. The paper evaluates three implementations (Integrated Memory, MemoryFusion, PDelta3-GDN2-CLVR+Local32): on SmolLM2-135M layers 0–2 pass with cumulative ΔNLL=+0.01209 while layer 3 is rejected due to poor representation fidelity; on Qwen3.5-0.8B full-attention layers 3, 7, 11 are accepted with final ΔNLL=+0.02073; Integrated Memory keeps perplexity within −0.07% to +0.93% and can reduce total cache by up to 6.01%, and a 200-item downstream sanity check on converted Qwen releases yields 28.5%–32.0% accuracy.
KEY POINTS
- arXiv:2609.21139v1 introduces TinyCeNN-LM, a quality-gated post-training conversion framework that replaces attention layers with CeNN-inspired cellular-recurrent modules containing bounded local processing, compact recurrent memory, routing/fusion, and accept-or-rollback validation.
- The paper evaluates three implementations (Integrated Memory, MemoryFusion, PDelta3-GDN2-CLVR+Local32): on SmolLM2-135M layers 0–2 pass with cumulative ΔNLL=+0.01209 while layer 3 is rejected due to poor representation fidelity; on Qwen3.5-0.8B full-attention layers 3, 7, 11 are accepted with final ΔNLL=+0.02073; Integrated Memory keeps perplexity within −0.07% to +0.93% and can reduce total cache by up to 6.01%, and a 200-item downstream sanity check on converted Qwen releases yields 28.5%–32.0% accuracy.
- This proposes a conservative, test-gated method to replace attention post-training with compact recurrent layers and validates acceptance by both representation fidelity and NLL, informing practical paths for attention alternatives without wholesale model retraining.
WHY IT MATTERS
This proposes a conservative, test-gated method to replace attention post-training with compact recurrent layers and validates acceptance by both representation fidelity and NLL, informing practical paths for attention alternatives without wholesale model retraining.