A small language model that learns from a single unlabeled stream of new domains, stores no old data, replays nothing, and is never told when the domain changes — while keeping most of what it learned before.
Stage E (Qwen 2B/9B + slot memory) is summarised below; report v1 will include it.
Every large language model is frozen the day its training ends. Teach it a new domain and it overwrites the old one — catastrophic forgetting. The standard fixes either store old data and replay it, or need a label telling the model which task it is on. Humans do neither: skills, gist and connections stay while details fade.
ANUPNET is built under three rules we do not break: no stored old data, no task labels, and no absolute numbers in decisions — every threshold is a ratio to the model's own running statistics.
| Method | Stored data | Avg forgetting ↓ |
|---|---|---|
| Capacity-matched control (same extra params, no mechanism) | – | 3.70 |
| Naive fine-tuning | – | 2.74 |
| Freeze trunk only | – | 2.58 |
| Dream gate (FFN rooms) | – | 0.87 ± 0.04 |
| ANUPNET (+ sleep + attention rooms) | – | 0.33 ± 0.05 |
| Replay 30% (stores old data — reference) | yes | 0.05 |
Forgetting = rise in validation loss (nats) on old domains after training on new ones; 800 steps per domain; 3 seeds for the two headline rows. What we are not claiming: ANUPNET does not beat replay when storing data is allowed; new-domain acquisition costs ~0.45 nats; results are at 30M parameters and 3 domains. All negative results are in the report.
The same three rules (no stored data, no task labels, no absolute thresholds) now run on top of a frozen open-weight model instead of our own 30M/110M trunk. The base model is the "hall" (never trained); small gated side modules ("rooms") learn skills; and a new hippocampus-style episodic memory gives each learned fact its own key–value slot. The model learns from a daily stream of chat, facts and tasks, sleeps once per "night" (self-generated rehearsal, no stored text), and can offload old rooms to CPU and wake them on demand ("yaad aaya") so serving cost stays flat as knowledge grows.
| Setting | Plain fine-tune baseline | ANUPNET |
|---|---|---|
| 9B hall, 6-day stream, retained facts | — | 13/14 and 19/20 |
| 2B hall, 10-day stream (50 items/day), learned facts | 28% (template collapse) | 60% (slot memory v3) |
| 2B hall, 20-day stream, learned facts | 22% | in progress (slot memory v4) |
| Skill retention across 6 task families at scale | 77–90% | 90%+ |
| Cross-room forgetting after new learning | present | 0 measured |
| Offload old room to CPU, then recall on question | — | lossless (12/14, 17/20 = same as no-offload) |
"Facts" = questions about people/places/deadlines taught once in the stream and asked days later; "skills" = situation → conclusion rules (e.g. ticket priority, expense category) never stated explicitly. Baseline = same rooms trained without slot memory. All runs on Google Cloud Run GPU jobs (NVIDIA RTX Pro 6000), 2B/9B = Qwen open-weight models.
What we are not claiming: Stage E results are single-seed so far; 30-day and 9B-scale confirmation runs are next and need sustained GPU.
Code, configs and the full experiment log (including failures) will be released with report v1. Compute partners and reviewers welcome — admin@lexinexiai.com.