Research program · ANUPNET

Language models that keep learning — without forgetting.

A small language model that learns from a single unlabeled stream of new domains, stores no old data, replays nothing, and is never told when the domain changes — while keeping most of what it learned before.

Technical report v0 (PDF) Pitch deck (PDF) Talk to the author

Stage E (Qwen 2B/9B + slot memory) is summarised below; report v1 will include it.

88%
less catastrophic forgetting than naive fine-tuning (2.74 → 0.33 nats, 3 seeds, 30M model)
0
stored data, replay buffers, or task labels used by the mechanism
4 / 4
domain boundaries detected at the exact step in an unlabeled stream, 0 false alarms, 3 seeds
110M
parameter model pretrained on 1B tokens as the base for the next stage
13 / 14
facts retained by a 9B model after a 6-day unlabeled stream, no stored data (Stage E)
The problem

Today's models are snapshots.

Every large language model is frozen the day its training ends. Teach it a new domain and it overwrites the old one — catastrophic forgetting. The standard fixes either store old data and replay it, or need a label telling the model which task it is on. Humans do neither: skills, gist and connections stay while details fade.

ANUPNET is built under three rules we do not break: no stored old data, no task labels, and no absolute numbers in decisions — every threshold is a ratio to the model's own running statistics.

How it works

A hall, rooms, dreams and sleep.

1
Hall and rooms
The trunk ("hall") is frozen after the first domain. Every FFN and attention block gets small gated side modules ("rooms") — one per domain, opened only when its domain is recognised.
2
Dreams as contrast
A gate cannot learn old-vs-new without seeing old inputs. Instead of a data buffer, the model generates its own text ("dreams") and uses it only to train the gates — never the weights.
3
Sleep
When a room closes, it learns to stay silent on dreams, and older rooms' keys are refreshed so they still recognise their own domain as the input distribution drifts.
4
Novelty and familiarity
A loss jump above the model's own noise band opens a new room before the step. A dormant room whose gate wakes up on the input is reopened instead — so revisits go back to the right memory.
Results · 30M model, wiki → code → stories

Measured, seeded, with controls.

MethodStored dataAvg forgetting ↓
Capacity-matched control (same extra params, no mechanism)–3.70
Naive fine-tuning–2.74
Freeze trunk only–2.58
Dream gate (FFN rooms)–0.87 ± 0.04
ANUPNET (+ sleep + attention rooms)–0.33 ± 0.05
Replay 30% (stores old data — reference)yes0.05

Forgetting = rise in validation loss (nats) on old domains after training on new ones; 800 steps per domain; 3 seeds for the two headline rows. What we are not claiming: ANUPNET does not beat replay when storing data is allowed; new-domain acquisition costs ~0.45 nats; results are at 30M parameters and 3 domains. All negative results are in the report.

Stage E · Sep 2026

Frozen open-weight hall (Qwen 2B / 9B) + rooms + slot memory.

The same three rules (no stored data, no task labels, no absolute thresholds) now run on top of a frozen open-weight model instead of our own 30M/110M trunk. The base model is the "hall" (never trained); small gated side modules ("rooms") learn skills; and a new hippocampus-style episodic memory gives each learned fact its own key–value slot. The model learns from a daily stream of chat, facts and tasks, sleeps once per "night" (self-generated rehearsal, no stored text), and can offload old rooms to CPU and wake them on demand ("yaad aaya") so serving cost stays flat as knowledge grows.

SettingPlain fine-tune baselineANUPNET
9B hall, 6-day stream, retained facts—13/14 and 19/20
2B hall, 10-day stream (50 items/day), learned facts28% (template collapse)60% (slot memory v3)
2B hall, 20-day stream, learned facts22%in progress (slot memory v4)
Skill retention across 6 task families at scale77–90%90%+
Cross-room forgetting after new learningpresent0 measured
Offload old room to CPU, then recall on question—lossless (12/14, 17/20 = same as no-offload)

"Facts" = questions about people/places/deadlines taught once in the stream and asked days later; "skills" = situation → conclusion rules (e.g. ticket priority, expense category) never stated explicitly. Baseline = same rooms trained without slot memory. All runs on Google Cloud Run GPU jobs (NVIDIA RTX Pro 6000), 2B/9B = Qwen open-weight models.

What we found
1
Template collapse.
With many same-pattern facts ("X's deadline is <month>"), dense side modules keep only the most recent one — facts fall from 62% (day 5) to 22% (day 20). This is the failure mode slot memory is built to break.
2
Slot memory needs frozen keys and owner-only updates.
Keys must come from the frozen hall (not from layers the rooms modify) and each slot's value may only be updated by its own fact — otherwise Adam drifts every slot a little every day. v3 doubled day-10 fact retention over baseline (60% vs 28%).
3
Offloading is free.
Rooms moved to CPU and woken on demand retain every answer of the always-on run.
4
Open (v4, running).
Making retrieval robust to paraphrase and near-duplicate questions; auto-triggered sleep (the model decides when to consolidate).

What we are not claiming: Stage E results are single-seed so far; 30-day and 9B-scale confirmation runs are next and need sustained GPU.

Roadmap

From 30M to a frozen 9B hall.

Done
A · Pretrain
110M params, 1.05B tokens (FineWeb-Edu + Cosmopedia), val loss 3.30, 6 GPU-hours on one NVIDIA RTX Pro 6000.
Done
B · Instruction tune
Q&A format on SmolTalk; the model answers questions in English.
Done
C · Frozen open-weight hall
Switched the hall to Qwen 2B/9B; rooms and slot memory on top (Stage E).
In progress (Stage E)
D · Continual stream
Running now on frozen Qwen 2B/9B instead of 125M: rooms + slot memory over 10–30 day streams. 30-day and 9B confirmation runs next.
Next
E · Product pilot
WhatsApp business assistant that remembers each customer permanently without retraining (Agent Lexi on ANUPNET).

Code, configs and the full experiment log (including failures) will be released with report v1. Compute partners and reviewers welcome — admin@lexinexiai.com.

Go To Top