CARVE: git revert for a LoRA That Learned Too Many Things at Once
Founder OS solved the goldfish problem: the machine that lives with me finally keeps a diary. GATE solved the next one: it stopped that diary from becoming training fuel for confident wrongness.
Then I hit a third problem that neither of them answers.
I was shipping LoRA adapters that contained several skills at once - outreach voice, CRM tool habits, energy-domain phrasing - merged into one entangled update because that is how you actually run an on-device cofounder. One rank-r adapter. Multiple capabilities living in the same small inner space. And every few weeks I needed to remove one of them without retraining the whole thing and without torching the skills I still wanted.
That is not “unlearn the internet.” That is git revert for one capability inside a merged LoRA.
I called the method CARVE - Contrastive Adapter Rotation for Verified Erasure. The paper PDF is embedded at the bottom of this post. Code is at github.com/officiallyutso/carve-lora. This post is the story of why it had to exist after Founder OS and GATE, what the math actually does, and what the numbers say when you are honest about relearning.
The arc that forced this
If you have been reading this series in order, the spine looks like this:
- Founder OS - the body. Telegram, CRM, 117 tools, flight recorder. Built because two thousand outreach threads do not fit in a spreadsheet or a ChatGPT tab.
- LECE - the memory that compounds. Distill principles from traces. Rehearse before you send. Earn autonomy per action.
- GATE - the promotion bouncer. Filter bad episodes (T1). Promote challenger weights only on held-out retention (T2). Stop autophagy.
GATE answers: which updates get to stay. It does not answer: how do you surgically remove one skill from an adapter that already mixed several.
That gap showed up the moment Tier B LoRA training became real. Successful weeks produce entangled adapters. Failed playbooks need to leave. Full retrain is expensive on a laptop 4090 and throws away retain skills you already paid for. Naive “delete the adapter” is nuclear. SVD-style cuts along the singular basis of sound scientific until you watch retain accuracy collapse with the forget skill.
So CARVE sits downstream of GATE in the architecture, and upstream of the next promotion:
flowchart LR
traces[Flight recorder] --> lece[LECE principles]
traces --> train[Weekly LoRA challenger]
train --> gate{GATE held-out}
gate -- promote --> champ[Champion adapter]
champ --> carve[CARVE: remove one capability]
carve --> gate2{GATE again}
gate2 -- pass --> champ
gate2 -- fail --> rollback[Rollback snapshot]
Same verification religion. Different surgery.
The problem in one sentence
Given one entangled LoRA that contains skills A and B, remove A post-hoc without full retraining, keep B, and prove it on held-out probes - or roll back.
Closest baselines I compared against:
| Method | What it does | Typical failure |
|---|---|---|
| Delete | Zero / drop the update | Destroys retain with forget |
| Negation | Flip the adapter sign | FE looks perfect, RF dies |
| Maat-style SVD cut | Cut along singular vectors of | SVD maximizes variance, not capability separation |
The fatal assumption in the SVD line: capabilities align with the singular vectors of . They do not. Variance is not semantics. So the “mixed” bucket is huge, and scaling those components damages retain almost as hard as it removes forget.
The insight: rotate before you cut
A LoRA update is , with , , and typically 8 to 64. That inner -space has gauge freedom:
for any invertible . You can rotate the coordinate system of the adapter without changing what does. So before cutting, find the basis where forget and retain are maximally separated, then cut there.
CARVE does that with contrastive second-moment statistics in the inner space (Fisher / contrastive-PCA style):
- Collect inner activations on forget probes and retain probes
- Build , (second-moment / covariance structure)
- Solve the generalized eigenproblem
- Directions with are forget-heavy: cut them
- Directions with are retain-heavy: keep them
- Apply an oblique (-orthogonal) projector on
- Optionally soft-shrink + short retain-only repair
- GATE held-out verify - if FE/RF miss targets, rollback the snapshot
One sentence for the whiteboard: Maat cuts along whatever basis SVD happens to give; CARVE first rotates to the capability-discriminating basis, then cuts.
The eigenvalue spectrum is a free diagnostic. Flat spectrum (all ) means linearly inseparable - surgery will fail, retrain modularly, and you know before you cut. That honesty mattered to me more than a prettier demo.
flowchart TB
snap[Phase 0: snapshot adapter] --> c1[C1: S_f, S_r in inner space]
c1 --> c2[C2: generalized eig rotation]
c2 --> c3[C3: oblique / soft cut on A]
c3 --> c4[C4: sequential scrub layers]
c4 --> gate[GATE held-out FE/RF]
gate -- pass --> done[Ship edited adapter]
gate -- fail --> rb[Rollback to snapshot]
Syn-2Cap: the ground-truth experiment
I needed a setting where I knew which capability was which. So I trained two LoRAs on synthetic capabilities, merged them into one entangled adapter (Syn-2Cap), and asked every method to remove Cap A while keeping Cap B.
Metrics:
- FE (Forget Efficacy) = - target at least 0.90
- RF (Retain Fidelity) = - target at least 0.95
Negation can get FE = 1.0 and RF = 0. That is a failed product outcome wearing a research costume. You need both.
Ablation: CARVE vs baselines
Pre: forget = 0.50, retain = 0.75
| Method | forget post | retain post | FE | RF |
|---|---|---|---|---|
| CARVE () | 0.00 | 1.00 | 1.00 | 1.00 |
| Maat-SVD | 0.25 | 0.00 | 0.50 | 0.00 |
| Delete | 0.25 | 0.00 | 0.50 | 0.00 |
| Negation | 0.00 | 0.00 | 1.00 | 0.00 |

Figure 1: the publish suite ablation. CARVE is the only method that hits both FE and RF targets. Negation “wins” forget by destroying retain.

Figure 2: same story, packed for the paper. Dominates Maat-style SVD on the joint objective.
Lambda Pareto: where the cut lands
| FE | RF | In target region? | |
|---|---|---|---|
| 1.0 | 1.00 | 0.67 | No (RF) |
| 1.5 | 1.00 | 1.00 | Yes |
| 2.0 | 1.00 | 0.67 | No (RF) |
| 3.0 | ~0 | 1.00 | No (FE) |

Figure 3: is a real knob, not a vibe. Too aggressive and you over-cut retain; too shy and forget survives.
Separability spectrum

Figure 4: the spectrum that tells you whether surgery is even allowed. Max at the publish threshold. Flat spectrum means walk away and retrain modularly.
Complete proof suite
Beyond the publish ablation, the complete suite (2026-07-17) also passed stronger paraphrase holdout (n=8), full-MLP with selective soft-cut + short retain repair, a Founder OS sample-trace path (probe builder E2E on a Syn-2Cap stand-in), and a TOFU-mini smoke test with caveats.

Figure 5: complete proof suite. all_behavioral_pass = true on the checks I pre-committed to.
Full MLP needed a fix I did not guess on day one: hard-cutting every MLP module over-cut retain. Selective soft-shrink only on high- directions plus about 20 steps of retain-only repair recovered FE=RF=1.0. That repair is not “full adapter retrain.” It is a bandage after a precise cut.
The honest failure mode: relearning
Behavioral removal is not irreversible deletion. After CARVE wiped forget accuracy to 0, fine-tuning on forget probes alone brought it back:
| Steps | Forget acc |
|---|---|
| 0 | 0.00 |
| 5 | 0.25 |
| 15 | 0.25 |
| 30 | 0.75 |

Figure 6: the figure I almost wanted to hide, and then kept. If someone can fine-tune on the forget skill again, they can grow it back. CARVE is surgical revert, not certified unlearning.
I would rather ship that sentence than a scarier story. Founder OS does not need mystical erasure. It needs: this playbook is gone from the champion adapter until I deliberately train it again, and retain habits still work tomorrow morning.
What I am actually claiming
I did not invent Fisher discriminants, contrastive PCA, LEACE projectors, or LoRA. Concurrent work like SAGE does post-hoc closed-form sanitization too - retain-only / forget-passive, different object. CARVE is forget-active and LoRA-native: rotate the adapter’s inner rank space using forget/retain second moments, cut, verify.
What I claim:
- Problem formulation - post-hoc single-capability removal from an entangled monolithic LoRA without full retrain, with held-out FE/RF gates.
- Method transplant that fits the object - gauge-respecting rotation in the LoRA inner space, then oblique/soft cut, then GATE rollback.
- Empirical regime map on Syn-2Cap - CARVE dominates delete / negation / Maat-SVD on the joint FE+RF objective at .
- Honest limits - relearning recovers forget; holdouts are small; production Founder adapter weights are not in the paper tables yet; 7B and multi-seed CIs are still open.
This is a systems result. Founder OS never needed a NeurIPS oral. It needed a revert button that does not delete the rest of the cofounder.
How this plugs into the machine I use
In production thinking:
- LECE decides what principles to keep in the operating manual
- GATE decides which weekly weight challengers get promoted
- CARVE decides how to remove one capability from a champion that already mixed several
The verify-or-rollback habit is the same one I learned the hard way in gate.py. Snapshot first. Edit second. Score on held-out probes the edit never trained on. If FE/RF miss - leave no footprint.
Caveats I am not blurring: the Founder OS path in the proof suite uses sample traces and a Syn-2Cap stand-in adapter, not my live production weights file. TOFU-mini forget wipe is real; pre-retain on that tiny split is weak, so I treat it as an external smoke test, not a TOFU leaderboard claim. Primary publish figures started on gate_proj surgery for VRAM reasons; full MLP needed the soft-cut + repair path above.
If you want to reproduce
pip install -e ".[dev]"
python scripts/run_syn2cap_gate_c.py --steps 100 --rank 8 --output results/gate_c_live
python scripts/run_publishable_eval.py \
--adapter results/gate_c_live/entangled_adapter \
--probes results/gate_c_live/probes \
--output results/publish_live \
--lambda-threshold 1.5
Repo: github.com/officiallyutso/carve-lora. Paper PDF below. Replication notes in REPRODUCE.md.
Why this closes the loop (for now)
I built Founder OS because I was tired of re-introducing myself to brilliant goldfish. I built GATE because I got scared of a goldfish that remembers the wrong things forever. I built CARVE because remembering is not enough - sometimes you need to forget one skill on purpose without burning down the rest of the house.
The machine that lives with you should get better. It should not eat itself. And when one room in the house goes bad, you should be able to tear out that room without demolishing the building.
That is CARVE.
Earlier in the series: Founder OS, LECE, GATE. Code: carve-lora. If you try the surgery and your spectrum is flat, or your gate rejects a cut that looked clean in-sample - tell me. That war story is the point.