CARVE: git revert for a LoRA That Learned Too Many Things at Once

Founder OS solved the goldfish problem: the machine that lives with me finally keeps a diary. GATE solved the next one: it stopped that diary from becoming training fuel for confident wrongness.

Then I hit a third problem that neither of them answers.

I was shipping LoRA adapters that contained several skills at once - outreach voice, CRM tool habits, energy-domain phrasing - merged into one entangled update because that is how you actually run an on-device cofounder. One rank-r adapter. Multiple capabilities living in the same small inner space. And every few weeks I needed to remove one of them without retraining the whole thing and without torching the skills I still wanted.

That is not “unlearn the internet.” That is git revert for one capability inside a merged LoRA.

I called the method CARVE - Contrastive Adapter Rotation for Verified Erasure. The paper PDF is embedded at the bottom of this post. Code is at github.com/officiallyutso/carve-lora. This post is the story of why it had to exist after Founder OS and GATE, what the math actually does, and what the numbers say when you are honest about relearning.

The arc that forced this

If you have been reading this series in order, the spine looks like this:

  1. Founder OS - the body. Telegram, CRM, 117 tools, flight recorder. Built because two thousand outreach threads do not fit in a spreadsheet or a ChatGPT tab.
  2. LECE - the memory that compounds. Distill principles from traces. Rehearse before you send. Earn autonomy per action.
  3. GATE - the promotion bouncer. Filter bad episodes (T1). Promote challenger weights only on held-out retention (T2). Stop autophagy.

GATE answers: which updates get to stay. It does not answer: how do you surgically remove one skill from an adapter that already mixed several.

That gap showed up the moment Tier B LoRA training became real. Successful weeks produce entangled adapters. Failed playbooks need to leave. Full retrain is expensive on a laptop 4090 and throws away retain skills you already paid for. Naive “delete the adapter” is nuclear. SVD-style cuts along the singular basis of ΔW\Delta W sound scientific until you watch retain accuracy collapse with the forget skill.

So CARVE sits downstream of GATE in the architecture, and upstream of the next promotion:

flowchart LR
    traces[Flight recorder] --> lece[LECE principles]
    traces --> train[Weekly LoRA challenger]
    train --> gate{GATE held-out}
    gate -- promote --> champ[Champion adapter]
    champ --> carve[CARVE: remove one capability]
    carve --> gate2{GATE again}
    gate2 -- pass --> champ
    gate2 -- fail --> rollback[Rollback snapshot]

Same verification religion. Different surgery.

The problem in one sentence

Given one entangled LoRA that contains skills A and B, remove A post-hoc without full retraining, keep B, and prove it on held-out probes - or roll back.

Closest baselines I compared against:

MethodWhat it doesTypical failure
DeleteZero / drop the updateDestroys retain with forget
NegationFlip the adapter signFE looks perfect, RF dies
Maat-style SVD cutCut along singular vectors of ΔW\Delta WSVD maximizes variance, not capability separation

The fatal assumption in the SVD line: capabilities align with the singular vectors of ΔW\Delta W. They do not. Variance is not semantics. So the “mixed” bucket is huge, and scaling those components damages retain almost as hard as it removes forget.

The insight: rotate before you cut

A LoRA update is ΔW=s BA\Delta W = s\,BA, with A∈Rr×dinA \in \mathbb{R}^{r \times d_{\mathrm{in}}}, B∈Rdout×rB \in \mathbb{R}^{d_{\mathrm{out}} \times r}, and rr typically 8 to 64. That inner rr-space has gauge freedom:

BA=(BR−1)(RA)BA = (B R^{-1})(R A)

for any invertible RR. You can rotate the coordinate system of the adapter without changing what ΔW\Delta W does. So before cutting, find the basis where forget and retain are maximally separated, then cut there.

CARVE does that with contrastive second-moment statistics in the inner space (Fisher / contrastive-PCA style):

  1. Collect inner activations z=Ahz = Ah on forget probes DfD_f and retain probes DrD_r
  2. Build SfS_f, SrS_r (second-moment / covariance structure)
  3. Solve the generalized eigenproblem Sfw=λSrwS_f w = \lambda S_r w
  4. Directions with λ≫1\lambda \gg 1 are forget-heavy: cut them
  5. Directions with λ≪1\lambda \ll 1 are retain-heavy: keep them
  6. Apply an oblique (SrS_r-orthogonal) projector on AA
  7. Optionally soft-shrink + short retain-only repair
  8. GATE held-out verify - if FE/RF miss targets, rollback the snapshot

One sentence for the whiteboard: Maat cuts along whatever basis SVD happens to give; CARVE first rotates to the capability-discriminating basis, then cuts.

The eigenvalue spectrum is a free diagnostic. Flat spectrum (all λ≈1\lambda \approx 1) means linearly inseparable - surgery will fail, retrain modularly, and you know before you cut. That honesty mattered to me more than a prettier demo.

flowchart TB
    snap[Phase 0: snapshot adapter] --> c1[C1: S_f, S_r in inner space]
    c1 --> c2[C2: generalized eig rotation]
    c2 --> c3[C3: oblique / soft cut on A]
    c3 --> c4[C4: sequential scrub layers]
    c4 --> gate[GATE held-out FE/RF]
    gate -- pass --> done[Ship edited adapter]
    gate -- fail --> rb[Rollback to snapshot]

Syn-2Cap: the ground-truth experiment

I needed a setting where I knew which capability was which. So I trained two LoRAs on synthetic capabilities, merged them into one entangled adapter (Syn-2Cap), and asked every method to remove Cap A while keeping Cap B.

Metrics:

Negation can get FE = 1.0 and RF = 0. That is a failed product outcome wearing a research costume. You need both.

Ablation: CARVE vs baselines

Pre: forget = 0.50, retain = 0.75

Methodforget postretain postFERF
CARVE (λ=1.5\lambda=1.5)0.001.001.001.00
Maat-SVD0.250.000.500.00
Delete0.250.000.500.00
Negation0.000.001.000.00

Syn-2Cap ablation FE/RF bars for CARVE vs Maat-SVD, Delete, Negation

Figure 1: the publish suite ablation. CARVE is the only method that hits both FE and RF targets. Negation “wins” forget by destroying retain.

CARVE vs baselines side-by-side

Figure 2: same story, packed for the paper. Dominates Maat-style SVD on the joint objective.

Lambda Pareto: where the cut lands

λ\lambdaFERFIn target region?
1.01.000.67No (RF)
1.51.001.00Yes
2.01.000.67No (RF)
3.0~01.00No (FE)

FE-RF Pareto across lambda

Figure 3: λ\lambda is a real knob, not a vibe. Too aggressive and you over-cut retain; too shy and forget survives.

Separability spectrum

Eigenvalue spectrum SURGERY_VIABLE

Figure 4: the spectrum that tells you whether surgery is even allowed. Max λ≈1.68\lambda \approx 1.68 at the publish threshold. Flat spectrum means walk away and retrain modularly.

Complete proof suite

Beyond the publish ablation, the complete suite (2026-07-17) also passed stronger paraphrase holdout (n=8), full-MLP with selective soft-cut + short retain repair, a Founder OS sample-trace path (probe builder E2E on a Syn-2Cap stand-in), and a TOFU-mini smoke test with caveats.

Proof completion FE/RF bars

Figure 5: complete proof suite. all_behavioral_pass = true on the checks I pre-committed to.

Full MLP needed a fix I did not guess on day one: hard-cutting every MLP module over-cut retain. Selective soft-shrink only on high-λ\lambda directions plus about 20 steps of retain-only repair recovered FE=RF=1.0. That repair is not “full adapter retrain.” It is a bandage after a precise cut.

The honest failure mode: relearning

Behavioral removal is not irreversible deletion. After CARVE wiped forget accuracy to 0, fine-tuning on forget probes alone brought it back:

StepsForget acc
00.00
50.25
150.25
300.75

Relearning curve after CARVE

Figure 6: the figure I almost wanted to hide, and then kept. If someone can fine-tune on the forget skill again, they can grow it back. CARVE is surgical revert, not certified unlearning.

I would rather ship that sentence than a scarier story. Founder OS does not need mystical erasure. It needs: this playbook is gone from the champion adapter until I deliberately train it again, and retain habits still work tomorrow morning.

What I am actually claiming

I did not invent Fisher discriminants, contrastive PCA, LEACE projectors, or LoRA. Concurrent work like SAGE does post-hoc closed-form sanitization too - retain-only / forget-passive, different object. CARVE is forget-active and LoRA-native: rotate the adapter’s inner rank space using forget/retain second moments, cut, verify.

What I claim:

  1. Problem formulation - post-hoc single-capability removal from an entangled monolithic LoRA without full retrain, with held-out FE/RF gates.
  2. Method transplant that fits the object - gauge-respecting rotation in the LoRA inner space, then oblique/soft cut, then GATE rollback.
  3. Empirical regime map on Syn-2Cap - CARVE dominates delete / negation / Maat-SVD on the joint FE+RF objective at λ≈1.5\lambda \approx 1.5.
  4. Honest limits - relearning recovers forget; holdouts are small; production Founder adapter weights are not in the paper tables yet; 7B and multi-seed CIs are still open.

This is a systems result. Founder OS never needed a NeurIPS oral. It needed a revert button that does not delete the rest of the cofounder.

How this plugs into the machine I use

In production thinking:

The verify-or-rollback habit is the same one I learned the hard way in gate.py. Snapshot first. Edit second. Score on held-out probes the edit never trained on. If FE/RF miss - leave no footprint.

Caveats I am not blurring: the Founder OS path in the proof suite uses sample traces and a Syn-2Cap stand-in adapter, not my live production weights file. TOFU-mini forget wipe is real; pre-retain on that tiny split is weak, so I treat it as an external smoke test, not a TOFU leaderboard claim. Primary publish figures started on gate_proj surgery for VRAM reasons; full MLP needed the soft-cut + repair path above.

If you want to reproduce

pip install -e ".[dev]"
python scripts/run_syn2cap_gate_c.py --steps 100 --rank 8 --output results/gate_c_live
python scripts/run_publishable_eval.py \
  --adapter results/gate_c_live/entangled_adapter \
  --probes results/gate_c_live/probes \
  --output results/publish_live \
  --lambda-threshold 1.5

Repo: github.com/officiallyutso/carve-lora. Paper PDF below. Replication notes in REPRODUCE.md.

Why this closes the loop (for now)

I built Founder OS because I was tired of re-introducing myself to brilliant goldfish. I built GATE because I got scared of a goldfish that remembers the wrong things forever. I built CARVE because remembering is not enough - sometimes you need to forget one skill on purpose without burning down the rest of the house.

The machine that lives with you should get better. It should not eat itself. And when one room in the house goes bad, you should be able to tear out that room without demolishing the building.

That is CARVE.


Earlier in the series: Founder OS, LECE, GATE. Code: carve-lora. If you try the surgery and your spectrum is flat, or your gate rejects a cut that looked clean in-sample - tell me. That war story is the point.

Full whitepaper

Scroll to read the PDF inline, ordownload the PDFfor offline reading.

--claps