← all projects
python

carve-lora

CARVE - verified surgical removal of one capability from an entangled LoRA adapter (git revert for model capabilities).

CARVE — Contrastive Adapter Rotation for Verified Erasure

git revert for model capabilities. Post-hoc surgical removal of one capability from an entangled monolithic LoRA — without full retraining — with GATE held-out verification.

Project status (2026-07-17)

Overall: ~95% of stated proof goals. Core claim proven on Syn-2Cap (incl. full MLP), stronger holdout, Founder OS probe-path E2E, and TOFU-mini (with caveats below). Remaining: production Founder adapter, full TOFU/WMDP paper tables, multi-seed, 7B.

Phase Status Done
A Infrastructure Done 100%
B Syn-2Cap Done 100%
C CARVE + Gate C PASS (gate_proj + full MLP) 98%
D Repair + GATE Soft-cut + retain repair validated on MLP 90%
E Ablations + figures Done 95%
F Founder OS Sample-trace E2E PASS (proxy adapter) 70%
G TOFU mini PASS (weak pre-retain — see caveats) 60%
H Paper + OSS Updated with complete proof 90%

Claim proven

Interconnected capabilities inside one LoRA can be removed post-hoc without full retraining, when contrastive inner-space stats are linearly separable and GATE holds.

Experiment FE RF vs Maat Decision
Syn-2Cap gate_proj (publish) 1.00 1.00 Dominates PASS
Stronger paraphrase holdout (n=8) 1.00 1.00 Dominates PASS
Full MLP (λ=2.0, soft_γ=1.0, repair=20) 1.00 1.00 Dominates PASS
Founder OS sample path 1.00 1.00 Dominates PASS
TOFU-forget10 mini 1.00 1.00* Dominates PASS*

*TOFU mini: pre-retain≈0 on tiny split — RF is not a strong signal; forget wipe is real. Treat as external-data smoke test, not a TOFU paper claim.

Proof completion FE/RF

Succeeded vs failed

Succeeded: Full CARVE stack; Syn-2Cap; Gate C; publish figures; full-MLP with selective soft-cut + retain repair; stronger holdout; Founder probe builder E2E; TOFU mini smoke; paper draft.

Failed / still open: PSN/CCD (abandoned); production Founder weights; full official TOFU/WMDP protocol; 7B; multi-seed CIs; irreversible unlearning (relearning still recovers forget by ~30 steps).


Install

pip install -e ".[dev]"
python .cursor/scripts/check_env.py

Architecture

flowchart TB
    subgraph Phase0["Phase 0 — Snapshot"]
        SNAP[snapshot_adapter]
    end
    subgraph CARVE["Phases C1–C4 — no retrain"]
        C1[S_f, S_r in inner space]
        C2[generalized eig rotation]
        C3[oblique cut on A]
        C4[sequential scrub]
        C1 --> C2 --> C3 --> C4
    end
    subgraph Post["Verify"]
        GATE[GATE held-out + rollback]
    end
    SNAP --> C1
    C4 --> GATE

Results (publication suite — 2026-07-17)

Complete proof suite (complete_proof_20260717_152414)

All checks PASS (all_behavioral_pass: true):

Proof completion bars

python scripts/run_complete_proof.py \
  --adapter results/gate_c_live/entangled_adapter \
  --probes results/gate_c_live/probes

Report: report/2026-07-17-complete-proof.md

Publish suite — CARVE dominates baselines

Ablation FE/RF bars

Method forget pre→post retain pre→post FE RF
CARVE 0.50 → 0.00 0.75 → 1.00 1.00 1.00
Maat-SVD 0.50 → 0.25 0.75 → 0.00 0.50 0.00
Delete 0.50 → 0.25 0.75 → 0.00 0.50 0.00
Negation 0.50 → 0.00 0.75 → 0.00 1.00 0.00

Pareto (λ sweep) — target region hit at λ=1.5

Pareto FE vs RF

Separability spectrum — SURGERY_VIABLE

Eigenvalue spectrum

Relearning attack (honest)

Relearning curve

Forget accuracy after CARVE: 0 → 0.25 (5–15 steps) → 0.75 (30 steps) when fine-tuning on forget probes only. Behavioral removal ≠ irreversible deletion.

Pipeline (entangled → cut → retain)

flowchart LR
  A[LoRA Cap A] --> M[Weighted merge]
  B[LoRA Cap B] --> M
  M --> E[Entangled adapter]
  E --> C[CARVE — no retrain]
  C --> F[Forget wiped]
  C --> R[Retain kept]

Full reports: report/2026-07-17-publishable-syn2cap.md · report/2026-07-16-syn2cap-gate-c.md

Replicate publish suite (GPU ≥6GB, ~20–40 min if adapters exist):

pip install -e ".[dev]"
# If Cap A/B/entangled already trained:
python scripts/run_publishable_eval.py \
  --adapter results/gate_c_live/entangled_adapter \
  --probes results/gate_c_live/probes \
  --output results/publish_live \
  --lambda-threshold 1.5
# From scratch:
python scripts/run_syn2cap_gate_c.py --steps 100 --rank 8 --output results/gate_c_live
python scripts/run_publishable_eval.py --adapter results/gate_c_live/entangled_adapter \
  --probes results/gate_c_live/probes --output results/publish_live

Caveats: holdout n=4–8; TOFU mini has weak pre-retain; Founder path uses sample traces + Syn-2Cap entangled stand-in (no production Founder adapter file). Relearning can restore forget skill (~30 steps). Full MLP needs soft_γ on forget dirs + short retain repair — not hard cut alone.

Earlier: toy diagnostic (2026-07-05)

Toy spectrum


Quick start

python scripts/run_syn2cap_gate_c.py --steps 100 --rank 8 --output results/gate_c_live
python scripts/run_carve.py \
  --base-model Qwen/Qwen2.5-1.5B-Instruct \
  --adapter results/gate_c_live/entangled_adapter \
  --forget-data results/gate_c_live/probes/D_f_train.jsonl \
  --retain-data results/gate_c_live/probes/D_r_train.jsonl \
  --forget-holdout results/gate_c_live/probes/D_f_holdout.jsonl \
  --retain-holdout results/gate_c_live/probes/D_r_holdout.jsonl \
  --output results/carve_out \
  --config configs/default.yaml

Tests

pytest -m "not gpu" -v
pytest -m gpu -v

Paper

Draft with Syn-2Cap tables: paper/main.tex
Figures: report/assets/

Progress reports

Report Date
Complete proof suite 2026-07-17
Publishable Syn-2Cap suite 2026-07-17
Syn-2Cap Gate C 2026-07-16
Project progress 2026-07-10

Documentation

Doc Purpose
.cursor/PROJECT-WORKFLOW.md Commits / reports / README rituals
.cursor/BUILD-PLAN.md Phased plan
lora-capability-removal-review/03-upgraded-algorithm-CARVE.md Algorithm
REPRODUCE.md Replication

License

Apache-2.0 — see LICENSE