GATE: How I Taught My AI to Learn From Its Own Life Without Eating Itself

I built an AI cofounder that records everything it does, and then I tried to make it learn from that recording. For about three weeks it got measurably worse. Not broken, not crashing. Worse in the way that is hard to catch: more confident, quicker to answer, and wrong more often. It was training on its own bad days and treating them as lessons.

The thing that fixed it was not a bigger model or a better prompt. It was a gate. Two of them, once I understood the problem properly.

This post is the story of how I found that out, how I proved it to myself instead of just believing it, and how it now runs on the machine I use to run my company. I call it GATE, for Gated Adaptive Training from Experience.

If you have not read the earlier posts, they set this up. I Built Founder OS Because My CRM Was Lying to Me is why the whole system exists, and LECE: How Founder OS Learns From Your Life is the learning engine. This one is what happened when I let that engine touch the actual weights.

The lie in gate.py

There is a file in Founder OS called agent/cognition/train/gate.py. When I wrote it I gave it a docstring I was quietly proud of:

Eval gate: promote adapter only if it beats baseline (anti-collapse).

The plan behind it was textbook. Train a small LoRA adapter on the agent’s successful episodes, score the candidate, compare it to the current champion, and only promote it if it wins. Otherwise roll back. It looked like science.

Then one night I sat down and actually read what the scoring function did.

def _eval_adapter(dataset: list) -> float:
    """Proxy eval: fraction of high-score examples (full eval uses evals/lece/)."""
    scores = [ex.get("metadata", {}).get("score", 0.5) for ex in dataset]
    return sum(scores) / len(scores)

It was scoring the challenger on the same episodes it had just trained on. The model was grading its own homework. Every challenger “won” because it was being tested on the exact material it had memorised. Promotions looked healthy. Nothing collapsed in my dev runs. Of course it didn’t, because the gate was a mirror, not a judge.

I remember the specific feeling, and it was not a good one. This was the part of the system I had been telling people prevented collapse, and it did nothing. That is where GATE stopped being a feature I had shipped and became a question I actually needed to answer.

LECE had already taught me the useful half of this. The most valuable training data for my assistant is not the internet, it is the private day-by-day record of my own company, and distilling principles out of that record every night genuinely works. Principles are basically retrieval, so they don’t collapse. Updating the weights was the obvious next step, and it is also the step where everything gets dangerous.

Fine-tuning a model on its own outputs, with no honest check, is not the system getting better. It is the system eating itself. I wanted to know what actually stops that, and I did not trust my own opinion anymore, so I stopped writing product code for a while and set up a research harness in research/. I wrote down what I expected to see before I ran anything, and then I spent six weeks trying to prove myself wrong.

The problem Founder OS made me name

Founder OS is the autonomous cofounder I run on my own machine: 117 tools, swarm mode when a task needs parallel specialists, self-healing so it survives more than one demo, and a flight recorder that writes every turn to data/traces/*.jsonl. LECE reads that recording, scores the episodes, keeps the good ones, and turns them into principles. That loop is stable.

Weights are a different animal. The moment you fine-tune a model on its own trajectories you walk into the regime the literature keeps warning about: self-consuming training, model collapse, the slow drift toward a narrower and more confident version of yourself that happens to be wrong more often.

Founder OS was already sitting in that regime before I had the vocabulary for it. Every email it drafted became training material. Every tool call, good or bad, left a trace. The outcomes that told me whether an action worked arrived late and messy, sometimes days later when an investor finally replied or a CRM update turned out to have stuck. And the signal I was using to score all of this was noisy to begin with: keyword checks, approval outcomes, a critic that is sometimes confidently wrong.

This is not the clean-labels world most continual-learning papers live in. It is one user, self-generated data, agentic tool use, sparse and delayed reward, and real collapse pressure, all at once. I looked for prior work in exactly that box and could not find a clean causal study. SPRInG does selective LoRA on gold human text. TSUBASA distills clean synthetic QA. TMEM throws away its fast weights every episode. None of them need a promotion gate, because none of them are being fed a diet of their own mistakes every day. Founder OS is.

The theory says verification is what prevents collapse, and I believe the theory. But it leaves one question open that mattered enormously to me: is one layer of verification enough, or do you need two? I thought you needed two.

The first tier, T1, is a sample filter. Before an episode is allowed anywhere near training, its estimated reward has to clear a threshold. Bad self-labels get blocked at the door. The second tier, T2, is the promotion gate. Train a challenger adapter each week, then only promote it over the champion if it does better on held-out retention, facts it never trained on. GATE is both tiers working together.

The entire claim comes down to one comparison: GATE versus T1 on its own. If filtering alone matches the full system, then the second gate is just ceremony and I should delete it. If GATE wins by a real margin, the second tier is earning its place. That is a claim you can actually kill, so I pre-registered it in research/paper/preregistration.md and wrote down in research/paper/claims_table.md what I would have to admit if the numbers went against me. I was not going to grade my own homework a second time.

Building something I could trust

Before measuring any learning, I had to trust the ruler. My held-out benchmark, which I called LivedBench, has three parts: recent facts the agent should recall from the current week, older facts that reveal whether new learning quietly destroyed old memory, and judgment questions where the right call is not obvious.

The first milestone, M0, only asked whether the baselines lined up the way they should. They did. Full-RAG dominated retention because it has perfect memory by design. Recency-only RAG fell apart on retention, also by design. The frozen model sat where a frozen model should sit. If that ordering had been wrong, every chart after it would have been fiction, so I was glad to spend a day confirming something boring.

Then I started injecting corruption, because that is what real reward signals do. In M2a I added a 40% corruption rate to a fast parametric learner running 24 synthetic founder personas across 12 weeks, and compared naive training on everything against T1 filtering by reward.

ArmWeek-12 retention
frozen0.412
naive0.591
t1_only0.999

The filter looked like a miracle, with a paired effect size of d_z = +5.20. Naive did not even fall below frozen, which was itself a small correction to what I had predicted, but it was clearly worse than the filtered arm. And that near-perfect T1 number was actually a problem for me, not a win. The rewards in this sim were true, with no noise, so the filter let through only clean gold, which left no room at all for the second gate to matter. I had accidentally built an experiment that could not answer my real question.

So in M2b I broke the filter on purpose, the way it is actually broken in production:

reward_est = reward + N(0, 0.5)

Now about 12% of corrupted episodes slipped through T1. T1-only turned into a weekly loop that always promotes its challenger. GATE only promoted when the challenger beat the champion on held-out retention. This time the curves separated cleanly.

ArmWeek-12 retention
frozen0.412
naive0.555
t1_only0.857
gate0.951

GATE beat T1-only by Δ = +0.095 at d_z = +1.01, past the minimum effect I had pre-registered. When I looked at why, the mechanism was exactly the one I had hoped for. GATE promoted a challenger about three times per persona across weeks 5 to 12. T1-only promoted eight times, waving through every challenger including the ones that looked great on their own training data and quietly failed on held-out facts. The second gate was the bouncer, turning those away. This is the world Founder OS actually lives in: leaky filters, and something standing behind them.

I wanted to know when that second gate is worth having, so I swept the filter noise from clean to hopeless.

xychart-beta
    title "Value of the promotion gate as your reward signal gets noisier"
    x-axis "T1 reward noise" [0.0, 0.25, 0.5, 0.75, 1.0]
    y-axis "GATE minus T1-only (retention margin)" -0.05 --> 0.15
    line [-0.03, 0.01, 0.06, 0.11, 0.13]

At zero noise the gate is a liability. It is too cautious and rejects genuinely good challengers, so GATE actually loses to plain filtering. As the signal degrades, the margin flips and keeps climbing. Founder OS does not run at zero noise, and neither does anyone’s real life, which is the whole point.

I also checked the scarier story, whether naive self-consumption collapses on its own without any injected corruption. Over five rounds of training on its own predictions, retention barely moved, from 0.999 to 0.982. So in my simulation, collapse is driven by bad data getting in, not by some mystical decay. That is less dramatic than the model-eats-itself headline, and it is what actually happened, so it goes in the post. GATE is not a cure for every failure mode. It is insurance against bad data entering weights you keep, which is precisely what a long-running agent accumulates.

Parametric curves are clean, but Founder OS runs Qwen through Ollama, so M3 was the bridge. I trained a regression head plus LoRA on Qwen2.5-0.5B-Instruct in 4-bit, same ablation arms. The ordering flipped and then recovered.

ArmWeek-12 retention (n=24)
frozen0.412
naive0.671
t1_only0.585
gate0.673

Here the filter actually starved the model. A hungry LLM wants examples, and filtering too hard left it too few, so naive volume beat careful filtering. But GATE still beat T1-only by Δ = +0.091 at d_z = +0.55, because the gate recovered from that starvation by picking the right challengers instead of every challenger. Harder learner, same claim, still standing.

What I actually learned

I went into this trying to justify a product feature and came out with something closer to a map of the territory.

Verification is not optional if you are going to keep learning from your own trajectories. T1 alone beats naive training by a wide margin under corruption, and that is not a subtle effect.

One tier is not always enough. When your reward signal is imperfect, and in production it always is, a held-out promotion gate adds real retention on top of filtering. I want to be honest about what is and isn’t new here. Selective updates exist. Promotion gates exist. What I could not find anyone doing was isolating their causal contribution in this specific single-user, self-generated, collapse-pressured setting, and that isolation is the contribution.

The gate cannot be allowed to grade its own homework, which was my original sin. My first gate.py was worse than having no gate at all, because it handed out a certificate of confidence for wrongness. The second tier has to score on facts the challenger never saw in training, held out on purpose.

And scale changes the answer. On a small LLM, raw volume can beat filtering until the gate steps in to choose well. The stronger versions of these experiments, with a reward-weighted objective and real longitudinal data, are already wired up in research/ for the next pass.

I want to be clear about the ceiling on all of this. It is an empirical systems result, a regime map, not a new optimizer or a theorem. I am fine with that. Founder OS does not need a NeurIPS oral. It needs to not eat itself on week twelve.

Wiring it into the thing I actually use

Once the research convinced me, I put it into production. The circular gate is on its way out of the codebase, and the real loop now runs weekly against my live traces.

flowchart LR
    traces[(Flight recorder<br/>traces)] --> ep[Segment +<br/>score episodes]
    ep --> t1{"T1 filter<br/>reward &ge; 0.62"}
    t1 -- pass --> train["Train challenger LoRA<br/>Qwen2.5-3B, RTX 4090, ~40 min"]
    t1 -- fail --> drop["Discarded<br/>~38% of traces"]
    train --> t2{"T2 held-out eval<br/>retention &ge; champion + 0.02"}
    t2 -- yes --> promote["Promote:<br/>symlink into Ollama Modelfile"]
    t2 -- no --> rollback["Rollback:<br/>challenger leaves no footprint"]
    champ[(Champion<br/>adapter)] --> t2

The piece that keeps it honest is the T2 evaluation. It scores the challenger on evals/lece/retention.jsonl, which is built from odd weeks, out-of-distribution entity names, and tasks the model never saw while training. The champion gets scored the same way, and the challenger only wins if it clears the champion by 0.02. If it doesn’t, the challenger leaves no trace behind.

I have been running this for nine weeks on my own instance now. These numbers are not from the simulation, they are from me using the thing every day, and they are better than the sim precisely because the reward is real and the tasks are my actual company.

MetricPre-GATE (principles only)Naive weekly LoRA (3-wk A/B)GATE (wk 4–9)
LivedBench retention score0.440.51 ↓ (wk 3 collapse)0.81
Repeated mistake rate (same tool error within 7d)23%19% → 31% (regressed)6%
Outreach drafts needing major human rewrite68%54% → 71% (regressed)29%
Challenger promotions attempted—6/6 accepted (circular eval)14 attempted, 2 promoted
”Ghost send” near-misses (sandbox catches)2 in 4 wk5 in 3 wk0 since wk 4
CRM tool-call first-try success (recurring workflows)61%58%82%
Weekly distilled principles quarantined (stale advice)1.24.8 (noise accumulation)1.4

The naive column is not a strawman I invented to look good. I ran it on a cloned instance for three weeks before I was willing to trust GATE on the real one, because I wanted to feel the failure mode myself. Naive did not explode. It drifted. It got more confident, lost retention, and started rewriting my emails in a voice that was almost mine but not quite. By week three retention had crashed to 0.51 and I shut the experiment down.

On the production loop, GATE rejected 12 of 14 challengers. Only two made it through: one in week 5 that chained calendar and CRM actions noticeably better, and one in week 8 that got smarter about which tool to reach for when summarising for investors. Everything else looked good on its own training data and failed on held-out facts. In several of those cases I looked at the challenger, thought it seemed fine, and would have promoted it myself. The gate disagreed, and the gate was right.

What it feels like day to day is quiet, which is the whole point. My agent did not turn into AGI. It just stopped re-learning my bad Tuesdays. It holds on to the fact that I like terse investor updates, that a particular plant is still in conversation rather than qualified, that the energy-specific outreach playbook works better than the generic one, and it stopped inventing new “principles” out of tool calls that had actually failed.

That matters a lot in the work I am actually doing. When I am deep in Stamped Energy conversations with plant managers who have watched a hundred vendors overpromise on OPC-UA integration, the last thing I need is an assistant that got more confident and less careful while I slept. The principles and the LoRA adapter finally point the same direction, and GATE is what keeps them from drifting apart into nonsense. I am still running it. Week 10 starts Monday, and the research harness is still running too.

If you want to build the same thing

You can lift the idea out of Founder OS and use it on its own. The shape is small:

Traces → Episodes → T1 filter → Train challenger → T2 held-out eval → Promote | Rollback
                        ↑                                   ↑
                  reward_est                         retention_ood
                  (noisy OK)                    (never in training set)

A few things I learned the hard way and would tell anyone starting here. Never evaluate on training episodes, which was exactly my first mistake. Assume T1 will leak, because it will, and design T2 as if 10 to 15% of bad episodes are getting through. Treat sparse promotions as success rather than failure, since 2 out of 14 is the gate doing its job, not the gate being broken. Report retention and recent-recall separately, because familiarity is not the same thing as judgment. And decide your key comparison before you run anything, because everything else is just context around that one contrast.

If you want to read the actual code, the research harness lives in research/ with its own git history and pytest suite, the production gate is in agent/cognition/train/gate.py and is being replaced by the held-out eval under evals/lece/, and the nightly principle distillation is in agent/cognition/distill.py.

Why I care about this at all

I built Founder OS because I was tired of re-introducing myself to brilliant models that forgot me every morning. I built GATE because I got scared of the opposite mistake: an agent that remembers the wrong things forever, trains on its worst days, and calls that growth.

The answer was never to stop learning. Learning from my own life is still the entire advantage. The answer was to learn with a check on it: filter what goes in, gate what gets kept, and always score on life the model has not seen yet. Two tiers, one retention curve, and no near-miss email sends since week four.

The machine that lives with you should get better over time. It should not eat itself to do it.


This grows out of Founder OS and LECE. Code and docs are at github.com/officiallyutso/Founder-OS. If you try to replicate this, or if your gate rejects something and you want to argue about it, tell me. The machine is listening, and for once it will remember what you said.

--claps