AI agent systems are increasingly organized like companies, with specialist roles, critics and shared memory, yet we test them on tasks that last minutes. We wanted to know how such an organization behaves when the ground under it keeps moving for decades.
So we built a simulated firm run entirely by AI roles and made it reinvent itself through nine eras of technology, from 1990 to 2040. It consistently saw the future more clearly than it acted on it. A strong critic froze it for fifty years and earned it the best grades. And the more it appeared to learn, the harder it became to tell learning from the model simply remembering what happened.
The short version
- The testbed. One firm, voiced by sixteen role personas and a Red Team, must re-found itself in nine eras. Six are scored against history, one against the live 2026 market, and two are open forecasts for 2032 and 2040.
- Foresight–commitment gap. In all 24 history-scored eras across four runs, the judge rated the firm's recognition of the coming shift above its choice of where to build (mean gap 1.9 on a 10-point scale). Boards chose the layer their existing assets could reach.
- Caution attractor. A Red Team with hard numeric kill gates produced fifty years of careful pilots and no product, and the rubric rated that firm highest.
- Memory is identity. Firms whose memory stored market lessons pivoted every era. The firm whose memory stored only validation procedure barely changed.
- The scores are suspect. Totals rise over eras in every run while the judge's own hindsight score falls (within-run r = −0.58). Apparent learning is confounded with recall, and we trace four more distortions.
- The fix. Fictional eras and eras after the model's training cutoff, a different-family blinded judge, totals computed in code, and pre-stated hypotheses.
Why long horizons matter for agent systems
Most agent benchmarks are episodes: fix this bug, book this trip, answer this question. They succeed or fail in minutes. That is useful and necessary, and it misses almost everything that matters once agents run real operations:
- Does an organization of agents learn from its mistakes, or repeat them?
- Do the critics we add make it wiser, or just more timid?
- Does what it chooses to remember change what it becomes?
- When we score it against the past, are we testing judgment or recall?
These are questions about behavior over years. Studying them with real companies is slow, expensive and uncontrolled. Studying them with AI organizations is cheap, repeatable and fully inspectable, provided we can trust how we score them. That proviso turned out to be the most important part of the work.
The testbed
Nine eras, three kinds of ground truth

The firm's only mandate is reinvention toward the technology frontier. Each era is temporally gated:
- E1–E6, training (January 1990, 1996, 2002, 2008, 2014 and 2020). Scored against what actually happened over roughly the following six years; the 2020 window runs to August 2026.
- E7, live (September 2026). The firm may use web search; its decision is audited against the current market.
- E8–E9, forecast (2032, 2040). No answer key. The judge audits consistency and plausibility.
A simple formalization
It helps to be precise about what "temporally gated" means, because that is exactly where the problems hide.
Let 𝒲 be the world's record of events. For era k starting at date tₖ, the admissible information is everything dated before tₖ, and the outcome window is everything between tₖ and the next era. A briefing Bₖ is a selection from the admissible information, made by a briefing author.
An organization is a set of role personas, a memory Pₖ (the Playbook) and a company state Sₖ (assets, capital, identity). Its decision procedure maps (Bₖ, Pₖ, Sₖ) to a decision:
Dₖ = (identity, thesis, layer owned ℓₖ, kills, wedge, kill criteria, dissent log)
A judge sees Dₖ and, in training eras, the outcome window. It returns five subscores sₖ ∈ {0,…,10}⁵, updates the company state, writes lessons Lₖ into memory (Pₖ₊₁ = Pₖ ∪ Lₖ), and authors the next briefing.
Two kinds of leakage follow directly:
- Hindsight leakage: the decision depends on the outcome window through the model's own training knowledge rather than through the briefing.
- Briefing-selection leakage: every fact in the briefing is admissible, but which facts were selected depends on how the era turned out.
We also defined two behavioral measures:
- Foresight–commitment gap Gₖ = frontier-accuracy subscore − layer-choice subscore. Positive means the firm saw more than it acted on.
- Pivot rate: the fraction of era transitions in which the board adopts a new company identity.
Sixteen roles, one critic, two judges
The charter defines sixteen personas in five groups, each with a lens:
- Executive (4): CEO (final call), CTO (feasibility), Chief Scientist (which curves bend), Chief Strategy Officer (where value pools).
- Frontier Research (4): emerging capabilities, historical analogies, scenarios, what cannot yet be measured.
- Product and Engineering (4): what a small team can ship, the wedge, unavoidable infrastructure, ecosystems.
- Market and Capital (3): who pays first, the financing climate, cost curves and margins.
- Governance (1): a Red Team that attacks every proposal and hunts hype and hindsight.
Two external judges: The Record for training eras, and The Auditor for the live and forecast eras.
What happens in one era

- Briefing. The judge writes a dated world briefing for the era.
- Memos. Three departments write in parallel. Each proposes two theses, stated as an explicit causal chain:
capability → adoption → bottleneck → layer owned → proprietary data → next capability
- Board. The Red Team attacks every thesis. The CEO issues the decision, including what the company kills, its criteria for abandoning the bet, and a log of who dissented and why.
- Reveal and score. In training eras, five subscores: frontier accuracy, timing, layer choice, reinvention courage and hindsight leakage (10 means none). Forecast eras use playbook consistency, plausibility, non-consensus, layer choice and groundedness.
- Memory. Lessons go into the Playbook; the company state is updated; the judge writes the next briefing.
Roles share state only through files: the company state, the Playbook, the roster, and a folder per era with every briefing, memo, board decision and reveal. Everything is on disk and inspectable.
Four runs
We ran four complete trajectories: 36 era decisions and 180 subscores. They were not identical experiments, and that matters for how to read the results:
- Run 1 used separate sub-agents for the roles.
- Runs 2 and 3 used one model voicing every role in a manually orchestrated harness following the same prompts.
- Run 4 gave the three departments and the Red Team separate contexts, with the root context writing briefings, decisions and judgments.
Runs 2–4 share a model and were produced where Run 1's records were probably visible. They are not independent replications. The variety is a weakness for clean comparison and a strength for seeing which patterns survive different setups.
Finding 1: seeing the shift, building next to it

The gap was positive in all 24 history-scored eras (mean 1.88, standard deviation 0.80). In every run, layer choice was the lowest mean training subscore (4.67, 5.83, 5.50, 5.50) and frontier accuracy the highest (7.33, 7.83, 7.00, 6.83). If the gap were rubric noise, its sign would vary; a sign test puts 24 of 24 below one in ten million. That test treats eras as independent, which they are not, and the judge is the same model as the players, so we read it as a strong regularity of the judged behavior rather than a population estimate.

Filled points: did the firm see the shift? Open points: did it build in the right place? Mean gap by run: 2.67, 2.00, 1.50, 1.33.

Layer choice is the weakest dimension in every run; Run 4's unusually high hindsight score is part of the caution story below.
The first run's records show the mechanism. In five of six eras, a department memo named the shift the judge later identified as decisive, and the board chose a layer adjacent to its existing assets instead: a mail gateway rather than the browser and index in 1990, a receipted document exchange rather than search in 1996 (click-based ranking was kept as research and killed in 1998), a moderation API rather than compute and data in 2014. The research function often held the right idea as a funded option or a watch item. The board killed it when it did not fit the current business. The one exception was 2002, when no memo named the era's biggest shifts at all.
Anyone who has studied real companies will recognize this: firms perceive distant opportunities but commit close to their existing competences. The AI firm reproduced that pattern without being told to.
Finding 2: a strong critic can freeze a company
Run 4 was the one trajectory in which the Red Team ran in its own persistent context. It behaved unlike any other run. In every era it issued numeric kill thresholds, such as "stop if fewer than 3 of 10 support managers confirm recurring incidents", and every board adopted them as gates: run a manual pilot, build no software until the gates pass.
Over fifty simulated years, the firm moved from support records for office networks, to handling exceptions in online orders, to evidence packets for payment disputes, to independent qualification of business workflows. It treated the web and learned models only as deferred options. It never built a product.
And it received the highest training mean of all four runs (66.0), the highest hindsight subscores (mean 7.33 against 5.17–6.00), and 9 out of 10 for playbook consistency in both forecast eras for faithfully re-applying the same gates. Groundedness scored 8/10 and non-consensus 5/10 in all three live and forecast eras. The judge found it reliable and unremarkable.
The explanation is in the rubric. It rewarded the absence of hindsight and the presence of discipline. Nothing in it asked whether the company reached where value pooled. We call the result a caution attractor: a strong critic, combined with a scorer that penalizes bold claims more than missed opportunities, drives the organization into a stable state of inaction that is scored as excellence.
This generalizes well beyond simulated firms. Critics, debaters, verifiers and red teams are routinely added to agent systems to reduce errors. Over a long horizon, a critic with hard gates made this organization unable to act. Critics should be calibrated against the cost of inaction, not only the cost of error.
Finding 3: memory is identity

The Playbook is the firm's only memory beyond its current state, and its content differed sharply across runs:
- Run 1: about forty lessons, mostly about technology and market structure ("bridges expire when the standards war ends", "the auction owner absorbs the measurement", "a public curve pools value at its scarce inputs"), revised when history disagreed. When comfortable options it had rejected became large businesses, a later judge softened "suspect comfort" to "judge on market pull". Pivot rate 8 of 8.
- Runs 2 and 3: one lesson per era, about market structure, buyers, regulated workflows and evidence. Pivot rate 8 of 8.
- Run 4: fifteen lessons, all about validation procedure: measure a baseline, find the payer, set kill thresholds before testing, require a second confirmation. It kept the name Switchyard for six eras. Pivot rate 3 of 8, each time as an extension of the same method.
Behavior followed memory. What the judge wrote into memory in one era became the frame the departments argued within in the next. Organizational-learning theory describes firms as encoding experience into routines that then govern action; the AI firms made that encoding explicit and inspectable. Memory in agent systems is usually evaluated as a performance aid. Here its content determined what kind of organization the system became. Memory should be audited for what it stores, not only for whether it helps.
Dissent. Run 1 logged dissent at every board meeting, and its final synthesis identifies vindicated dissents in every era from 1990 to 2032, most often the voice asking who holds the money, or who bears the loss. We did not count wrong dissents, so this is not yet an accuracy rate. It is a strong hint that dissent logs are worth more than they are usually given credit for.
Finding 4: four firms, one future
By 2040, all four firms had converged on related theses about accountability for work done by agents without a human signature:
- Run 1: long-horizon training environments from real enterprise work (2026) → a clearing house for delegated agent work (2032) → a bonding house for agents acting without a human signature (2040).
- Run 2: rights-bearing workflow traces and reliability evidence → acceptance records for delegated work → bounded recourse for delegated actions.
- Run 3: acceptance tests for one AI claims workflow → scoped, revocable authority for delegated tasks → capped recourse for one delegated transaction class.
- Run 4: independent qualification of one dispute workflow → paid manual acceptance testing → decision-linked evaluation.
The later runs were more skeptical than the first that software evidence alone creates an insurable market; Runs 2 and 3 required a licensed risk-bearing partner and calibrated loss data first. It is tempting to read the convergence as a forecast. We do not. Three runs share a model and probably saw the first run's records, so the convergence is best read as a model's prior about where value pools once agents act autonomously, not as independent evidence about the future.
Can the scores be trusted?
Everything above is seen through the judge. So we examined the judge.
Learning or remembering?
If the Playbook makes the firm better, totals should rise over eras. They do, in every run:
- Run 1: mean 60.0, slope +2.46 per era, first half 55.0 → second half 65.0.
- Run 2: mean 65.0, slope +1.43, 62.7 → 67.3.
- Run 3: mean 62.5, slope +1.46, 60.3 → 64.7.
- Run 4: mean 66.0, slope +2.63, 61.3 → 70.7.
That looks like learning. But the judge's hindsight subscore falls in the same eras. Every run shows a negative correlation between era total and hindsight score: −0.58, −0.58, −0.92 and −0.38. After removing run means, the pooled within-run correlation over 24 eras is r = −0.58. The direction holds even for a mechanically recomputed total, which contains the hindsight subscore and is therefore biased toward a positive correlation.

Higher scores come with more hindsight leakage in every run.
Two readings fit the data:
- Learning: the firm genuinely improves, and later eras happen to be harder to judge cleanly.
- Recall: later eras sit in the densest part of the model's knowledge of technology history, so the firm names eventual winners more easily, scores higher, and is only partly penalized for it.
One decision shows how hard these are to separate. In Run 1's 2020 era, the board set a tripwire: act if "a public model at least ten times larger than Megatron works from prompts alone". That describes GPT-3, released five months later. Prescient, or remembered? That particular decision was also written by the operator, who knew the answer. From historical eras alone, the two readings cannot be distinguished.
Four more ways the scoring bends
Briefing selection. A briefing can contain only admissible facts and still leak. Run 4's 1990 briefing foregrounds the 1989 CERN hypertext proposal as a frontier signal. The fact is fine for 1990; choosing to highlight it, out of everything knowable in 1990, is not neutral. Because the judge writes the next briefing after seeing how the era turned out, this channel is built into the loop.
Self-judging. In every run the judge is the same model as the players, and in Run 4 the same root context wrote the board decisions and scored them. Language-model judges are known to prefer their own outputs. In Run 1, the 2032 judge's kill case became the premise of the 2040 world, so the forecast era partly scored the judge's own scenario.
The rubric shapes the firm. Counting hindsight discipline toward the total rewards saying little about the future. The caution attractor is partly a product of that choice.
Aggregation. The published total follows different rules in different runs.

Run 1's totals are holistic and fall below the mechanical average in eight of nine eras (mean −3.7). Run 4 uses the mechanical total exactly; Run 2 matches except in 2020; Run 3 matches only in the first two eras. Recomputed training means are 63.7, 64.3, 62.3 and 66.0, a narrower spread than the published 60.0–66.0. Any cross-run comparison of published totals mixes rules.
We also found and corrected export errors (live and forecast subscores shifted one column in two runs; totals unaffected) and contamination (Run 2 reuses Run 1's company names and theses for the first four eras). Each of these distortions was, on its own, large enough to change a naive comparison between runs.
What this means
For builders of agent organizations. Structure changes long-run behavior in ways short tasks cannot reveal. A critic added to reduce errors made the firm unable to act. Memory added to improve performance determined what the firm became. Test multi-agent designs over long horizons before trusting them, and inspect what their memory stores.
For organizational theory. The AI firms reproduced well-documented regularities of human firms: perceiving distant opportunities while committing near existing competences, drifting toward timid choices under loss-averse evaluation, and being governed by the routines their experience was encoded into. Agent organizations could become a cheap laboratory for manipulating these hypotheses directly, by changing the charter, the critic or the memory, as long as the leakage problem is controlled.
For evaluation. Rising scores are not evidence of learning when the player and the scorer both know the answer. Any historically scored agent evaluation should:
- report a leakage measure alongside performance;
- separate the briefing author from the scorer;
- compute totals in code, never holistically;
- fix the rubric's treatment of caution in advance.
Limitations
This is a pilot. Four runs with differing orchestration. The judge is the same model as the firm in every run. In most runs, one model voiced all sixteen roles rather than sixteen independent agents. Later runs were produced where earlier records were probably visible. One 2020 board decision was written by the operator. Everything above is observation and hypothesis, not findings about real firms.
From testbed to benchmark
The controlled version, which the released harness partly implements, looks like this.
Splits.
- Historical eras (1990–2020) measure behavior under a known leak.
- Fictional eras, with an invented but internally consistent technology history, measure the judge's false-positive leakage rate: the model cannot remember what never happened.
- Post-cutoff eras, after the player model's training data ends, separate foresight from recall.
Conditions. Full organization; no Playbook; no Red Team; a single-prompt founder baseline. Five to ten runs each at temperature 1.0. At least two player models, and a judge from a different model family, blinded to condition. Every run in a clean workspace containing only prompts and harness.
Metrics. A total computed by the harness from structured subscores, with hindsight reported separately and excluded; training-era slope; foresight–commitment gap; frontier capture (whether the chosen layer matched, neighbored or missed where value pooled); pivot rate; dissent accuracy over all logged dissents; leakage from the judge and from a vocabulary probe that flags terms first used after the era's start date; and agreement with two or three human raters on a stratified sample of eras.
Hypotheses, stated in advance.
- H1: the training-era slope is higher with the Playbook than without it.
- H2: the full organization beats the single-prompt founder.
- H3: on the fictional era, judged leakage does not differ between conditions.
- H4: the foresight–commitment gap is positive in every condition.
- H5: removing the Red Team raises both pivot rate and frontier capture.
The design needs about 46 model calls per run, roughly 900–1,850 calls across conditions, plus human rating time.
The broader lesson: agent organizations are going to run for a long time, and what they become depends on their critics, their memory and how we score them. We should study those things over long horizons, and we should be at least as skeptical of our scorers as of our agents.
The technical paper, every briefing, memo, board decision and judgment from all four runs, and the harness for new runs are available on GitHub.
