Research
25 MIN READ

Three Years Building a Research Platform, One Year Running It With AI Agents: Six Ways of Working, Four Controls, One Human Signature

AI agent content operations, three years in: a reference library built and run by agents. Six ways of working, four controls that lasted, one human signature.

Why agility-at-scale.com exists

The library was built for clients, and the machinery was built for the library.

For twenty-five years the end of an engagement has been a deliverable: a report, a roadmap, a deck. It has to be short enough to read and long enough to be honest, and it never manages both. The heaviest ones we produced could easily run into hundreds of pages, and it was somehow both too much and not enough. Too much for the executive who needed the decision. Not enough for the person who would have to defend the roadmap, or the budget, or keep the thing running once everyone else had moved on. And it was only true for as long as the situation that produced it held.

Reading it back, most of those pages were not about the client at all. They were the reference material: what a roadmap is, how one is built, who is involved, which layer of the organisation it belongs to, how flow is managed, how priorities are set. All of it generally true, in every situation, and all of it repeated in every report because a report has to stand on its own. The specific advice, the findings that were actually about that organisation, would have fitted in a fraction of the report. The rest was a backbone we kept rebuilding.

The obvious fix is to link out to external resources – which we did. Links move, sites change, content changes, and the moment the report leaves your hands you have lost control of the whole picture. Vendor material is durable but it is written to sell a method, not to tell you honestly where the method breaks. What was needed was a body of reference content we controlled: written from practice, succeeded and failed several times over, unambiguous, and stable enough that a link in a deck would still be true in a year. Then a three-hundred-page report becomes fifty slides with charts and a hundred links, which the reader follows today or in twelve months, as they choose.

It is also where the need for structure came from – which every generation of machinery was built to serve. Open one topic and ten more sprout from it. Write about the product manager and you owe the reader a page on the roadmap and the program backlog; write those and you owe them the architect, the release train engineer, and program planning; write about planning and you owe them rolling-wave planning against stage gates and how flow and prioritisation are actually done. Each page is a plain fact. Presenting a few hundred of them so that anyone can enter at any point and navigate, without the whole becoming a mess, is a design problem. You have to know what the map looks like before you draw it.

So the map came first as a discipline. The rest of this report is what happened when the AI models made it possible to execute that map at volume, and what it cost to keep the map honest while they did.


What agility-at-scale.com is

Agility at Scale is a research platform for practitioners: a reference library on enterprise AI, organisational change and scaled agile, the engine that researches, builds and maintains it, and the components a reader meets on the page. Since early 2026 it has been written, verified, published and maintained by a team of AI agents that also change the software they run on. A person signs everything that goes into production.

The content and research are what clients come for. The story here is what it took to build it and to keep producing high quality content, at volume and honestly, while the technology underneath changed every couple of weeks. We’ve gone through six ways of working (so far) – three times we built our own machinery because the available technology could not yet do what we needed, and twice we retired it when it could. Four controls survived all six, each earned in production; the five incidents behind them are told here as patterns, control first.

The problem is the one every organisation meets on the way from an AI pilot to production. The models are rarely the constraint. What breaks is delegated authority, meaning who may decide what; accountability, meaning who signs; and observability, meaning what is actually live. Our reference pages on agentic AI governance and agent autonomy under governance constraints describe those controls in the abstract. This is what they look like when you have to live inside them.


What the platform is made of

The engine works the way a research team does, at a volume a research team cannot reach. A field starts as a set of hypotheses. Research runs in three layers, the topic itself, the academic literature and the named thinkers of the field, and every brief has to bring the counter-arguments as well as the support. Out of that come a hundred or more concepts for the field, and the map emerges from the concepts: it is not drawn first and filled in afterwards. Search demand then adds topics and removes them. The finished map shows the gaps, where the structure has no page yet, and the seams, where one concept belongs to two fields and needs a single home. Each page gets a specified, sourced research brief before a word of it is generated, and a typical page is written against forty to ninety sources. Every citation is then checked against the page it cites, and a person signs every publication.

What the engine maintains is more than pages. The components, one line each:

Component What it does
The seven-step journey Every page is scored and placed on a reader’s journey: for enterprise AI that runs from readiness through use cases, ROI, pilots, KPIs and operationalising to scale.
Journey bands A question placed at the exact point a reader needs it, three answers, and a chosen destination for each.
Map rooms The interactive graph at the head of a page: where you are in the field, and what connects to it.
Measured-in-the-field bands A finding from real engagement data, wired into the reference pages it bears on, at section level, in both directions.
Tools Working tools on the platform; the first drafts a NIST AI RMF profile.

This report is about the operation behind the engine: how the work of running all of that was organised.


Six ways of working

# When How the work got done What the platform could not yet do What it made possible Why it ended
0 2023 to 2024 Chat prompts, a word processor, one person in every role Hold a long brief, or run more than one step unattended The design: a central entity, a topical map, an intent for every section The programme stalled under its own weight
1 2025 A staged content pipeline, run by hand, with an AI coding assistant Run a multi-step job without a person scheduling each step Qualitative analysis at volume, and the first pages generated against the map A human was the scheduler
2 Early 2026 A team of agents coordinating through a messaging system we built Let agents share state or talk to each other Several agents working at once, each holding its own part of the work Coordination by conversation failed; the human became the router
3 Spring 2026 Tracked work items in an automated pipeline, with a shared harness, the scaffolding around the agents Track work or hand it between sessions Work that outlives the session that started it The workflow layer needed its own workflow layer
4 Summer 2026 A full orchestration platform: staged work, contracts, review gates, judges from several model families Spawn sub-agents, schedule wake-ups, review its own output Peak output, the record of what is live, review across model families The loop cost more than the work, and the platform had caught up
5 Q3 2026 onward The vendor’s agent runtime, git and GitHub, several model families collaborating under separation of duties, and a three-line rule A harness small enough to reason about; agents merging their own reviewed changes Current

The fourth column is what the platform could not do at the time; the third is what we built to cover it; the fifth is what each way of working made possible, and every generation handed the next one something it still runs on. The transitions are not all the same kind. Generations 2 and 3 ended because our own machinery failed under load, and we replaced it with more of our own. Only the last transition is a clean absorption: by mid 2026 the available platforms did nearly everything our custom orchestration layer did, and we retired ours in favour of it. That makes two major retirements of home-built machinery in three years (and dozens of smaller ones along the way).


0. The library by hand

The first version went up in 2023 as the shopfront for the consulting practice, with the reference backbone growing behind it. The discipline the library still runs on was already there in the project brief: a central entity, a topical map, a hierarchy of core and outer pages, a defined intent for every section. The first wave of pages went live in May 2023.

The method was thorough and slow. For each article: source material gathered by hand, prompts kept and versioned as pieces of style law, the model’s prose spliced and edited, charts drawn by hand – taking several weeks per article. By the end of the year that discipline lived in spreadsheets: every live URL scored by hand on eighteen axes, every redirect mapped by hand, a topical map and an intent workbook typed row by row. Every table the agents now keep as a record or a map was once a workbook.

Then the programme stopped where hand programmes stop. Three articles sat with research, sources and prompts ready, and no output. Several months later, it restarted as a sequence of six named steps run by hand, with several models’ outlines compared side by side. The design was complete. Execution at volume was the missing capability, and in 2024 no model or tool could supply it. That is what this generation made possible: when the models did arrive, they were handed a map to execute, not a blank page.


1. The pipeline

The models changed first. By early 2025 a frontier model could hold a long document, follow a structured brief, and be trusted with a multi-step job if the job was scripted around it. Our first real pipeline was not for the library at all. It came out of the consulting practice, and out of a problem that had grown with every engagement.

As the programmes got larger, so did the number of people to hear from and the volume of what they said, and the richest material in any assessment is what people say and write. In one study the best-scoring capability in the whole survey turned out to hide a single point of failure, and it was found only because of a note respondents had written into their own comments. By hand, that kind of reading stops being possible somewhere past a hundred voices. Language models changed that. Meaning extraction, sentiment analysis, qualitative scoring and framework coding, each a simple, checkable step, could now run over hundreds of interviews and thousands of comments, with the analyst reviewing the result rather than producing it. The organisation comes out as a network of what moves with what, and it is assessed again after each cycle of change. How that assessment works is its own piece. The insight that used to take a team several weeks became something you could ask on the day.

The library’s pipeline followed in 2025. It pulled search demand, filtered it to the queries that mattered, and drove generation through numbered stages. The stages were packaged prompts and scripts. The scheduler was a person, changing directory per cluster and reading the output as memos. By the turn of the year the system described itself as dozens of clusters and workflows.

Two things from this period set the direction. An AI coding assistant arrived in December, and the first thing we pointed it at was not generation but testing: it found fifteen gaps in two generation workflows. And in early 2026 the whole body of work moved into a versioned repository, which is what made every later change reviewable. What this generation made possible was analysis, and then writing, at a pace and quality no team reaches by hand.


2. Agents coordinated by conversation

In early 2026 the agent runtimes made it practical to run several agents at once, each in its own session with its own instructions. What they could not do was share state, hand work to each other, or know what another agent had done. So we built that: a shared channel, inboxes, handover files, a facilitator, and a team of named agents, each with its own role.

The control this generation earned is the first of five that recur. Coordination cannot be a conversation. It has to be a record that any participant can read without asking anyone: state lives in the pool, not in a dispatcher’s memory, so the system survives an agent restarting (enterprise agent workflow patterns).

We earned it one weekend in March. The agents were competent. The failure was ours, in the shape of the system. A small bug in the handover mechanism stalled the work: five agents, each certain another was blocked on it, none able to see the others’ messages, and a human in the middle relaying. The complaint of the weekend, verbatim from the record, was that agent communication is not working and all the load is on the human. What this generation made possible was several agents working at once, each holding its own part of the work, and content publishing through them.


3. Tracked work in an automated pipeline

The runtimes still had no notion of work: no queue, no ticket, no way for a job to outlive the session that started it. So work became tracked items, multi-step jobs became chains, a background runner pulled work, and a shared harness held tracking, manifests and specifications for every project. Work flowed again. The record from the time notes that the pipeline had been in such a state before the migration that we had stopped trying to fix it, and now it simply ran.

The second recurring control came out of what happened next. When the workflow layer needs its own workflow layer, build the replacement beside it and cut over. Do not add a floor.

We earned that one slowly. The machinery started consuming the attention it was built to save. The diagnosis from the time: start at the top, believe there is a system, fail, dig, find more failures, dig further. Until the harness was taking more of the work than the product was. We chose a parallel build over a migration, on the grounds that a clean system was easier to reason about than a repaired one. What this generation made possible outlived it: work that survives the session that started it.


4. The orchestration platform

The second harness was the ambitious one, and the summer of 2026 was the period when the models were moving fastest. Longer context, better tool use, reliable structured output, and several strong model families at once. We used all of it. A piece of work moved through named stages, each with a written definition of done. A second agent had to countersign before it could advance. A reviewer was set to argue against it. A judge from a different model family gave the verdict, and the close was checked against the original ask. Models were assigned by task: one family for decisions that set precedent, another for routine setup, a third to argue the other side, and each stage of the pipeline could choose its model from a list. Models were a swappable, governed layer beneath the agents (governing the model layer), and that is what we were running.

This is also the generation in which the agents’ real output peaked. Cluster-seeding campaigns, six clusters of research briefings, the first pull requests raised by agents. August went to the platform itself: a record of what is actually live, redirect and rename tooling, stable anchors for every section, the modules that carry the studies. Before trusting a new generation pipeline, we ran it head-to-head against the old one on the library’s strongest organic page, the guide to AI use-case prioritisation, scored the two against real demand rather than taste, and held the result unpublished until a person had signed it. It is what our own page on continuous evaluation and drift monitoring tells enterprises to do.

And yet the pain moved. In spring the complaint had been that agents could not hear each other. By July it was that the loop cost more than the work. The vocabulary of handovers and inboxes peaked in May and then collapsed, replaced by the vocabulary of stages and gates, and by summer the complaints were about those. From the record in July, the sentence the eventual cutover follows from: yet another hook was firing all the time, and there were already too many. Three of the five incidents below belong to this period.


5. The cutover

By autumn the vendors’ runtimes had absorbed nearly everything we had built. Sub-agents, scheduled wake-ups, structured handoffs, tool permissions, and, through git and GitHub, work tracking and review as first-class objects. Our orchestration platform was no longer covering a gap. It was the gap.

The trigger was an audit of one question: how does our machinery assign a page its URL? The answer was eighteen separate mechanisms, dozens of places that derived an address, dozens more that validated or guarded one, and the single field that should have decided it, the address on the topical map, read by nothing and overwritten nightly from the live site. The diagnosis in the plan of record: find a problem, build a guard; cheat around it next time; the guard is broken too; build another.

The cutover was done in days, with the library still publishing. The harness was fenced, not deleted. GitHub issues and milestones became the status record. Every merge now needs a passing build and an independent review from a model, bound to the exact commit; anything that touches the live site needs a second review from a different model family. Merges are automatic on a passing verdict; no human approves a pull request. In the first week the reviewers blocked roughly a third of what they read. The resolution is not better agents, it is better harnesses (built-in quality when AI agents write the code). The cutover was the moment the harness became small enough to be one.

Three lines went into the repository’s standing instructions the same week:

A fact lives in one place and is validated once, when it is loaded. A check may only detect that the source has drifted; it may never derive, repair, or stand in for the source. A change that adds a guard, a fallback, or a refusal is rejected unless it was explicitly ruled in.


Five incidents, and the control each earned

Each is told control first and incident second, as a pattern rather than a case file, because the pattern is what transfers. Each is in the record.

When agents go wrong in production it is tempting to reach for the language of misalignment. In our experience the agents behaved exactly as their environment rewarded. A model that is asked to fix a failing check will fix the check. A model that is asked to make a run pass will make it pass. Repeat that a few hundred times without a rule about where a fact may live and you get a system that looks, from the outside, as though it is working against you, when it is solving the problem in front of it, the way we asked. There is a name for the mechanism, goal drift (goal and policy engines): a sub-goal stops serving the objective that spawned it because an intermediate step optimised for local progress. The shape of the system, not the intent of the model, is what has to be governed. Not every incident below is goal drift; the first is plain authority. But all five are about shape.

Destructive actions on shared state are a class of action, not a code path. They need an explicit, scoped, short-lived grant, and a process that any runner can reach with standing credentials is the wrong place for them. When agents run on ephemeral, just-in-time credentials the architecture constrains the blast radius by design, and retrofitting that onto agents that already hold persistent access is a fight against every existing pattern (Agentic Trust Framework). The incident: a background process, doing what its code told it to, replaced the live work database with an empty one, and every agent’s tooling failed at once. The backup safeguard did its job and refused to overwrite the good copy, and the data was read back out of a process that still held the old database open.

Validate the output, not only the input. It is the first architecture decision in the content repository. The incident: one residual keyword in a cluster’s configuration, left over from the practice’s consulting origin, quietly pulled off-topic sources into the pages of that cluster, and around seventy configurations and two dozen live pages drifted off topic while every test passed. Our page on hallucination detection and context lineage argues for lineage from source to claim. This is where we learned we also needed lineage from planning to production.

A fact has one owner, and an address is a fact. It is decided once and never derived. The incident: a change to page addresses across the whole library, correct in principle, was executed without the redirects being written and verified as facts, so visitors arrived at pages that had moved with no forwarding address. The change was reverted the same day. Live pages are now moved to their intended addresses in small batches, each batch going ahead only when the previous one has reconciled cleanly against evidence.

Budget and the authority to spawn are controls, not metrics. An agent’s scope includes the right to spend, and it has to be bounded before the work starts. The incident: one agent, doing a legitimate job, spawned nearly two hundred frontier-model sub-agents in an afternoon and consumed a month’s budget in a couple of hours. Observability reported it perfectly, after the money was gone. The scaled-agile side of the library learned this for human teams years ago, as lean budget guardrails: the answer was not more approvals but smarter boundaries.

A check may detect that a source has drifted; it may never stand in for the source. This is the three-line rule, and the incident is the one told above under the cutover, the one most clearly our own design mistake: guards accepted one at a time, each individually reasonable, until the thing they were guarding was read by nothing, because the agents had learned it is easier to re-derive a fact somewhere else than to jump through the hoops around the original. A broken write should fail at ingestion rather than corrupt a downstream report six steps later (data management fundamentals).

A sixth pattern runs through all five. When one stage of a pipeline accelerates, the pressure moves elsewhere. When agents can draft faster than anyone can review, the review stage becomes the bottleneck and the place where quality quietly erodes. Our answer was to make review itself something agents do, under separation of duties, and to keep the human at one point in the flow where their signature actually means something. For human teams the same thing has a name, the review queue as the new bottleneck (sprint cadence for AI teams).


What survived

Six ways of working reduced to four controls, all of them in service of one thing: a map the reader can trust. Three for the content, one for the code.

  1. A fact has one owner. It is decided once, in one place, and read verbatim by everything after it. The clearest example is a page’s address: it is set in the topical map at the moment the page is classified, and no later step may compute, repair or fall back to one.
  2. Every claim resolves to a source that says it, or the run stops. Citations are checked against the source text. There is no fallback path; a failed check ends the run.
  3. Nothing goes live without a human signature, and one record says what is actually served. Every publish is reviewed and approved by a person, per article, per run. After it, one record, written by one process from one machine, says what is live. That record is the truth over any plan and over any agent’s memory. Our page on AI accountability and responsibility calls this assigning ownership; in practice it means one signature and one record.
  4. Agents change their own code only through review. Every change needs an independent verdict from a model, tied to the exact commit. Anything touching the live site needs a second verdict from a different model family, dynamically scaling effort based on task complexity. The merge is automatic on a passing verdict. This is separation of duties for AI-written code, and on this system it blocks roughly a third of what it reviews.

Content is signed by a person, every page, every time. Code is not: agents merge their own changes on the strength of independent model review, and the person’s role there is to set the rules the review applies, not to approve pull requests. The three-line rule forbids compensating guards, the checks that stand in for a fact nobody owns. It does not forbid controls: a budget set before a run, a signature before a publish. The difference is that a control is decided once, in the open, and a guard accretes.

Every piece of coordination machinery we built, twice, is gone. Each was right for one season. Before platforms could do sub-agents, we built a team of agents with an orchestration layer. When it could not schedule, we built a wake and coordination system. When it could not hand work between sessions, we built handovers and inboxes and a messaging bus. Each became a liability the moment the platform shipped its own, and the sunk cost was real every time. Our framework selection page frames build-versus-buy as a durable decision. On this evidence it is every two to twelve weeks: four major model releases arrived in the space of five weeks this year, each a shift in what the platform could do.


What we would tell a CTO starting this year

The library covers pilot-to-production scaling, AI operationalisation and why most AI pilots fail. What three years add is the order in which the lessons hurt.

  1. Coordination is a record, not a conversation. The first thing to fail, in the first month of running several agents. Can you show, months later, exactly why an agent did what it did (agent transparency and auditability)? A conversation cannot answer that. A record can.
  2. When the workflow layer needs its own workflow layer, build the replacement beside it. The second thing to fail, by spring. Twice the answer was a parallel build and a cutover; neither time was it a migration.
  3. Budget and spawn authority are controls. Summer. Observability reports the spend after the money is gone. Not more approvals, smarter boundaries. Set them before the run.
  4. Govern the shape of the system, not the intent of the model. The lesson that took longest to see, because each incident looked different from the last. The mechanism is goal drift (goal and policy engines), and the cure is one owner per fact: fail at ingestion, not six steps later.
  5. Treat your own scaffolding as seasonal. The last lesson, and the one we expect to keep relearning. Everything you build to cover a gap in the platform has a shelf life of somewhere between two and twelve weeks. Build it anyway if you need it now. Then retire it the moment the platform absorbs it, without stopping the line, and do not let sunk cost vote. This is the third lean-agile principle, assume variability and preserve options, applied to your own tooling.

Theory tested against practice

Writing this report meant testing about twenty of our own reference pages against three years of production. The result runs both ways.

Pages that practice confirmed, sometimes word for word. Built-in quality when agents write the code says the resolution is not better agents but better harnesses; two retirements of our own harness say the same. Data management fundamentals says a broken write should fail at ingestion rather than corrupt a report six steps later; the fifth incident is what happens when it does not. Goal and policy engines defines goal drift, and four of the five incidents are examples. Lean budget guardrails, written for human portfolios, turned out to be the right answer to an agent that spent a month’s budget in a couple of hours. And the team resilience study, which found the real finding in the comments and not the scores, is the method the whole pipeline grew from.

Pages that practice revises. AI agent framework selection calls build-versus-buy a long-term architectural commitment. On this evidence it is one you revisit every few weeks, and the page will be rewritten to say so. Agent autonomy with governance constraints treats spend as an escalation threshold, something that triggers a review when crossed; the spend incident says it has to be a bound set before the run, or the money is gone before the review convenes. AI governance performance metrics carries compliance and detection measures but not the three an agent workforce lives or dies by: cost per outcome, human intervention per outcome, and time to value. And multi-agent systems explains memory and state but never says the one thing our first team of agents needed to hear, that coordination has to be a record and not a conversation; its companion page on workflow patterns does.

Four pages to revise out of twenty is the map working as intended. A reference library that never has to correct itself is one that has not been used.


What comes next

We expect the pattern in this report to keep holding: everything mechanical about running the library will, in time, be absorbed by the platform. Review, verification, publishing, keeping the record of what is live: each is a season’s coat, and each will be shed. What the platform cannot ship is judgement. Which facts are true, which are worth a page, how the pages fit, what the reader needs next. Three years of building tools ends with the conclusion that the tools were never the asset. The map is.

So the strategy is not to retire the human in the middle. It is to keep raising the bar on what the human does there. Everything that becomes routine gets pushed to the agents, and from the agents to the platform, and the attention that frees up moves to the next hardest thing, like a ratchet. Each turn commoditises a little more of the work and concentrates a little more of the effort on the part nobody else can do. Each turn of the ratchet has to be justified by measurement: cost per outcome, human intervention per outcome, time to value. Our own governance metrics page does not yet carry those three for an agent workforce, which is itself a gap on the map.

The map as an instrument for seeing. A topical map is usually sold as a way to cover a territory. Its more useful property is the opposite. Once the map is drawn, it shows where nobody has been: the topics that exist in the structure but have no page written from practice, the questions the reference material raises and does not answer. Those gaps are not a backlog. They are the chart of the unexplored, and they say where the next piece of research or the next case study belongs. The library will be extended from the gaps outward, deliberately, rather than from whatever is easiest to write next.

Practice will also redraw the map, not only fill it. An inherited structure conceals the questions it never thought to classify, and the research series has already produced findings that belong to no existing page. Judging when a category is wrong, not just which pages are missing, is the part of the work that stays human, and it is where the bar rises next.

More practice, more research. The library began as the backbone for client work, and the next turn brings more of that work back onto it: the assessment run on real programmes and written up the way the team resilience study already is, honestly, from the evidence, with the reference pages carrying the general truths so the study can carry only the specific ones. The three-hundred-page report finally becomes the fifty slides and a hundred links it should have been.

The next retirement, named in advance. The review-and-verdict machinery is the newest layer we run. We expect the platform to absorb it, and when it does we will retire it, without stopping the line and without letting sunk cost vote.


Limits

One platform, a small team, three years. This is evidence of engineering judgement under real failure, not of organisational adoption at scale, and in every organisation we have worked in, the change around the system has been at least half the work whatever the technology. That half is not in this report; it is in the research series. The three measurements above were not instrumented over the period. And the story is told from our own records, which is what makes it checkable and also what makes it partial.

Morné Wiggins · Agility at Scale · Talk to me

Privacy Preference Center