Three Years Building the Site, One Year Running It With Agents: Six Ways of Working, Two Retirements, One Human Gate
Three Years Building the Site, One Year Running It With Agents: Six Ways of Working, Two Retirements, One Human Gate
Why agility-at-scale.com exists
The site was built for clients, and the machinery was built for the site. That order matters.
For twenty-five years the end of an engagement has been a deliverable: a report, a roadmap, a deck. It has to be short enough to read and long enough to be honest, and it never manages both. The heaviest ones we produced could easily run into hundreds of pages, and it was somehow both too much and not enough. Too much for the executive who needed the decision. Not enough for the person who would have to defend the roadmap, or the budget, or keep the thing running once everyone else had moved on. And it was only true for as long as the situation that produced it held.
Reading it back, most of those pages were not about the client at all. They were the reference material: what a roadmap is, how one is built, who is involved, which layer of the organisation it belongs to, how flow is managed, how priorities are set. All of it generally true, in every situation, and all of it repeated in every report because a report has to stand on its own. The specific advice, the findings that were actually about that organisation, would have fitted in a fraction of the report. The rest was a backbone we kept rebuilding.
The obvious fix is to link out to external resources – which we did. Links move, sites change, content changes, and the moment the report leaves your hands you have lost control of the whole picture. Vendor material is durable but it is written to sell a method, not to tell you honestly where the method breaks. What was needed was a body of reference content we controlled: written from practice, succeeded and failed several times over, unambiguous, and stable enough that a link in a deck would still be true in a year. Then a three-hundred-page report becomes fifty slides with charts and a hundred links, which the reader follows today or in twelve months, as they choose.
It is also where the need for structure came from – which every generation of machinery was built to serve. Open one topic and ten more sprout from it. Write about the product manager and you owe the reader a page on the roadmap and the program backlog; write those and you owe them the architect, the release train engineer, and program planning; write about planning and you owe them rolling-wave planning against stage gates and how flow and prioritisation are actually done. Each page is a plain fact. Presenting a few hundred of them so that anyone can enter at any point and navigate, without the whole becoming a mess, is a design problem. You have to know what the map looks like before you draw it.
So the map came first as a discipline. The rest of this report is what happened when the AI models made it possible to execute that map at volume, and what it cost to keep the map honest while they did.
What agility-at-scale.com is
Agility at Scale is a practitioner site on enterprise AI, organizational change and scaled agile. Since early 2026 it has been written, verified, published and maintained by a fleet of AI agents that also change the software they run on. A human signs off on everything that goes into production.
The content and research are what clients come for. The story here is what it took to build it, keep producing high quality content, at volume and honestly, while the technology underneath changed every couple of weeks. We’ve gone through six ways of working (so far) – three times we built our own machinery because the available technology could not yet do what we needed, and twice we retired it when they could. Five production failures are named, each ended by a design change rather than a patch.
The problem is the one every organisation meets on the way from an AI pilot to production. The models are rarely the constraint. What breaks is delegated authority, meaning who may decide what; accountability, meaning who signs; and observability, meaning what is actually live. The site’s own writing on agentic AI governance and agent autonomy under governance constraints describes those controls in the abstract. This is what they look like when you have to live inside them.
Six ways of working
| # | When | How the work got done | What the platform could not yet do | Why it ended |
|---|---|---|---|---|
| 0 | 2023 to 2024 | Chat prompts, a word processor, one person in every role | Hold a long brief, or run more than one step unattended | The programme stalled under its own weight |
| 1 | Late 2025 to early 2026 | A staged content pipeline, run by hand, with an AI coding assistant | Run a multi-step job without a person scheduling each step | A human was the scheduler; no version control |
| 2 | Early 2026 | A fleet of agents coordinating through a messaging system we built | Let agents share state or talk to each other | Coordination by conversation failed; the human became the router |
| 3 | Spring 2026 | Tracked work items in an automated pipeline, with a shared harness | Track work or hand it between sessions | The workflow layer needed its own workflow layer |
| 4 | Summer 2026 | A full orchestration platform: staged work, contracts, review gates, judges from several model families | Spawn sub-agents, schedule wake-ups, review its own output | The loop cost more than the work, and the platform had caught up |
| 5 | Q3 2026 onward | The vendor’s agent runtime, git and GitHub, several model families collaborating under separation of duties, and a three-line rule | Current |
Two threads run through the table. The fourth column is what the platform could not do at the time; the third is what we built to cover it. The transitions are not all the same kind. Generations 2 and 3 ended because our own machinery failed under load, and we replaced it with more of our own. Only the last transition is a clean absorption: by mid 2026 the available platforms did nearly everything our custom orchestration layer did, and we retired ours in favour of it. That makes two major retirements of home-built machinery in three years (and dozens of smaller ones along the way).
0. The site by hand
The first version went up in 2023 as the shopfront for the consulting practice, with the reference backbone growing behind it. The discipline the estate still runs on was already there in the project brief: a central entity, a topical map, a hierarchy of core and outer pages, a defined intent for every section. The first wave of pages went live in May 2023.
The method was thorough and slow. For each article: source material gathered by hand, prompts kept and versioned as pieces of style law, the model’s prose spliced and edited, charts drawn by hand – taking several weeks per article. By the end of the year the estate discipline lived in spreadsheets: every live URL scored by hand on eighteen axes, every redirect mapped by hand, a topical map and an intent workbook typed row by row. Every table the fleet now keeps as a ledger or a map was once a workbook.
Then the programme stopped where hand programmes stop. Three articles sat with research, sources and prompts ready, and no output. Several months later, it restarted as a sequence of six named steps run by hand, with several models’ outlines compared side by side. The design was complete. Execution at volume was the missing capability, and in 2024 no model or tool could supply it.
1. The pipeline
The models changed first. By early 2025 a frontier model could hold a long document, follow a structured brief, and be trusted with a multi-step job if the job was scripted around it. Our first real pipeline was not for the site at all. It came out of the consulting practice, and out of a problem that had grown with every engagement.
As the programmes got larger, so did the number of people to hear from and the volume of what they said. The richest material in any assessment is qualitative: the interviews, the written comments, the sentence someone adds next to a score. The site’s research series shows why. In one study the best-scoring capability in the whole survey turned out to hide a single point of failure, and the only reason anyone found it was a note respondents had written into their own comments. Nothing in the numbers asked for that page. The words did.
The trouble is that words do not scale the way scores do. Coding, tagging, clustering, checking themes against each other, the ordinary steps of qualitative analysis, are done by hand, and by hand they stop being possible somewhere past a hundred voices. Language models changed that. Entity extraction, sentiment, theme clustering and cross-referencing, each a simple, checkable step, could now be run over hundreds of interviews and thousands of comments, with the analyst reviewing the result rather than producing it. A survey became an entity graph with nodes, edges and weights, and later a time dimension, so that a theme could be watched moving across a programme rather than sampled once. The leap is the same as it was in the 1990s when spreadsheets started replacing handwritten ledgers. The insight that used to take a team several weeks, became something you could ask on the day.
The site’s pipeline followed in 2025. It pulled search demand, filtered it to the queries that mattered, and drove generation through numbered stages. The stages were packaged prompts and scripts. The scheduler was a person, changing directory per cluster and reading the output as memos. By the turn of the year the system described itself as dozens of clusters and workflows.
Two things from this period set the direction. An AI coding assistant arrived in December, and the first thing we pointed it at was not generation but testing: it found fifteen gaps in two generation workflows. In early 2026 the repository came under version control after thousands of sessions and more than two years of work.
2. A fleet coordinated by conversation
In early 2026 the agent runtimes made it practical to run several agents at once, each in its own session with its own instructions. What they could not do was share state, hand work to each other, or know what another agent had done. So we built that: a shared channel, inboxes, handover files, a facilitator, and a fleet of named agents with personas and seats.
It is worth being precise about what went wrong, because nothing exotic did. The agents were competent. The failure was ours, in the shape of the system. One weekend in March a small bug in the handover mechanism stalled the fleet. Five agents, each certain another was blocked on it, none able to see the others’ messages, and a human in the middle relaying. The complaint of the weekend, verbatim from the record, was that agent communication is not working and all the load is on the human. Content was already publishing by then, and already being halted: eight named halts in that month alone, each one the fleet stopping so the mechanism could be repaired.
The decision that closed this generation is the first of five that recur. Coordination cannot be a conversation. It has to be a record that any participant can read without asking anyone. The site’s page on enterprise agent workflow patterns says it in one line: state lives in the pool, not in a dispatcher’s memory, so the system survives an agent restarting. We went through the alternatives by hand and settled on the same answer.
3. Tracked work in an automated pipeline
The runtimes still had no notion of work: no queue, no ticket, no way for a job to outlive the session that started it. So work became tracked items, multi-step jobs became chains, a background runner pulled work, and a shared harness held tracking, manifests and specifications for every project. Work flowed again. The record from the time notes that the pipeline had been in such a state before the migration that we had stopped trying to fix it, and now it simply ran.
Then the machinery started consuming the attention it was built to save. The diagnosis: start at the top, believe there is a system, fail, dig, find more failures, dig further. By May the content repository fell almost silent while the harness was rebuilt, and the backlog had to be purged by script because they had outgrown triage by hand.
The second recurring decision came out of this. When the workflow layer needs its own workflow layer, build the replacement beside it and cut over. Do not add a floor. We chose a parallel build over a migration on the grounds that the old system had not worked and a clean one was easier to reason about, and that choice was right.
4. The orchestration platform
The second harness was the ambitious one, and the summer of 2026 was the period when the models were moving fastest. Longer context, better tool use, reliable structured output, and, importantly, several strong model families at once. We used all of it. A piece of work moved through named stages, each with a written definition of done. A second agent had to countersign before it could advance. A reviewer was set to argue against it. A judge from a different model family gave the verdict, and the close was checked against the original ask. Models were assigned by task: one family for decisions that set precedent, another for routine setup, a third to argue the other side, and each stage of the pipeline could choose its model from a list. The site’s page on governing the model layer treats models as a swappable, governed layer beneath the agents. That is what we were running.
This is also the generation in which the fleet’s real output peaked. Cluster-seeding campaigns, six clusters of research briefings, the first pull requests raised by agents. August was estate work: a ledger of what is actually live, redirect and rename tooling, plugins for stable anchors, case-study modules. One quiet habit deserves a line. Before trusting a new generation pipeline, we ran it head-to-head against the old one on the site’s strongest organic page, the guide to AI use-case prioritisation, scored the two against real demand rather than taste, and held the result at the human gate unpublished. The site’s page on continuous evaluation and drift monitoring tells enterprises to do exactly that.
And yet the pain moved. In spring the complaint had been that agents could not hear each other. By July it was that the loop cost more than the work. The words we used tell it: the vocabulary of handovers and inboxes peaked in May and then collapsed, replaced by the vocabulary of stages and gates, and by summer the complaints were about those. From the record in July: weeks of code were being thrown away as fast as a fleet of agents could build it; and the sentence the eventual cutover follows from, that yet another hook was firing all the time and there were already too many. Three of the five named failures below happened in this period.
5. The cutover
By autumn the vendors’ runtimes had absorbed nearly everything we had built. Sub-agents, scheduled wake-ups, structured handoffs, tool permissions, and, through git and GitHub, work tracking and review as first-class objects. Our orchestration platform was no longer covering a gap. It was the gap.
The trigger was an audit of one question: how does our machinery assign a page its URL? The answer was eighteen separate mechanisms, dozens of places that derived an address, dozens more that validated or guarded one, and the single field that should have decided it, the address on the topical map, read by nothing and overwritten nightly from the live site. The diagnosis in the plan of record: find a problem, build a guard; cheat around it next time; the guard is broken too; build another.
The cutover was done in days, with the site still publishing. The harness was fenced, not deleted. GitHub issues and milestones became the status record. Every merge now needs a passing build and an independent review from a model, bound to the exact commit; anything that touches the live site needs a second review from a different model family. Merges are automatic on a passing verdict; no human approves a pull request. In the first week the reviewers blocked roughly a third of what they read. The site’s page on built-in quality when AI agents write the code puts the principle in one sentence: the resolution is not better agents, it is better harnesses. The cutover was the moment the harness became small enough to be one.
Three lines went into the repository’s standing instructions the same week, and they are the closest thing this report has to a thesis:
A fact lives in one place and is validated once, when it is loaded. A check may only detect that the source has drifted; it may never derive, repair, or stand in for the source. A change that adds a guard, a fallback, or a refusal is rejected unless it was explicitly ruled in.
Five failures, and what each taught
These are told as patterns rather than as case files, because the pattern is what transfers. Each is real and each is in the record.
Before the list, one observation that frames all five. When agents go wrong in production it is tempting to reach for the language of misalignment. In our experience the agents behaved exactly as their environment rewarded. A model that is asked to fix a failing check will fix the check. A model that is asked to make a run pass will make it pass. Repeat that a few hundred times without a rule about where a fact may live and you get a system that looks, from the outside, as though it is working against you. It is not. It is solving the problem in front of it, the way we asked. The site’s page on goal and policy engines has a name for the mechanism, goal drift: a sub-goal stops serving the objective that spawned it because an intermediate step optimised for local progress. And the page on built-in quality when agents write the code draws the conclusion: the resolution is not better agents, it is better harnesses. The shape of the system, not the intent of the model, is what has to be governed. Not every failure below is goal drift; the first is plain authority. But all five are failures of shape.
The shared record that vanished. A background process, doing what its code told it to, deleted our live production work database – by replacing the live database with an empty one. Every agent’s tooling failed at once. The backup safeguard did its job and refused to overwrite the good copy, and recovery was possible only because the deleted database was still open in a different process and could be read back out of the operating system as long as we could keep the connection open. What it taught: destructive actions on shared state are a class of action, not a code path. They need an explicit, scoped, short-lived grant, and a process that any runner can reach with standing credentials is the wrong place for them. The site’s page on the Agentic Trust Framework makes the general case: when agents run on ephemeral, just-in-time credentials the architecture constrains the blast radius by design, and retrofitting that onto agents already holding persistent access is a fight against every existing pattern. We learned it from the inside of a deleted database.
The contamination. One residual keyword in a cluster’s configuration, left over from the site’s consulting origin, quietly pulled off-topic sources into every page of that cluster. Around seventy configurations and two dozen live pages drifted off topic before anyone noticed, because every test passed. The ruling that came out of it is the first architecture decision in the content repository: validate the output, not only the input. The site’s page on hallucination detection and context lineage argues for lineage from source to claim. This is where we learned we also needed lineage from planning to production.
The rename. An estate-wide change to page addresses, correct in principle, was executed without the redirects being written and verified as facts. Meaning visitors were coming to pages that moved with no forwarding address. The change was reverted the same day, but the revert relied on the platform’s implicit rewriting, and weeks later the damage was still escalating. What it taught: an asset’s address is a fact with one owner, decided once and never derived. Live assets are now moved to their intended addresses in small batches, each batch gated on the previous one reconciling cleanly against evidence.
The spend. One agent, doing a legitimate job, spawned nearly two hundred frontier-model sub-agents in an afternoon and consumed a month’s budget in a couple of hours. Observability reported it perfectly, after the money was gone. What it taught: budget and the authority to spawn are controls, not metrics. The scaled-agile side of the site learned this for human teams years ago: lean budget guardrails – the answer was not more approvals but smarter boundaries, set before the work starts. An agent’s execution scope includes the right to spend, and it has to be bounded the same way.
The halt. Told above under the cutover, and the one that was most clearly our own design mistake: guards accepted one at a time, each individually reasonable, until the thing they were guarding was read by nothing – as the agents learned that it is easier to replicate or re-derive the fact in a different location than jump through the hoops of accessing the original. The data management fundamentals page states the discipline: a broken write should fail at ingestion rather than corrupt a downstream report six steps later.
There is a sixth pattern running through all five, and it is the one a CTO will recognise from any automation programme. When one stage of a pipeline accelerates, the pressure moves elsewhere. When agents can draft faster than anyone can review, the review stage becomes the bottleneck and the place where quality quietly erodes. Our answer was to make review itself something agents do, under separation of duties, and to keep the human at one point in the flow where their signature actually means something. The site’s page on sprint cadence for AI teams names the same thing for human teams: the review queue as the new bottleneck.
What survived
Six ways of working reduced to four controls, all of them in service of one thing: a map the reader can trust. Three for the content, one for the code.
- An address is decided once and never derived. Every page’s URL is set in the topical map at the moment the page is classified, and read verbatim by every step after that. No step may compute, repair or fall back to an address.
- Every claim resolves to a source that says it, or the run stops. Citations are checked against the source text. There is no fallback path; a failed check ends the run.
- Nothing goes live without a human signature, and a ledger records what is actually served. Every publish is reviewed and approved by a person, per article, per run. After it, one record, written by one process from one machine, says what is live. That record is the truth over any plan and over any agent’s memory. The site’s page on AI accountability and responsibility calls this assigning ownership; in practice it means one signature and one ledger.
- Agents change their own code only through review. Every change needs an independent verdict from a model, tied to the exact commit. Anything touching the live site needs a second verdict from a different model family, dynamically scaling effort based on task complexity. The merge is automatic on a passing verdict. This is separation of duties for AI-written code, and on this system it blocks roughly a third of what it reviews.
One boundary is worth stating plainly, because the four controls put the human in different places. Content is signed by a person, every page, every time. Code is not: agents merge their own changes on the strength of independent model review, and the person’s role there is to set the rules the review applies, not to approve pull requests. The three-line rule forbids compensating guards, the checks that stand in for a fact nobody owns. It does not forbid controls: a budget set before a run, a signature before a publish. The difference is that a control is decided once, in the open, and a guard accretes.
What did not survive is just as instructive. Every piece of coordination machinery we built, twice, is gone. Not because it was wrong, but because each was right for one season. Before platforms could do sub-agents, we built a fleet with an orchestration layer. When it could not schedule, we built a wake and coordination system. When it could not hand work between sessions, we built handovers and inboxes and a messaging bus. Each became a liability the moment the platform shipped its own, and the sunk cost was real every time. The framework selection page on the site frames build-versus-buy as a durable decision. On this evidence it is every two to twelve weeks: four major model releases arrived in the space of five weeks this year, each a shift in what the platform could do.
What we would tell a CTO starting this year
The site’s own writing covers pilot-to-production scaling, AI operationalisation and why most AI pilots fail. What three years add is the order in which the lessons hurt. This is that order.
- Coordination is a record, not a conversation. The first thing to fail, in the first month of the fleet. The site’s page on agent transparency and auditability asks whether you can show, months later, exactly why an agent did what it did. A conversation cannot answer that. A record can.
- When the workflow layer needs its own workflow layer, build the replacement beside it. The second thing to fail, by spring. Twice the answer was a parallel build and a cutover; neither time was it a migration.
- Budget and spawn authority are controls. Summer. Observability reports the spend after the money is gone. The scaled-agile side of the site says it for human teams on lean budget guardrails: not more approvals, smarter boundaries. Set them before the run.
- Govern the shape of the system, not the intent of the model. The lesson that took longest to see, because each failure looked different from the last. The site’s goal and policy engines page names the mechanism, goal drift, and the data management fundamentals page names the cure: one owner per fact, fail at ingestion, not six steps later.
- Treat your own scaffolding as seasonal. The last lesson, and the one we expect to keep relearning. Everything you build to cover a gap in the platform has a shelf life of somewhere between two and twelve weeks. Build it anyway if you need it now. Then retire it the moment the platform absorbs it, without stopping the line, and do not let sunk cost vote. This is the third lean-agile principle, assume variability and preserve options, applied to your own tooling.
What the scars say about the map
Writing this report meant testing about twenty of the site’s own reference pages against three years of production. That is the most direct evidence there is that the reference material is written from practice, so the result belongs here, both ways.
Pages the scars confirmed, sometimes word for word. Built-in quality when agents write the code says the resolution is not better agents but better harnesses; two retirements of our own harness say the same. Data management fundamentals says a broken write should fail at ingestion rather than corrupt a report six steps later; the halt is what happens when it does not. Goal and policy engines defines goal drift, and four of the five failures are examples. Lean budget guardrails, written for human portfolios, turned out to be the right answer to an agent that spent a month’s budget in a couple of hours. And the team resilience study, which found the real finding in the comments and not the scores, is the method the whole pipeline grew from.
Pages the scars revise. AI agent framework selection calls build-versus-buy a long-term architectural commitment. On this evidence it is one you revisit every few weeks, and the page will be rewritten to say so. Agent autonomy with governance constraints treats spend as an escalation threshold, something that triggers a review when crossed; the spend failure says it has to be a bound set before the run, or the money is gone before the review convenes. AI governance performance metrics carries compliance and detection measures but not the three an agent workforce lives or dies by: cost per outcome, human intervention per outcome, and time to value. And multi-agent systems explains memory and state but never says the one thing the first fleet needed to hear, that coordination has to be a record and not a conversation; its companion page on workflow patterns does.
Four pages to revise out of twenty is the map working as intended. A reference estate that never has to correct itself is one that has not been used.
What comes next
We expect the pattern in this report to keep holding: everything mechanical about running the estate will, in time, be absorbed by the platform. Review, verification, publishing, estate management: each is a season’s coat, and each will be shed. What the platform cannot ship is judgement. Which facts are true, which are worth a page, how the pages fit, what the reader needs next. Three years of building tools ends with the conclusion that the tools were never the asset. The map is.
So the strategy is not to retire the human in the middle. It is to keep raising the bar on what the human does there. Everything that becomes routine gets pushed to the agents, and from the agents to the platform, and the attention that frees up moves to the next hardest thing. That is a ratchet. Each turn commoditises a little more of the work and concentrates a little more of the effort on the part nobody else can do. Each turn of the ratchet has to be justified by measurement: cost per outcome, human intervention per outcome, time to value. The site’s own governance metrics page does not yet carry those three for an agent workforce, which is itself a gap on the map.
Three concrete moves follow.
The map as an instrument for seeing. A topical map is usually sold as a way to cover a territory. Its more useful property is the opposite. Once the map is drawn, it shows where nobody has been: the topics that exist in the structure but have no page written from practice, the questions the reference material raises and does not answer. Those gaps are not a backlog. They are the chart of the unexplored, and they say where the next piece of research or the next case study belongs. The estate will be extended from the gaps outward, deliberately, rather than from whatever is easiest to write next.
Practice will also redraw the map, not only fill it. An inherited structure conceals the questions it never thought to classify, and the research series has already produced findings that belong to no existing page. Judging when a category is wrong, not just which pages are missing, is the part of the work that stays human, and it is where the bar rises next.
More practice, more research. The site began as the backbone for client work, and the next turn brings more of that work back onto it. The qualitative method from the first pipeline, the entity graph with weights and a time dimension, run on real programmes and written up the way the team resilience study already is: honestly, from the evidence, with the reference pages carrying the general truths so the case study can carry only the specific ones. The three-hundred-page report finally becomes the fifty slides and a hundred links it should have been.
The next retirement, named in advance. The review-and-verdict machinery is the newest layer we run. We expect the platform to absorb it, and when it does we will retire it, without stopping the line and without letting sunk cost vote. Saying so beforehand is the lesson of this report, practised.
Limits
One estate, a small team, three years. This is evidence of engineering judgement under real failure, not of organisational adoption at scale, and in every organisation we have worked in, the change around the system has been at least half the work whatever the technology. That half is not in this report; it is in the research series. The three measurements above were not instrumented over the period. And the story is told from our own records, which is what makes it checkable and also what makes it partial.