AI Strategy & Transformation
31 MIN READ

Build vs Buy AI: A Decision Framework for Enterprise Leaders

Build vs Buy Decision Framework: score each AI capability on core factors, adjust with peripheral ones, then run a three-year cost model to decide.

Most enterprises still ask whether to build or buy AI as if the company makes one choice for its entire technology stack. That framing produces expensive mistakes: teams build commodity capability at premium cost, or buy differentiated capability a competitor can license next quarter. This build vs buy decision framework treats sourcing as a decision made once per AI capability, using named criteria, priced over a three-year horizon, and recorded so it can be defended and revisited.

Build vs Buy AI: What the Decision Covers in Enterprise AI Strategy

Build vs buy AI is the decision to develop an artificial intelligence capability in-house, purchase it from a vendor, or engage a partner to deliver it, made separately for each capability rather than once for the whole company. Most leadership teams still make it once, at the portfolio level, which is where the expensive mistakes start.

Anup Srivastava, Chandrani Chatterjee and Jonathan Tanone make the same point in the Harvard Business Review: companies treat build versus buy as a single strategic choice, but the real question is whether this particular asset or capability should be built or bought (Harvard Business Review). They name five conditions that answer the question for each capability:

  • the type of asset being created
  • the time horizon the organization is working against
  • the organization’s ability to integrate what it acquires
  • how immediate the market need is
  • how much uncertainty management is willing to carry

The conventional framing, buying for speed, scale and ready-to-use assets, building for cost control and customization, treats every capability as if it sits at the same point on all five conditions, and that assumption breaks down fast: a commodity capability with slack time scores differently on those five conditions than a differentiated capability needed next quarter.

Enterprises are answering that per-capability question in the buy direction more often than expected. Menlo Ventures research, relayed by Beam.ai as “the great AI flip,” finds a majority of enterprise AI use cases are now purchased rather than built in-house: the opposite of the build-first instinct many technology leaders carry into the first planning meeting.

A third path exists alongside build and buy: partner, where a specialist delivers the capability under the enterprise’s own governance rather than through a licensed product or an internal team. MasternodeAI and TechAhead’s build, buy or partner frameworks both treat partnering as a distinct option, not a hedge between the other two, and it earns its own line in the criteria that follow.

The Build vs Buy Decision as a Per-Capability Choice

The build vs buy decision is scoped to one capability at a time, because a capability’s asset type, time horizon and integration demands rarely match the profile of the capability next to it on the roadmap.

A capability like fraud-scoring on payment data behaves differently under Srivastava, Chatterjee and Tanone’s five conditions than a capability like meeting-summary generation for internal calls. Fraud-scoring touches a regulated asset type, carries a long time horizon because retraining and audit cycles run for years, and demands tight integration with the core transaction system. Meeting summarization is a commodity asset type, needed now, and tolerant of a vendor swap next year if a better one appears. Treating both under one company-wide build or buy policy forces at least one of them into the wrong sourcing path.

AI capability is the unit that carries this scoring, not the company and not the department. A capability is a specific function the organization performs with AI, underwriting risk, routing support tickets, drafting contracts, not a technology category like “generative AI” or “machine learning.” Scoring the technology category produces one answer for the whole portfolio; scoring the capability produces a different answer for every row in it, which is why enterprises hold two, three, or a dozen build vs buy decisions open on the same roadmap at once.

The practical consequence is procedural: a build vs buy decision needs to be requested and answered per capability, with its own owner and its own record, before an architecture board approves a build or signs a vendor contract. Skipping that scoping step is how a single “we’re an AI-buy company” policy statement ends up licensing a commodity chatbot and licensing a differentiated pricing engine on the same terms: one overpaying for control it doesn’t need, the other buying away the edge it was trying to protect.

AI Capability and the Partner Option

AI capability decisions gain a third option beyond build and buy: partnering with a specialist who operates the capability inside the enterprise’s own governance and data boundary.

Partnering answers a specific gap that neither build nor buy closes cleanly: capabilities that need deep domain expertise the enterprise doesn’t have time to hire, but also need a level of data control a generic vendor contract won’t grant. A specialist partner embeds inside the enterprise’s infrastructure and governance, delivering the capability without the enterprise owning the underlying model or engineering team, and without the vendor holding the data outside the enterprise’s boundary.

The market shift toward buying that Menlo Ventures’ research documents, and that Beam.ai relays as “the great AI flip,” doesn’t erase the strategic choice to build for capabilities that define competitive advantage: it just means fewer enterprise AI use cases default to build than technology leaders assume walking into the planning cycle. Enterprise AI use cases split roughly into two groups: those where the underlying model is the differentiator, which stay rare, and those where the application of the model to a specific workflow is the differentiator, which cover most of what shows up on an AI roadmap. TechAhead’s build, buy or partner framework and MasternodeAI’s version of the same choice both place partner in this second group by default, reserving build for the capabilities that would lose their edge the moment a competitor licenses the same vendor product.

Build vs Buy AI Decision Criteria: Core Factors and Peripheral Factors

Build vs buy AI decisions rest on a small set of core decision factors, strategy, commodity status, competitive advantage, organizational maturity, cost and requirements fit, while risk, timing, politics and ongoing support act as peripheral factors that adjust the decision but never override it. Most build vs buy checklists treat all ten factors as equally weighted, which is precisely the assumption a 2008 study of large-organization decisions overturned.

P. Hung and G. Low interviewed senior IT decision-makers at ten organizations with turnover above $750 million, and split what they found into two tiers (Journal of Information Technology, via Semantic Scholar): a core tier that decides the sourcing choice, and a peripheral tier that adjusts it. Basit Shahzad, Abdullatif Abdullatif, Naveed Ikram and Adam Mashkoor’s study of large-scale software sourcing frames the stakes at portfolio scale: enterprise IT spending was on a path toward $572 billion by 2022, and every dollar of it eventually runs through a build or buy call at the project level Abdullatif Abdullatif (Build Software or Buy: A Study on Developing Large Scale Software). The 2026 AI-specific frameworks, HP’s decision tree, TechAhead’s Capability, Complexity, Criticality axes, and MasternodeAI’s differentiation, data sensitivity and internal maturity criteria, restate Hung and Low’s split for AI capabilities without citing it: core factors decide an AI build vs buy question, peripheral factors only adjust its timeline.

Core Factors: What Decides Build or Buy

Core factors are the conditions Hung and Low found senior IT executives treating as the actual determinants of a build or buy decision: strategy, commodity status versus competitive advantage, organizational maturity, cost, and requirements fit.

Hung and Low ran their interviews before enterprise AI existed as a category, which is exactly why the split travels well: it measures how organizations reason about sourcing, not what they were sourcing. Strategy asks whether the capability sits inside the company’s stated direction. Commodity status versus competitive advantage asks whether competitors can buy the same thing off the shelf. Maturity asks whether the organization’s own teams can operate what they build. Cost and requirements fit close the set: what it costs against the alternatives, and how well a vendor’s product actually matches the specific requirement rather than a nearby one.

Strategy and Competitive Advantage

Strategy and competitive advantage decide whether a capability is worth owning even at higher cost or effort than buying it. A capability scores high on strategy when it sits directly inside the stated direction of the business, a logistics company’s routing algorithm, a bank’s credit model, rather than adjacent to it. Competitive advantage asks a sharper question: would a competitor gain ground by buying the identical product an enterprise is considering? If the answer is yes, buying closes the gap a competitor would otherwise have had to build themselves.

This factor matters most at the top of the priority list because it overrides the other four when it triggers: a capability that scores high on competitive advantage gets built even when it costs more and takes longer than the buy alternative, because the alternative is a market where every competitor buys the same edge. A regional insurer that builds its own claims-fraud model instead of licensing a generic fraud API keeps a detection pattern competitors can’t see or copy; the insurer that buys the same API as three of its competitors is paying for parity, not advantage.

Maturity, Cost and Requirements Fit

Maturity, cost and requirements fit decide whether the organization can actually execute a build even when strategy says it should. Maturity measures whether the internal team has run a comparable project before, data engineering, MLOps, model monitoring, not whether the organization has AI talent in general. Cost compares the build path’s engineering time, infrastructure and ongoing maintenance against a vendor’s licence and integration fee, not against the vendor’s list price alone. Requirements fit checks how close the nearest vendor product actually sits to the specific requirement, because a product that covers 80% of a regulated workflow still leaves the hardest 20% for the enterprise to build anyway.

These three factors work as a gate, not a score: a capability can pass strategy and competitive advantage and still fail here, and when it does, the organization buys or partners regardless of how strategic the capability looked on paper. A hospital system with strong data science talent but no prior production ML deployment fails the maturity gate even for a strategic capability, and a vendor with a proven radiology model closes that gap faster than a first production build would. Requirements fit produces the opposite failure: an enterprise with excellent internal maturity discovers no vendor product covers its specific compliance workflow, which pushes a moderately strategic capability toward build by elimination rather than by strong strategic pull.

Peripheral Factors: Risk, Timing, Politics and Ongoing Support

Peripheral factors, risk, timing, political dynamics and ongoing support, adjust a build vs buy decision that the core factors have already pointed toward, and they only flip the decision when a core factor was already close to even.

Risk covers what happens if the chosen path fails outright: a build that never reaches production, or a vendor that exits the market mid-contract. Timing measures how urgent the capability is against the board’s patience, not against an abstract project schedule: a capability needed for a Q1 regulatory deadline weighs timing heavily even if its core-factor score is mixed. Political factors register who inside the organization has a stake in the outcome: a build path protects an internal platform team’s mandate, a buy path protects a procurement relationship, and both pressures show up in the room before the criteria do. Ongoing support asks who answers the page at 2 a.m. when the capability breaks: an internal build commits engineering time indefinitely, a vendor contract commits a support-level agreement instead.

None of these four factors decide the sourcing question on their own, and none of them should: an organization that lets timing pressure override a competitive-advantage finding buys away an edge to hit a deadline, then discovers the deadline was negotiable and the edge wasn’t. Hung and Low placed these four factors as secondary for a reason that still holds for AI capabilities; they describe the conditions around the decision, not the capability itself.

How the 2026 AI Decision Frameworks Map onto Core and Peripheral Factors

Three named 2026 frameworks, HP’s decision tree, TechAhead’s Capability, Complexity, Criticality axes, and MasternodeAI’s criteria set, all restate Hung and Low’s core and peripheral split for AI capabilities, without naming the split they’re restating.

HP’s Enterprise AI Services decision tree asks three questions: is the capability core to competitive advantage, how sensitive or regulated is the data behind it, and what scale of investment and speed to market does it require. The first two questions are core factors under Hung and Low’s split; the third, speed to market, is peripheral. TechAhead’s 2026 framework scores each initiative on Capability, Complexity and Criticality before recommending build, buy, partner or a staged transition; Capability and Criticality map to strategy and competitive advantage, Complexity maps to maturity. MasternodeAI’s criteria, competitive differentiation, data sensitivity, and the maturity of the internal machine learning team, map even more directly onto the same three core factors under new names.

2026 framework Question it asks Hung and Low factor it restates Tier
HP decision tree Is this core to competitive advantage? Strategy / competitive advantage Core
HP decision tree How sensitive or regulated is the data? Requirements fit Core
HP decision tree What scale and speed does this need? Time and timing Peripheral
TechAhead C-C-C Capability and Criticality score Strategy / competitive advantage Core
TechAhead C-C-C Complexity score Maturity Core
MasternodeAI Competitive differentiation Competitive advantage Core
MasternodeAI Internal ML team maturity Maturity Core

The pattern across all three frameworks is consistent: every 2026 AI-specific criterion traces back to one of Hung and Low’s five core factors or four peripheral factors, which means the frameworks aren’t offering new criteria; they’re offering new labels for the same ten-factor structure applied to AI. The one place all three frameworks agree in practice, without stating it directly, is precedence: a peripheral factor such as delivery time can break a tie between two capabilities that score evenly on the core factors, but none of the three frameworks lets a peripheral factor overrule a core factor that has already returned a clear answer. A capability that scores strongly on competitive advantage still gets built even when speed to market argues for buying: the frameworks just don’t state the rule as plainly as Hung and Low’s original research does.

Why the Same Ten Factors Keep Reappearing Under New Names

Every AI-specific build vs buy framework published in 2026 restates Hung and Low’s original ten factors, which means an enterprise doesn’t need a new framework for every capability category it evaluates.

A programme lead evaluating a document-summarization capability and a programme lead evaluating a fraud-detection model can use the identical ten-factor list, because the factors don’t change with the technology; only the answers do. The two capabilities diverge sharply on exactly one of the ten factors, competitive advantage, while the other nine move far less between them, which is the divergence a single averaged portfolio score hides. Running each capability through the same core-then-peripheral sequence, rather than adopting a new vendor’s named framework for each new capability type, keeps the criteria consistent across a portfolio an architecture board has to compare in one meeting.

The risk in the opposite direction is real: a different named framework for every capability category produces board decks that can’t be compared to each other, because HP’s three questions, TechAhead’s three axes and MasternodeAI’s three criteria use different words for the same five core factors. An architecture board that adopts Hung and Low’s original ten-factor structure as the shared vocabulary, and treats the 2026 frameworks as translations rather than replacements, keeps every capability’s decision defensible against the same standard a year later when someone asks why it was built instead of bought.

Total Cost of Ownership for Build vs Buy AI: Three-Year Cost Model

Total cost of ownership for an AI capability is modeled across three cost lines over a three-year horizon, per-request buy costs, infrastructure-and-labor build costs, and the technical debt hidden inside AI-generated code, priced as unit economics, in cost per request, not in the acquisition price alone. The number on the vendor’s price sheet and the number on the internal build estimate both undercount, in opposite directions, and the gap between them only shows up after eighteen months of usage data.

Mohit Sewak’s 2026 three-year total cost of ownership analysis makes the underlying point directly: the visible acquisition cost, the licence fee, the initial build budget, is a small fraction of what an enterprise AI capability actually costs across three years of operation. Modeling the buy path, the build path and the hidden costs of each separately, as unit economics, catches what a single acquisition number always misses: a capability that looks cheap at pilot volume and expensive at production volume is the most common total-cost-of-ownership surprise enterprises report.

Buy Path Costs: Per-Request Pricing and Reserved Capacity

The buy path’s real cost is a per-request or per-token price that multiplies with usage, not the flat licence figure that appears in the initial vendor quote.

Federico Vibrati and Federico Preli’s Google Cloud guide frames the buy-path cost choice as a trade-off between pay-as-you-go pricing and reserved capacity (Google Cloud). Pay-as-you-go pricing runs on a shared-quota model: requests inside an organization’s default throughput threshold get priority access targeting a 99.5% service-level objective, while a spike above that threshold drops into a best-effort lane with no such guarantee. That structure means the visible per-request price only holds steady below the threshold; cross it during a traffic spike, and cost per request rises exactly when the business needs the capability most, at a product launch or a regulatory filing deadline.

Reserved capacity inverts the trade: a fixed monthly commitment buys a guaranteed throughput ceiling regardless of spikes, at a higher baseline cost than pay-as-you-go during quiet months. The break-even between the two options is a usage-volume question, not a preference question: an enterprise with steady, predictable request volume overpays for reserved capacity it rarely exhausts, while an enterprise with spiky, launch-driven volume overpays for pay-as-you-go every time it crosses the threshold. Modeling the buy path’s total cost of ownership means picking the pricing structure that matches the capability’s actual traffic pattern, not the one that looked cheapest in the sales deck.

Build Path Costs: Infrastructure, Labor and Break-Even Usage

The build path’s total cost of ownership includes infrastructure, engineering labor and ongoing operation, and it beats the buy path only above a specific, calculable usage level.

Daniel Huynh’s AI Total Cost of Ownership Calculator, built to compare in-house deployment against AI APIs, makes the build path’s hidden cost explicit AI APIs (Hugging Face): open-source models are “free” only in licence terms. Self-hosting carries a large initial infrastructure cost, GPU capacity, serving infrastructure, monitoring, plus the labor cost of the engineers who operate it, and that initial acquisition cost amortizes only as usage climbs. Below the break-even usage level, every request costs more self-hosted than the same request would cost through a vendor API; above it, the self-hosted cost per request keeps falling while the API’s per-request price stays roughly flat.

The break-even level itself moves with the specific model and infrastructure choice, which is why Huynh’s calculator exists as a tool rather than a rule of thumb; but the direction of the trade doesn’t move: a capability at pilot volume almost never clears the break-even threshold, and a capability that has scaled to production volume across a full enterprise user base frequently does. Enterprises that model this cost line accurately budget for a period where the build path costs more than the buy alternative would have, and treat that period as the price of reaching a usage level the buy path can’t discount its way past.

Hidden Build Costs: Technical Debt in Brownfield Environments

AI-generated code raises developer productivity by up to 55% in controlled tasks, and it compounds technical debt in brownfield environments fast enough to erase that gain within a year.

MIT Sloan Management Review’s research on generative AI coding tools documents both halves of that finding directly (MIT Sloan Management Review): the productivity studies showing up to 55% gains ran in controlled environments, on isolated tasks, not inside the complex existing systems most enterprise engineering happens in. When AI-generated code ships rapidly into a brownfield environment, an established system with years of accumulated structure and undocumented dependencies, and an inexperienced developer merges it, the code compounds problems that were already present instead of resolving them. The review’s recommendation is direct: organizations deploying AI coding tools at scale need clear guidelines, a standing priority on technical debt management, and training that teaches developers to use the tools responsibly rather than simply faster.

For a build-path total cost of ownership, this cost line rarely appears in the initial estimate because it doesn’t show up in the first sprint: it shows up in the sixth, when a feature that should take two days takes eight because the codebase underneath it has quietly become harder to change. A build path costed only on initial development time and infrastructure, without a technical-debt line for brownfield deployment, understates its own three-year total by a margin that grows every quarter the capability stays in a legacy environment.

What Falling Model Prices Change in the Three-Year Model

Falling model prices lower the buy path’s future cost per request, but the drop is smaller and slower than most three-year models assume, because demand in the emerging market for intelligence doesn’t fall with price the way it typically does in a mature commodity market.

Mert Demirer, Andrey Fradkin, Nadav Tadelis and Sida Peng’s NBER working paper measures the LLM market directly on API usage data from OpenRouter and Microsoft Azure Andrey Fradkin (NBER Working Paper 34608): open-source models now run 90% cheaper than closed-source models of comparable intelligence, and prices keep falling as open-source entrants multiply. That number changes the buy path’s future cost line in a three-year model: a capability priced against today’s closed-model rate may cost a fraction of that by year two if an equivalent open-source option is usable by then.

The same paper’s second finding limits how far that discount stretches: short-run price elasticity for LLM demand sits just above one, which means falling prices don’t multiply spend the way a classic Jevons-Paradox effect would predict; usage grows roughly in proportion to the price drop, not faster than it. A three-year buy-path model that assumes falling prices will both cut cost per request and hold total spend flat is assuming two things Demirer and his co-authors’ data doesn’t support at once; the realistic model prices the discount into future years while still expecting total spend to grow with adoption, not to shrink because unit prices did.

The Hybrid Path: Buy the Foundation Model, Build the Intelligence Layer

This hybrid build and buy path buys the foundation model and the infrastructure underneath an AI capability while building the domain logic and proprietary workflows on top of it, which is closer to how most enterprises actually source AI than a pure build or pure buy choice. Framed as a single word, “hybrid” sounds like a compromise; framed as a ladder of specific technical steps, it’s a sequence an enterprise climbs deliberately, one rung at a time.

Monterail’s 2026 formulation names the pattern precisely: buy the foundation, build the intelligence layer, partner for execution. Models and infrastructure get bought because no enterprise engineering team can out-invest a foundation-model lab on pretraining. Domain logic and proprietary workflows get built because that is exactly where an enterprise’s own data and process knowledge live, and no vendor can build it for them.

Buy the Foundation, Build the Intelligence Layer

Buying the foundation model and building the intelligence layer separates what an enterprise can’t out-build, the base model, from what a vendor can’t build for it, the enterprise’s own workflow logic.

The foundation model is the pretrained base: the enterprise buys access to it, through Azure OpenAI, AWS Bedrock, Vertex AI or a comparable provider, rather than training an equivalent model from raw data: a project that costs far more than any single enterprise AI capability would justify. The intelligence layer sits above that foundation: the prompts, retrieval systems, fine-tuning data, evaluation harnesses and business logic that turn a general-purpose model into a capability that does one enterprise’s specific job. A generic foundation model can draft a contract clause; the intelligence layer is what makes it draft a clause that matches one enterprise’s actual contract templates, risk tolerance and approval workflow.

This split explains why the build vs buy question resolves differently at each layer of the same capability. At the foundation layer, the answer is almost always buy, because the foundation model is a commodity input with genuine competition between providers. At the intelligence layer, the answer shifts toward build, because the workflow logic is exactly the kind of asset Srivastava, Chatterjee and Tanone’s competitive-advantage condition flags for ownership: it encodes decisions specific to one enterprise that a generic vendor product can’t replicate without access to the same proprietary data and process knowledge.

The Customization Ladder Between Buy and Build

The AWS Generative AI Customization Spectrum lays out an eight-step staircase between prompt engineering and training a custom model, and each step up raises effort, cost, data requirements and control together.

Prompt Engineering and Retrieval Augmented Generation

AWS’s core principle for climbing the ladder is direct: start with the simplest approach and go deeper only when the use case demands it (AWS Machine Learning Blog). Prompt engineering and Retrieval Augmented Generation sit at the bottom two rungs of that ladder, and both work entirely on top of a bought foundation model without touching its weights. Prompt engineering shapes what the model does through instructions, examples and structure in the input alone: no additional infrastructure, no additional data pipeline. Retrieval Augmented Generation adds one component: a retrieval system that pulls the enterprise’s own documents into the model’s context at request time, so the model answers from current, proprietary information instead of only from what it learned during pretraining.

These two rungs cover most enterprise AI capabilities that need enterprise-specific knowledge but not enterprise-specific reasoning patterns: a support assistant that answers from a product manual, a research tool that summarizes a document library. The cost and data requirements at this level stay low enough that most capabilities should start here regardless of how strategic they look on paper, because the AWS principle treats climbing the ladder as a decision the use case has to earn, not a default. A capability that reads its answers correctly through Retrieval Augmented Generation alone has no need for the cost and complexity two rungs up.

Fine-Tuning and Custom Model Training

Fine-tuning and custom model training sit at the top of the ladder, and both change the model itself rather than what surrounds it. Fine-tuning adjusts a pretrained model’s weights on enterprise-specific examples, teaching it a pattern, tone or task structure that prompting alone can’t reliably produce. Custom model training goes further still, building a model from data the enterprise controls end to end: the step Lak Lakshmanan’s data-scaling finding governs directly: a linear decrease in a model’s error rate requires an exponential increase in the data behind it.

Both rungs demand infrastructure, labeled data and ongoing retraining that the lower rungs don’t, which is why they pay off only for capabilities where prompting and retrieval have already been tried and have already fallen short of the requirement: not as a first move. Teams that skip straight to fine-tuning spend real money in compute and real time in delayed launches solving a problem a well-structured prompt would have closed in an afternoon. A fraud-detection model trained on an enterprise’s own transaction history, where the patterns worth learning exist only inside that enterprise’s data, justifies the climb; a customer-service tone adjustment that a well-written prompt already handles does not.

The four rungs trade in the same currency, effort, cost, data need and control, in a fixed order:

Ladder rung Effort Data required Control gained
Prompt engineering Lowest None beyond examples in the prompt Output framing only
Retrieval Augmented Generation Low An indexed document set Answers grounded in current enterprise content
Fine-tuning Moderate to high Labeled examples of the target task Model behavior tuned to enterprise patterns
Custom model training Highest Enterprise data at a scale that supports pretraining Full ownership of the trained model

Moving down this table to solve a problem is always available; moving up it is a use case earning the next rung, not a default engineering choice.

Buy to Learn, Build to Last: The Hybrid Path Over Time

“Buy to learn, build to last” describes how the hybrid path changes over a capability’s lifetime, not just across its architecture layers, according to Helium42’s pattern across more than 500 European enterprise deployments.

Helium42’s implementations across 500-plus European businesses show enterprises buying a vendor product first, for a capability they don’t yet understand well enough to specify a build for, and building a replacement or extension only once the buy-path deployment has taught the organization what the requirement actually is. Buying first de-risks the learning phase: a vendor product gets a capability into production fast enough to generate real usage data, at a cost the organization can absorb before committing to a build it might specify wrong.

Building later locks in what the learning phase revealed, once the requirement has stopped moving. A capability that looked like a generic document-classification problem in its first vendor deployment often turns out, after six months of production use, to need a proprietary taxonomy and an escalation workflow no off-the-shelf classifier supports; at which point the organization has the specification a build needs and the usage evidence to justify the investment. Treating buy and build as sequential phases of the same capability, rather than as a one-time either/or choice, is what makes the hybrid path a time dimension as well as an architecture dimension.

Sovereignty and Sustainability: Evidence from the Government LLM Framework

Governments choosing between buying, building and hybrid approaches to large language models evaluate the choice on six dimensions, not the enterprise’s usual cost-and-control pair alone.

Jiahao Lu, Mohan Kankanhalli and their co-authors evaluate sovereign, commercial and open-source models across sovereignty, safety, cost, resource capability, cultural fit and sustainability, arguing national AI strategies are typically pluralistic rather than a single sourcing choice applied everywhere Mohan Kankanhalli (Lu et al., “Buy versus Build an LLM: A Decision Framework for Governments”). Governments rely on commercial models for non-sensitive or commodity tasks while pursuing greater control, build or a domestic hybrid, for critical, high-risk or strategically important applications, and building doesn’t require acting alone: public research institutions, universities, state-owned enterprises and joint ventures can all carry the domestic-capability side of a hybrid approach.

Two of Lu and Kankanhalli’s six dimensions translate directly into enterprise practice even outside government: sovereignty, in an enterprise context, becomes control over where sensitive data and model weights physically sit and who can access them; sustainability becomes whether the organization can keep operating the built or bought capability as vendor relationships, regulations and model generations change over the years the capability stays in production. An enterprise weighing a hybrid path for a regulated capability gains more from asking these two questions early than from any cost comparison alone, because a cost-optimal choice that fails on sovereignty or sustainability gets unwound at a much higher cost later.

How to Run a Build vs Buy AI Decision: From Explicit Criteria to an Auditable Record

Running a build vs buy AI decision means making the criteria explicit, testing the capability against a sourcing test and an integration check, writing a decision record with a named owner and success metric, and setting a review date before the model market moves the decision again. Most of that process happens informally in a single meeting, which is exactly the point at which an auditable decision quietly becomes an unauditable one.

Janardan Misra, Vikrant Kaulgud, Adam Burden and Sanjay Podder’s structured decision-support research is direct about what a build vs buy process needs to survive scrutiny: explicit, systematically comparable criteria, not fragmented expertise and informal reasoning, produce decisions with real transparency and auditability Vikrant Kaulgud (arXiv:2606.29816). That finding turns into a four-step sequence any architecture board can run for a specific capability, starting with the criteria themselves.

Make the Criteria Explicit and Comparable

Making the criteria explicit means writing down, before the meeting, which of the core and peripheral factors apply to this specific capability and what evidence answers each one.

Misra and his co-authors’ structured approach grounds this step in an ontology of decision factors spanning strategic considerations, application characteristics, cost and budget constraints, and risk dimensions, applied here as a checklist rather than a debate. What makes that checklist produce an auditable record is when it gets filled in: writing the criteria down before the discussion, rather than during it, is what produces decision auditability, giving a record that shows which factor moved the decision rather than a summary written after the fact to justify whatever the room decided.

The comparability half of this step matters as much as the explicitness half. A criteria list that uses different language for every capability, one deck scoring “strategic fit,” another scoring “core to the mission”, can’t be compared across a portfolio review, even though both decks are being explicit. Using the same five core and four peripheral factors as the shared vocabulary across every capability under review, even when a specific framework like HP’s or TechAhead’s supplied the original language, is what lets an architecture board put ten capabilities’ decision records side by side and see which ones were reasoned the same way and which ones weren’t.

Apply the Sourcing Test and the Three-Layer Check

The sourcing test asks whether building this capability diverts the organization from skills its suppliers already do better; Ravi Venkatesan’s 1992 question, still the right one for AI capability sourcing.

Venkatesan’s Harvard Business Review article, “Strategic Sourcing: To Make or Not To Make,” argues that manufacturing focus means learning how not to make the parts that divert a company from cultivating its own skills; parts a supplier could make more efficiently Not To Make (Harvard Business Review, 1992). Applied to an AI capability, the test asks a pointed version of the same question: does building this in-house grow a skill the organization needs to keep, or does it consume engineering time the organization would rather spend on the capability that actually differentiates it?

DevOpsDigest’s Build, Buy or Integrate framework operationalizes the answer into three questions run in sequence: does the capability represent genuine strategic differentiation, can the organization support it securely and sustainably once it’s live, and does it integrate into a unified platform rather than existing as a disconnected feature bolted onto existing systems. A capability that passes the sourcing test but fails the third layer, strategic, supportable, but architecturally isolated, still creates a build vs buy problem the criteria alone didn’t catch: an internally built capability with no path into the organization’s platform accumulates the same integration complexity a poorly chosen vendor product would have, just under a different name on the budget line.

Write the Decision Record: Path, Owner and Success Metric

A decision record names the capability, the chosen path, the production owner and the success metric; without all four, the decision can’t be defended once its makers move on.

The AWS Generative AI Path-to-Value framework explains why the owner and success-metric fields matter more than they first appear to: most generative AI initiatives clear the proof-of-concept stage, which demonstrates technical feasibility, only to stall before production, because data access gets constrained by security requirements, integration with existing systems turns out more complex than planned, and governance and approval friction adds delay; all while teams struggle to define a consistent success metric connecting the capability to a business outcome AWS Generative AI Path-to-Value (AWS Machine Learning Blog). A decision record that names a production owner and a success metric at the moment the sourcing decision is made, rather than after the pilot succeeds, is what prevents that stall from turning into an indefinite one.

When the recorded path is buy, Pertama Partners’ 2026 vendor-selection weights close the record with the criteria the next team needs: solution fit, integration complexity and total cost of ownership, in that order, rather than price alone. The record stops there for a buy decision; vendor due diligence from this point belongs to procurement, not to the architecture board that made the sourcing call.

Set the Review Date and the Exit Condition

Every build vs buy decision record needs a review date and an exit condition, because the model market that justified today’s answer moves fast enough to make it wrong within a year.

Demirer, Fradkin, Tadelis and Peng’s market data shows why the review date can’t be optional: they document frequent turnover among leading models and creators even as, at the same time, most firms concentrate their usage on a single model rather than spreading it across several. That combination is how lock-in forms quietly: not through a single bad decision, but through a good decision that was never revisited while the market underneath it kept moving. A capability sourced from the best available model eighteen months ago may now be running on a model two full generations behind what a competitor is shipping on, with no one on the team who made the original call still tracking whether that gap has become a genuine cost.

The exit condition closes the loop the review date opens: a stated threshold, a competitor model outperforming the current one by a defined margin, a price change past a set point, a support commitment expiring, that triggers a re-decision automatically instead of waiting for someone to notice. Without it, the review date becomes a calendar entry nobody acts on, and the build vs buy decision that was defensible on the day it was made quietly becomes the decision nobody remembers choosing.

Summary

This build vs buy AI decision framework scopes each choice to one capability, weighs it against core and peripheral factors, prices it across a three-year total cost of ownership model, and closes with a written record; skip that closing step, and even a well-reasoned sourcing choice has nothing to show an auditor, or a successor, for why it was made.

The Decision Sequence in One Pass

The sequence runs in one direction: scope the decision to a single capability, score it against Hung and Low’s core and peripheral factors, price it over three years, check whether a hybrid split serves it better than a pure path, then record and review the result.

Scoping the decision to one capability stops a single company-wide policy from mis-sourcing half the portfolio in one line item. The core and peripheral factors give that per-capability decision one fixed vocabulary, which is what lets a fraud model and a summarization tool run through the same comparison even when their answers land nowhere near each other. A decision record that never prices the three-year figure has nothing to hand an auditor who asks, a year later, why the capability was sourced the way it was.

The hybrid check matters because most enterprise AI capabilities don’t resolve to a pure build or a pure buy once priced and scored, which is exactly why the sequence runs this check as its own step rather than folding it into the cost model above. An unrecorded decision leaves the organization unable to check itself when any of the three forces below quietly change the answer.

What Changes the Answer Later

Three forces move a settled build vs buy decision without anyone deciding to revisit it: falling model prices, a capability’s own maturation from pilot to production volume, and model-market turnover: the reason a decision record needs a review date and exit condition in the first place.

Falling model prices are precisely the kind of shift a decision record with no review date is built to miss. Usage growth works the same way: a capability’s own volume can climb from pilot to production without prompting anyone to revisit the sourcing decision, and only the review date and exit condition already built into the process step catch that a decision priced at pilot volume is still governing production volume. These three forces don’t arrive on separate timelines; they compound, and a decision left unreviewed absorbs whichever combination of them lands first.

None of these three forces require a new framework to manage; they require the review date and exit condition every decision record should already carry, checked against the same core and peripheral factors that produced the original answer, on a schedule the organization actually keeps.

Anonymous. Counted, not tracked.

Where is your organisation with this right now?

What is the hardest part where you are?

Morné Wiggins · Agility at Scale · Talk to me

Privacy Preference Center