Malcolm Angus
← Sources

Show notes

Bain: proprietary data is a moat that leaks

2026-08-22


Decision 3 in Bain's How to Win with AI series, Proprietary Data, is the series' data chapter, tagline: "a key part of your competitive moat." The core claim is the familiar one, "that data is what gives your agents the context to reason about your specific business rather than generic patterns," and "it is a moat that competitors cannot simply purchase." What earns the note is that the piece immediately complicates its own claim in a way most data-moat arguments refuse to. These are my notes on it; the flagship, the AI-native explainer, and the decision-4 architecture deep dive have their own.

A moat that leaks, refilled daily

The honest twist arrives in one sentence: "your data cannot be synthesized by a competitor, yet its value decays as conditions evolve." Customer behavior shifts, operations change, last year's outcome records describe a business that no longer quite exists. Which relocates the moat from the archive to the operation: "the live signal that your operations generate every day is the part that stays fresh." The moat is not a vault of accumulated records; it is a stock with a leak, and the refill rate is set by how much of what the business does each day gets captured as usable data. A company hoarding a big historical dataset has a wasting asset. A company instrumenting its daily operations has a moat that renews.

A stock-and-flow sketch of Bain's decision 3: a tank labeled your accumulated data, customers, operations, decisions, outcomes, refilled by a highlighted box labeled today's live signal, what operations generate daily, with a leak arrow out the bottom labeled decays as conditions evolve. Caption: competitors cannot synthesize it, and time quietly drains it.

Put simply: the data moat is real, competitors cannot buy or synthesize it, and it leaks. Value sits less in the accumulated archive than in the daily operational signal that refreshes it, so the durable advantage is instrumentation, not hoarding.

Do only the data work the bets need

The second section is a spending discipline. Most enterprise data is not AI-ready, and the tempting response is a generic readiness program, fix all the data, build all the infrastructure. The piece rejects it: "the better approach is to anchor your data investments to your domain bets," the three to five concentrated domains from decision 2. Build the semantic layer, the pipelines, and the governance for the workflows those bets actually run, and let the rest of the estate wait. The piece is candid about the texture of this work: "it is unglamorous work. It does not generate headlines or demonstrate immediate ROI." That candor is the point; a readiness program scoped to everything produces depth nowhere and a stalled transformation with excellent architecture diagrams.

Two panels contrasting data investment strategies, the right one highlighted: generic infrastructure as a grid of many small identical boxes, a little readiness everywhere, depth nowhere, versus three deep domain blocks labeled domain one, two, three, deep readiness where the money is. Caption: unglamorous, no headlines, and the only version that pays back.

Put simply: do not run a general data-readiness program. Pick the domain bets first, then build exactly the semantic layer, pipelines, and governance those workflows need, and accept that the work is unglamorous and pays back through the bets, not through headlines.

Governance at the speed of the agents

The section with the sharpest edge is about governance clocks. "Traditional data governance was designed for a world where humans made decisions and data supported those decisions," committees, access reviews, after-the-fact audits, all running on human timescales. Agents act on live data continuously, so "governance in an agentic world must be embedded in the systems and pipelines themselves," the permissions, checks, and audit trails sitting inline where the agent acts rather than in a meeting that convenes later. The piece then makes its strongest empirical-sounding claim: "running AI through existing governance alone is the single most common reason that transformations stall." No supporting data accompanies it, which is worth noting, but the mechanism is credible: a control system that reviews decisions after the fact cannot govern an actor that completes ten thousand decisions between reviews.

Two panels contrasting governance models, the right one highlighted: review after the fact, a pipeline with a dashed weekly-review box beneath it, humans inspect later if ever, versus embedded in the pipeline, the same flow with highlighted inline gates between stages, every agent call passes the checks. Caption: old governance reviews decisions later; agents have already acted by then.

Put simply: governance built for human decision speed cannot supervise agents. Move the controls into the pipeline, permissions, checks, and audit at the point of action, because per Bain, bolting agents onto existing governance is the most common way transformations stall.

The compounding coda

The close ties the data decision to the series' larger machine. With a shared memory layer, learning accumulates across agents: "without it, every new agent starts from scratch. With it, the tenth agent is faster and smarter on day one than the third agent was after six months." And the page repeats the series' compression of the whole architecture argument, "Agent = Model + Harness," the rented engine inside the owned vehicle. Read against the first section, there is a quiet tension worth keeping: the decay argument says yesterday's data loses value on its own, which softens the series' day-one-compounding urgency more than the authors acknowledge. The reconciliation is the live signal: what compounds is not the pile of data but the instrumented loop that keeps producing it, which is the same conclusion as the warehouse argument reached from the architecture side.

Put simply: shared memory is what makes agent ten start ahead of agent three, and the moat that compounds is the instrumented loop, not the archive. The decay point and the compounding point only reconcile if the daily signal keeps flowing.