Malcolm Angus
← Sources

Show notes

McKinsey: "AI data readiness is the key to scaling impact"

2026-08-12


McKinsey's June 2026 report, AI data readiness: the key to scaling impact, starts from a deceptively simple observation: AI makes enterprise data look easy. A contract becomes a summary, a customer transcript becomes a recommended action, a policy becomes an answer. Underneath, AI systems are pulling that data apart and putting it back together again and again, across documents, systems, prompts, and workflows, and scaling AI means ensuring all of it is treated as truth. Only 7 percent of companies have fully scaled AI across their organizations, and the report's argument is that data, not the model, is what stands in the way. These are my notes on it.

Data is the constraint, not the model

AI systems lean on unstructured content, documents, emails, call transcripts, and video, and each file expands into multiple representations, text, tables, and images, that get reused across applications, so a small data issue spreads and quickly becomes a large one. Digitizing that content and making it searchable is not enough: the report's point is that searchability does not equal usability, and tools alone, vector databases, model gateways, retrieval pipelines, and retrieval-augmented generation, are not a silver bullet either. The evidence it leads with is Exhibit 1: more than two-thirds of high-performing companies name data as the primary obstacle to scaling gen AI, ahead of risk, the operating model, technology, strategy, and talent. AI data readiness, in the report's definition, is connecting structured and unstructured data into a governed, traceable, and reusable foundation, one that also includes metadata, lineage, and the tools and skills AI systems use to act on data consistently.

A horizontal bar chart from McKinsey's 2024 AI survey: among high-performing companies, data is the top-ranked challenge to scaling generative AI, highlighted in gold and labeled more than two-thirds, well ahead of risk and responsible AI, operating model, technology, strategy, talent, and adoption and scaling. Caption: data outranks tech, talent, and strategy as the top blocker to scale.

Put simply: Only 7 percent of companies have fully scaled AI, and more than two-thirds of top performers say the blocker is data, not the model. Readiness means making structured and unstructured data governed, traceable, and reusable.

Ready is relative to the use case

The report rejects the idea of readiness in general. There is no single bar; a company should define what "good enough" means for each use case, based on business need and the risk profile of the data and the process (Exhibit 2). An internal productivity copilot needs moderate quality and basic controls with a human reviewing. A knowledge-search assistant needs citation, lineage, and access controls. A customer-facing assistant needs runtime monitoring, guardrails, and controls on personal information. A regulated decision, a loan or a claim, needs very high quality, full lineage, compliance, auditability, and human approval. The objective is not perfect data before you start, but data that is good enough for the job in front of you.

Put simply: Readiness is not one standard. It scales with the stakes, from a low-risk internal copilot to a regulated decision that needs full lineage and human approval.

Four shifts, six disciplines

McKinsey frames four technology shifts that are stalling AI scaling: unstructured data is becoming a harder enterprise challenge to trace and defend; the risk surface expands as AI retrieves, recombines, and generates data, moving responsibility to the application layer; AI generates and reuses data faster than periodic governance can keep up; and the resulting fragmentation is making chief data officers central to AI enablement. To meet them, the report says organizations must evolve their core data disciplines rather than replace them, applying six of them across all four kinds of data: structured records, unstructured content, derived artifacts, and the governed tools and skills that let large language models and agents use data consistently.

A framework grid of the six data disciplines McKinsey says AI data readiness requires, evolved for AI rather than replaced: observability, data quality, metadata, data lineage, governance and controls, and platform and tooling, sitting over a gold band naming the four kinds of data they apply across, structured, unstructured, derived artifacts, and governed tools. Caption: evolve the core data disciplines for AI, do not replace them.

The six disciplines: observability, making the assembly of context and the production of outputs visible, not just the movement of data; data quality management, extended across extraction, chunking, retrieval, and generation rather than checked once at ingestion; metadata management, which becomes the control layer for unstructured artifacts, with fine-grained schemas and links to core entities like customers and contracts; data lineage, tracing every derived artifact, which chunk was retrieved and how a prompt was built, back to its source; governance and controls, enforced at retrieval and generation and not only at storage, because compliant storage does not guarantee compliant outputs; and platform and tooling architectures, the reusable extraction pipelines, shared embedding and indexing, common retrieval layers, and standardized guardrails that let the other five scale.

Put simply: Four shifts, unstructured data, an expanding risk surface, velocity, and fragmentation, demand six evolved disciplines: observability, data quality, metadata, lineage, governance, and platform, applied across four kinds of data.

The call to action

The report's closing charge is a change in mandate. Chief data officers must shift from owning data pipelines, models, and warehouses to owning the standards, the control plane, the reusable data products, and the governed tools and skills that make AI outcomes reliable and repeatable, wherever applications are built. It lays out six concrete steps: put structured and unstructured data into governed data products, with canonical schemas, entity alignment, and artifact-level lineage; build shared foundation services once and reuse them, the retrieval, policy, and monitoring stack that becomes the control plane; let business units build on that common infrastructure instead of each rolling their own; manage derived artifacts like embeddings and indexes as versioned, owned enterprise assets; govern semantic consistency so an "active customer" returns the same underlying data whether queried in SQL, filtered in search, or retrieved by a vector-based assistant; and measure readiness on four metrics, reuse, reliability, governance, and scalability.

A before-and-after diagram of the chief data officer's mandate per McKinsey: on the left, muted boxes for what the CDO used to own, data pipelines, models, and warehouses; a gold arrow labeled AI at scale points right to gold boxes for what the CDO now owns, standards and the control plane, reusable data products, and governed tools and skills. Caption: from owning pipelines and warehouses to owning standards and governed tools.

Put simply: Treat data as a core enterprise asset. Own the standards and the control plane, not just the pipelines, and measure readiness on reuse, reliability, governance, and scalability. Do that, the report concludes, and you can scale AI with consistency, safety, and speed.