top of page
web banner Fika Friday.png

Article 2.1 · The Data Maturity Gap: Why Your AI Is Only as Good as Your Filing System

  • Writer: Will Whawell
    Will Whawell
  • Jul 30
  • 6 min read

T3PS Legal Dynamics · Series 2: AI Readiness — It’s Not a Technology Question ·


Written by Will Whawell. Human intelligence throughout; AI assisted with the drafting.


Refreshed June 2026


There is a phrase that has followed computing since its earliest days, so well-worn by repetition that it risks sliding past without registering: garbage in, garbage out. In the age of AI, that phrase has not aged. It has metastasised. Because AI does not merely reproduce your poor data — it confidently scales it, patterns it, and presents it back to you with the authority of a system that has processed tens of thousands of documents. As Saifr's analysis of AI data quality puts it plainly: "A model will only ever be as good as its data." That is not a caveat. That is the entire architecture of the problem.


Most law firms currently adopting AI are doing so in the wrong order. They are procuring the software before auditing the substrate. They are installing the engine before checking whether there is any road.


The Two Layers of Legal AI

To understand why data quality is so foundational, you first need to understand what legal AI actually consists of. There are two distinct layers, and conflating them is the source of enormous confusion.


The first is the base model — the large language model trained on vast quantities of publicly available text. GPT-4 maybe 5, Claude, Gemini, and their derivatives have absorbed enormous amounts of legal language: case law, statute, legal commentary, academic writing. They are fluent in the grammar and vocabulary of law. They can draft a clause, summarise a judgment, and explain the rule against perpetuities with reasonable competence.


But fluency is not expertise. And general legal knowledge is not the same as firm-specific, matter-specific, jurisdiction-specific capability. This is where the second layer comes in: the content that firms provide to their AI systems to make them useful in practice. Template repositories, clause libraries, historical precedents, engagement letters, billing guides, jurisdiction-specific forms — this is the material that turns a sophisticated autocomplete engine into something resembling a junior fee earner.


The problem is that this second layer — the one that actually differentiates your AI deployment from every other firm's — depends entirely on the quality, structure, and consistency of your firm's own data. And for the majority of UK law firms, that data is in a state that would not pass a basic document management audit, let alone the requirements of an AI training pipeline.


Why General AI Fails Specialised Legal Work

Consider two practice areas where this gap is most acutely felt: residential conveyancing and litigation budgeting.


Conveyancing is procedurally dense, jurisdiction-specific, and heavily dependent on precise wording in precedent documents. The standard AI model knows what a TR1 is. It does not know how your firm handles requisitions on title in leasehold transactions in specific local authority areas, because that knowledge lives in a combination of your experienced fee earners' heads, incomplete file notes, and a shared drive folder last reorganised in 2019.


Litigation budgeting — my own specialist territory — presents an even more pointed version of the same problem. Phase/task/activity coding is the structural framework through which costs are categorised and reported under the CPR costs management regime. Courts expect it. Costs judges reference it. Litigation funders and ATE insurers use it to model risk. But in practice, most litigation solicitors do not record their time by reference to these categories in any consistent way. They do not because the billing system does not enforce it, the partner has never prioritised it, and frankly it feels like administrative overhead rather than legal work.


The result, from the costs consultant's perspective, is chaos. Hours spent wading through email threads, WhatsApp message exports, printed attendance notes, and file documents that exist only as paper — trying to reconstruct a chronology of work from materials that were never designed to be imported, filtered, or analysed. You cannot feed that into an AI system and expect structured output. You will get structured-looking output, which is categorically worse, because it will have the appearance of rigour without the substance.


The Content Library Problem

Progressive firms have begun investing in what are sometimes called content libraries: curated repositories of high-quality legal documents, templates, and precedents that can be used to fine-tune or augment AI systems. The logic is sound. If you provide your AI with sufficient volumes of good, consistent, well-labelled data, it will begin to learn the patterns that define your practice. It will start to reflect your firm's standards, your preferred language, your risk appetite in clauses.


But "sufficient volumes of good, consistent, well-labelled data" is doing enormous work in that sentence. Building a genuinely useful content library requires someone to make decisions about what goes in and what does not. It requires document standardisation: consistent naming conventions, version control, metadata tagging. It requires a view on which precedents represent current best practice and which are legacy documents retained out of institutional inertia. It requires, in short, the kind of information governance discipline that most firms have never applied to their document collections.


MIT research shows that 82% of machine learning projects stall due to data quality issues. In the legal sector, where document management has historically been treated as an administrative afterthought rather than a strategic asset, that figure is entirely credible.


The Data Governance Gap

The deeper issue is not that firms have poor data. It is that most firms adopting AI have not stopped to ask the question. There is no data readiness assessment being conducted before procurement decisions are made. There is no gap analysis between the current state of document management and the requirements of an AI deployment. The procurement conversation is dominated by functionality, pricing, and vendor demonstrations — and the demonstrations, naturally, are conducted on clean, curated datasets that bear no resemblance to the firm's actual information estate.


Data governance — the policies, processes, and accountabilities that determine how information is created, stored, classified, and retrieved — is not a glamorous topic. It does not generate the same boardroom excitement as a live demonstration of AI drafting a contract in thirty seconds. But it is the unglamorous foundation on which everything else depends.


This is precisely why firms with existing quality management infrastructure are so much better positioned than they typically realise. An organisation that has implemented ISO 9001 has already grappled with documented processes, defined responsibilities, and controlled outputs. An organisation certified to ISO 27001 has already classified its information assets, assessed risks, and implemented access controls. These are not peripheral activities. They are the direct antecedents of AI readiness. The work is not finished — AI introduces specific requirements around data labelling, training pipelines, and output monitoring that go beyond traditional quality and security management — but the bones are already there. For a process-mature organisation, the path to AI readiness is iterative, not transformational.


For a firm that has never invested in process maturity, however, the path is considerably steeper. And the temptation to shortcut it — to buy the AI and hope the data problem resolves itself — is both understandable and dangerous.


What Good Looks Like

A firm genuinely ready to extract value from AI has, at minimum, addressed the following:


Structured time recording. Time is recorded by reference to phase, task, and activity codes where relevant. Not because the costs consultant asked for it once, but because the practice management system is configured to require it and fee earners understand why it matters.


Controlled document management. Precedents, templates, and standard-form documents are held in a single, authoritative repository. There is a version control policy. There is a named owner for each document type. Documents are not duplicated across individual hard drives, email attachments, and shared drives simultaneously.


Metadata discipline. Documents are consistently named and tagged with matter type, jurisdiction, date, and status. Not aspirationally — actually. You can find things without ringing the fee earner who opened the file three years ago.


Defined data flows. The firm can map where client data enters its systems, how it moves between systems, and where it exits. This matters for AI both because it informs training data selection and because it is a prerequisite for the data protection and information security obligations that AI deployment triggers.


None of this is about technology. All of it is about discipline, culture, and sustained management attention. That is precisely why it is so hard. And precisely why it must precede procurement rather than follow it.


The firms that are going to extract genuine, sustained value from AI are not necessarily the ones with the largest budgets or the earliest adopters. They are the ones that took the time to build clean foundations. Everything else, however sophisticated, is just an expensive way to automate your existing chaos.


Questions worth sitting with:

1.    If you ran a data quality audit across your firm's document estate today, what proportion of your precedents and templates would meet the standard you would require of an AI training dataset?


2.    For your most data-intensive practice area, can you genuinely map where all relevant information resides — and is any of it locked in email threads, paper files, or individual fee earners' heads?


3.    If your AI produces confident-sounding output based on poor underlying data, who in your firm has the expertise and mandate to catch that before it reaches a client?


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page