How law firms can measure value, assign ownership and decide when to build, expand or stop.
The law firm's guide to building legal AI: Complete guide · Part 1: Ownership · Part 2: Integration · Part 3: Commercial value
A review tool produces its first draft in minutes. The demonstration ends there. On a live matter, someone still has to prepare the files, investigate missing evidence, correct the output and approve the work.
The investment case must include those steps. The useful unit of measurement is a lawyer-approved work product, delivered to an agreed standard.
For a cross-border practice, value might mean turning an urgent batch of foreign-language documents into a reviewable issues list before a deadline. It might mean absorbing more matters without delaying existing clients. Start by identifying the service improvement the firm wants, then measure whether the proposed system delivers it.
Use Building legal AI: what law firms should own, buy and connect to identify the components your proposed system needs, then include their implementation, review and maintenance requirements in the investment case below.
Before approving a wider build, ask for four outputs: an agreed quality result, a comparison with the current process, evidence that the team can use it, and a named owner with a maintenance budget. A successful demonstration supplies only part of that case.
Choose one outcome the practice cares about
Define a comparable work product: a diligence issues list for specified provisions, a loan comparison or an evidence chronology. Record the required standard, eligible documents and the route for incomplete work.
Measure the current process, including preparation, waiting, review and rework. A result for clean contracts should not stand in for a scanned multilingual bundle. Part 2 shows how to account for those steps.
Distinguish capacity, cost and revenue
Time released is capacity. It becomes revenue when the firm wins, delivers and collects payment for additional work; it becomes a cash saving when the relevant expenditure changes.
Capacity can still matter without either result: the team may meet an urgent deadline or spend more time on substantive advice. Name the benefit and measure it.
Another route is a specialist product or service. A&O Shearman's April 2025 announcement with Harvey described agents intended for external sale and software revenue sharing. That is one commercial model to assess against client demand, not a default for every firm. A&O Shearman’s announcement.
Include the work after generation
For each completed work product, record document preparation, analysis, lawyer review, rework and the allocated cost of running the system. Include supplier charges, support, legal maintenance and implementation costs without counting the same expense twice.
Here is a deliberately fictional illustration. It assumes one comparable review batch per month, equivalent acceptable quality and an internal labour cost of US$100 per hour. The hourly figure is an internal cost assumption, not a client charge-out rate. These are modelling assumptions, not Bluente prices or observed results.
Cost per monthly batch | Existing process | Proposed AI workflow |
|---|---|---|
Preparation, analysis, review and rework | 80 hours × $100 = $8,000 | 56 hours × $100 = $5,600 |
Incremental processing charges | — | $600 |
Allocated ongoing support and maintenance | — | $1,000 |
Allocated implementation cost | — | $500 |
Total modelled cost | $8,000 | $7,700 |
In this illustration, the workflow releases 24 hours of capacity but reduces modelled cost by only $300 per batch. Four additional hours of review would bring its cost to $8,100, above the baseline.
Compare modelled savings with cash and capacity separately. For this example, assume the $1,000 support allocation and $500 implementation allocation are fixed monthly amounts. If volume falls to half a batch, and labour and processing costs fall proportionally, the proposed route costs $4,600 against a $4,000 baseline. This is a sensitivity illustration, not a forecast.
Replace those assumptions with the firm's own evidence. Include any baseline technology costs that differ between options. Keep shared overhead out of the incremental comparison unless it changes, or allocate it consistently to both routes. Do not add a maintenance allocation for time already included in the labour line.
For mixed work products, segment the calculation. A short, clean contract and a scanned multilingual bundle should not be treated as equivalent units merely because each counts as one job.
Keep quality beside the economics
Agree the release criteria before the pilot. For example, the practice sponsor might require no unresolved material failures in the agreed test cases, a reconciliation of all input files, and a documented route for exceptions. Passing a test set is evidence for a bounded decision, not a guarantee of error-free future work.
Use a compact scorecard that makes both failures and effort visible:
Measure | What to record | Who should own it |
|---|---|---|
Material findings | Expected material issues found, missed or incorrectly asserted | Supervising lawyer |
Completeness | Files and pages received, processed, excluded and unresolved | Legal ops or the matter coordinator |
Review burden | Minutes to verify and correct each comparable work product | Reviewing lawyers |
Delivery | Receipt-to-approval time, including queues and retries | Legal ops |
Use | Eligible matters offered the workflow, completed through it and abandoned | Practice owner |
Economics | Incremental costs, capacity released and the use made of that capacity | Finance and the practice sponsor |
Choose thresholds with the practice group for the intended use. Separate clean files from difficult scans, and routine matters from exceptions. An overall average can hide the cases that make the tool unsuitable for a particular team.
A 2024 evaluation found hallucinations in the tested legal research products despite retrieval-based design. That historical result is a reason to test the actual task, not a current failure estimate for a different product or workflow. Magesh and colleagues, “Hallucination-Free?”.
Decide how the client shares in the benefit
Discuss the service offered, its scope and the basis of charges. Faster turnaround, clearer reporting or more predictable delivery may support a different commercial arrangement, but a technology investment does not establish entitlement to a particular fee.
In US guidance, ABA Formal Opinion 512 says hourly billing must reflect actual time spent, and fees and expenses must remain reasonable. Flat fees are also subject to reasonableness; AI efficiency does not automatically justify keeping an unchanged fee. Applicable rules and engagement terms require assessment in the relevant jurisdiction. ABA Formal Opinion 512, July 2024, fees discussion.
Build the commercial proposition around the work delivered and the responsibilities retained. If the firm offers an ongoing service, specify what is covered, what needs separate advice and how difficult cases will be handled.
Fund ownership, adoption and lawyer development
A pilot needs someone accountable for its future. Assign a legal owner to maintain the method and approve changes, a technical owner to maintain the service, and a practice sponsor to assess value. These are responsibilities; smaller firms may combine roles or use external partners.
Knowledge maintenance includes the asset and portability decisions in Part 1: Who owns your law firm's AI advantage?. Technical maintenance includes integrations, access changes, incident handling and retesting when a component changes.
Training needs its own allocation. Linklaters' July 2024 Copilot announcement described mandatory foundational AI training and training specific to Copilot alongside the rollout. Linklaters, “Linklaters rolls out Microsoft 365’s Copilot globally”.
For a proposed review workflow, training should teach lawyers how to inspect evidence, recognise an incomplete result and escalate uncertainty. Measure successful use in real work, rather than counting accounts or logins alone.
Also plan how junior lawyers develop the judgment required to supervise the tool. One approach is to ask them to analyse selected documents independently before comparing their conclusions with the system and a senior reviewer. The aim is to retain deliberate practice while changing how routine work is delivered.
Use a staged investment decision
Before funding a broad platform, compare options against the same task. Could an existing product meet the requirement? Would configuration or a narrow integration close the gap? What additional benefit would custom development provide?
Quality, adoption and full delivery economics should support expansion. A failed gate points to a different response: refine, redesign, change scope or retire.
Use three review points:
Before development: define the deliverable, evidence set, client constraints, owners, baseline and acceptance criteria.
Before live expansion: demonstrate acceptable quality, review effort, failure handling and delivery time on representative work.
After live use: compare actual volume, support demands, client value and total cost with the case for investment.
Scale when the evidence supports it. Refine the workflow when a specific problem is fixable. Retire or replace it when maintenance exceeds its value, adoption remains low or an available product meets the requirement more effectively.
Specify the evidence that would change the decision in advance. Otherwise, a pilot can continue consuming time because no one defined what completion—or failure—would look like.
Put a bounded request in front of the partner
The approval paper should state the service to be improved, eligible matters, quality results, current and proposed costs, expected use, owners and the next decision date. Attach the unresolved issues rather than hiding them in an average score.
Separate three requests: permission to investigate, permission to run a controlled pilot, and funding for ongoing use. A practice without available review capacity or a maintenance owner may be ready for discovery but not deployment.
For an hourly matter, ask how the client benefits from reduced time and how released capacity will be used. For a scoped recurring service, ask whether more predictable delivery supports the proposed engagement terms. For a client-facing product, assess support and commercial obligations separately from the internal tool. This keeps the revenue case tied to the service the firm can deliver.
Evaluate the document component on the same basis
For scanned and multilingual work, compare the time from receipt to an LLM-ready file and then to an approved work product. Include OCR exceptions, parsing corrections, translation where needed and manual formatting work. Evaluate document extraction on the difficult scans as well as the clean native files that enter the practice.
Bluente can be evaluated as the document infrastructure layer for OCR, parsing and translation. Its extraction-only OCR mode returns HTML without translation, allowing the firm to assess scanned-file preparation separately from language conversion.
Bring a defined task, representative documents and an acceptance standard to that evaluation. The result should give the practice sponsor evidence for a decision about service quality, capacity and continuing cost.
Assessing a multilingual workflow? Book a Bluente demo to evaluate document parsing and translation against your team's delivery requirements.
Explore the series: Part 1: Knowledge ownership · Part 2: Document workflows · Complete guide