Applying identical review depth to every document in a set does not produce uniform quality. It produces uniform shallowness, because the reviewer's hours are fixed and the set is not. The workable alternative is to tier documents by what the text does — operative provisions, schedules and definitions, historical annexes, correspondence — assign a defined depth of check to each tier, have the tier set by the person accountable for the matter rather than by the translator, and record the tiering decision with its reasoning. The tiering record is what makes the shallow pass on tier four defensible, because it shows the shallowness was chosen rather than run out of.
This guide covers how to draw the tiers, what changes between them, and how to write the decision down.
The Arithmetic of Uniform Review
Start with the constraint that governs everything else. A reviewer has a fixed number of hours. A cross-border transaction has a fixed number of documents. Divide the first by the second and you have the real review depth, whatever the policy says.
Take a diligence exercise with nine hundred foreign-language documents and one qualified reviewer for a week. Uniform treatment yields something under two minutes per document — which is not review, it is page-turning. Meanwhile the twelve documents that carry every material obligation in the deal receive the same two minutes as a 2019 utility invoice.
The failure is not laziness; it is a policy that sounds rigorous and is arithmetically impossible. "Everything gets full human review" is the sentence that guarantees nothing does. Tiering does not add rigour to the set as a whole. It concentrates the rigour you actually have where the consequences actually are, and it makes the trade-off explicit instead of accidental.
Tiering by What the Text Does
Draw the tiers on function, not on document type. A "contract" spans three tiers on its own.
Tier 1 — operative text. Provisions that create, limit or transfer obligations: the operative clauses, indemnities, limitation of liability, termination, governing law, payment terms, covenants and any warranty schedule that is incorporated by reference. If a dispute would turn on these words, they are tier 1.
Tier 2 — definitions, schedules and figures. Defined-term tables, price schedules, financial statements, calculation methodologies, tables of rates. Errors here are usually numeric or terminological rather than interpretive, but they propagate: a defined term rendered inconsistently breaks every clause that uses it.
Tier 3 — historical annexes and background. Superseded agreements, prior correspondence attached for completeness, historical minutes. Read for comprehension and for surprises, not for enforceability.
Tier 4 — routine correspondence and administrative material. Notices, transmittal letters, invoices, filings that exist only to demonstrate that something was filed.
The first cut is usually one hour of work and reshapes the entire budget.
The Person Who Sets the Tier
The tier must be set by whoever is accountable for the outcome — the instructing partner, the deal lead, the compliance owner. Not the translation vendor, who has no view of the matter. Not the reviewer, who has an incentive to tier down under deadline. Not a coordinator applying page counts.
The mechanics that work in practice: the matter owner tiers the index at kickoff, before any file is translated, working from document titles and a five-minute conversation about the deal's risk shape. Anything genuinely unclear goes to the higher tier by default. Tiers are then frozen except by explicit re-tiering, which anyone touching the document can request and only the owner can grant.
That last rule is the one that matters. Reviewers and translators will find documents that are misfiled — an annex that turns out to contain a change-of-control provision, a "routine" letter that concedes a liability. The escalation path must be a two-line email, not a change-control process, or people will simply carry on with the wrong tier.
What Changes Between the Tiers
Depth has to mean something concrete, or tiering becomes a label with no operational content. Define the checks per tier and write them into the workflow.
Tier 1 gets full post-editing in the sense used by ISO 18587 — a bilingual reviewer comparing source and target segment by segment — plus a second read by someone qualified in the destination jurisdiction, plus mechanical verification that every clause number, cross-reference and citation appears in both versions.
Tier 2 gets full bilingual post-editing plus automated numeric reconciliation: every figure, date, percentage and currency symbol extracted from both versions and compared. Most tier-2 errors are caught by a script, not by reading.
Tier 3 gets a monolingual review of the target for sense, fluency and anything anomalous, with spot bilingual checks on passages that look consequential.
Tier 4 gets automated checks only — terminology enforcement, numeric extraction, completeness — and no human read unless a check fires.
The TAUS post-editing guidelines are a useful reference point for describing these levels in language a vendor will recognise.
Signals That Move a Document Up a Tier
Static tiering by title will misclassify some documents. Cheap automated signals catch most of them, and they run over the whole set rather than the sampled part.
Terminology hits are the strongest: a document tiered as routine correspondence that contains three defined terms from the operative agreement is not routine. Numeric density combined with unusual units, unexpected clause-numbering patterns, and the appearance of jurisdiction or governing-law language all warrant a second look.
Automatic quality estimation adds a further filter. Neural evaluation metrics such as COMET score machine output against human judgement patterns, and reference-free variants can flag segments the model itself finds difficult. Used as a router rather than as a verdict, a low-confidence cluster is a reasonable trigger for human eyes. Where you need a shared vocabulary for describing what was actually wrong, the MQM error typology gives you categories — accuracy, terminology, fluency, style — that survive a conversation with a regulator better than "quality issues".
Documenting the Tiering Decision
This is the part that turns an operational shortcut into a defensible one, and it is usually skipped.
For each document, record the tier assigned, who assigned it, when, and the reason in one line. For the set as a whole, record the tiering criteria in force and the escalation events — every document re-tiered, in which direction, and what prompted it. Keep this in the same register as the rest of your translation process documentation, keyed on the same document IDs.
The reason is precise. If a translation error later surfaces in a tier-4 document, the question will be whether the process was reasonable. A record showing a considered allocation of review effort by a named accountable person, with a working escalation path and documented escalations that actually happened, is a good answer. No record at all leaves you arguing that a two-minute pass on nine hundred documents was a deliberate methodology, with nothing to show that it was.
Running the Model Across a Data Room
In practice, on a live diligence exercise, the sequence looks like this.
Translate everything first, machine-first, formatting intact — you cannot tier what you cannot read, and untranslated documents are the ones that hide problems. Bluente handles that first pass across 120+ languages and 22+ file types with tables, numbering and cross-references preserved, which matters because tier decisions are made from structure as much as content. Then tier the index, then apply the per-tier checks, then work the escalations.
Practitioners hit the same shape of problem from the budget side — a thread on large-scale multilingual translation on a constrained budget walks through exactly this trade-off between coverage and depth. The tiering answer is the same either way: decide where the depth goes before the deadline decides for you. Our notes on risk discovery in foreign-language contracts cover what to look for once the set is readable.
Sources and Further Reading
Large-scale multilingual translation on a budget, r/machinetranslation — practitioners trading coverage against review depth on real volume
Related Reading
Last reviewed 24 August 2026 by the Bluente document engineering team, who build and test the pipeline described here. This is general operational guidance, not legal advice; where a filing rule or a court sets a required standard of review, that rule governs.
Depth is a budget. Spend it where the obligations are. Try BluTranslate free.