Model training

    A private model tailored to your organisation

    Train a private model on work your team has already published and approved. New translations follow your house style, editorial voice and document-formatting rules — including the fonts and type treatments your brand requires.

    No credit card required · Your first translation in minutes.

    Trusted by teams at
    Franklin TempletonLaSalleKearneyYaraShimadzu

    Teach the model what makes your output yours

    A generic model can translate the words. A model trained for your organisation learns the decisions that make the finished work recognisably yours.

    House style

    Train on published articles, bulletin scripts, reports, releases and other approved work. The model learns your register, sentence rhythm, headline conventions, quotation style and the constructions your editors consistently choose.

    Format style

    Carry the visual rules of the publication or brand into every translated file: specific font families, weights and sizes, heading hierarchy, line length, table treatments and recurring page structures.

    One model per publication or brand

    Keep distinct editorial identities distinct. Each masthead, channel, imprint or brand can use a model trained on its own archive rather than being flattened into a group-wide average.

    When a generic translation is not enough

    The words may be correct while the finished document still feels unlike anything your organisation would publish.

    01

    “This reads like a machine wrote it”

    Correct wording is only the baseline. Training captures the register, cadence and sentence choices that make the output read like your newsroom, publication or organisation.

    02

    “It is correct, but it is not how we write”

    Every firm has sentence habits: how obligations are phrased, how conditions are ordered, how formal the register runs. None of it is expressible as a term pair.

    03

    “The regulator approved this phrasing”

    “In a highly regulated industry, if we put out a document in English that goes to the regulators and they approve that language as is, when it gets translated it has to have the same context.” Where whole constructions were approved rather than individual words, tuning is the mechanism that carries them across.

    Train on decisions your editors already made

    Your archive is more precise than a description of your style. It shows how the organisation actually writes, edits and formats finished work.

    Approved bilingual document pairs

    Filings, contracts, manuals, releases and other work your team has already translated, edited and signed off. The archive shows which constructions your reviewers keep — and which ones they quietly rewrite every time.

    • Bilingual pairs, aligned and cleaned before anything is trained
    • Source files where available, so text can be aligned accurately
    • As far back as you hold; recent work is weighted more heavily
    • Inconsistent history teaches inconsistency, so cleaning comes first
    • No fixed minimum — we scope it against the training you actually need
    Archive sweepaligned
    Source
    Approved

    The guidance behind the work

    Style guides, editorial standards, brand guidance, templates and typography specifications make explicit the rules the archive demonstrates. Together they give the model a clearer target.

    • Style guide and editorial standards
    • Brand guide: voice, capitalisation and market-specific treatments
    • Approved font families, weights and sizes by document element
    • Templates and examples for each publication, programme or audience
    Talk to sales
    Style guide+ brand guide

    When the guide and the archive disagree

    They will. It is the most useful thing the process turns up, and the reason it is worth doing even if you never buy the model.

    The guide is what you believe you do

    A written style guide records the rules a house thinks it follows. It is usually years old, written by someone who has left, and silent on most of what actually distinguishes the writing.

    The archive is what you actually did

    On a recent media account, three rules in the client’s own written spec were contradicted by their own published output — how a wire sign-off was handled, when headline case applied, whether a byline appeared on wire sports. Nobody had noticed, because no single person reads enough of the output to notice.

    You decide which one is the rule

    We report every conflict before training rather than silently picking a side. The output of that step is a corrected style guide you own, whatever you decide about the model.

    What your private model learns

    Training carries the editorial decisions that normally live in an editor's eye, a house guide and years of approved output.

    01

    How your writing sounds

    Register and sentence habit: how obligations are phrased, how conditions are ordered, how formal the voice runs, where your house prefers a short sentence and where it does not. None of this is expressible as a term pair.

    02

    Your house style, including the no-go list

    The rules your sub-editors apply without thinking: date and number formats, how figures are written out, headline case, quotation conventions, how a name is rendered on second mention — and the words that are simply not used here, whatever the dictionary says.

    03

    Your brand guide

    Product and entity names, capitalisation, approved taglines per market, the tone the brand is contracted to hold, and the do-not-translate list. Brand rules are usually the first thing generic translation output breaks, and the most visible when it does.

    04

    Context by document type

    A bulletin, long-form feature and corporate report should not use the same register. Separate models can learn the phrasing and level of formality appropriate to each body of approved material.

    One workspace, several houses

    A publishing group is not one house style. It is a title with a formal register, a title that writes short, and a brand book that forbids in one what it requires in another.

    A private model per title, chosen at translation time

    Each masthead, imprint or brand gets its own model — its own voice, editorial conventions and format specification — trained on that title’s archive rather than the group average. Choose the model when you translate, and the same source can be produced in two houses’ styles without being rewritten twice.

    • One private model per title, imprint or brand
    • Trained on that title’s archive, not the group average
    • Convert an approved piece from one house style into another
    • Keep each publication’s editorial identity separate
    Talk to sales
    Same sourcetwo houses
    Title A
    Title B

    Your format rules run as checks

    Anything that can be written as a rule becomes a check that runs on every file before it reaches a person.

    Typography

    Font family, weight and size, per element, checked against your template. If your spec says body copy is 10.5pt and a translated table has dropped to 9pt to make the text fit, that is a flagged deviation rather than something a reviewer is expected to notice at four in the afternoon.

    Structure

    Heading hierarchy, numbering continuity, cross-references, table and figure captions, section order, and the line-length limits that matter when the output has to fit a column, a caption or a cell.

    Convention

    Dates, currency, units, decimal separators, entity names, the do-not-translate list, and the no-go words from your style guide. Every hit reports where it is, so the fix is a jump rather than a re-read.

    What “house style” means depends on the house

    The mechanism is the same everywhere. What it is asked to hold on to is not.

    Media & broadcast

    Not subtitling — the document layer wrapped around the content. Bulletin scripts, listings, press material, inbound brand mandatories that have to reach local production teams, and the annual and sustainability reports a listed broadcaster still has to file. Names spelled the way they went out last season, because a returning strand cannot change its own spellings halfway through.

    Publishing

    Several mastheads under one roof, each with its own voice, its own banned list and its own idea of a headline. Tuning per title stops a group flattening into one register, and lets a piece filed for the newspaper desk be produced for the bulletin desk in the bulletin desk’s style rather than rewritten by hand.

    Legal

    Defined terms, how obligations are phrased, clause numbering that has to survive intact, and the constructions a regulator has already approved. Where whole phrasings were signed off rather than individual words, only tuning carries them across.

    Finance

    Disclosure language, figure and date conventions, entity names, and the boilerplate sections that must read identically year on year. Drift between two annual reports is the failure everyone notices and nobody catches in review.

    Medical & pharma

    Regulated terminology, dosage and unit conventions, and one document set that has to hold a clinician register in one section and a patient register in the next without either bleeding into the other.

    Industrial & manufacturing

    Part numbers, safety and warning language with fixed formats, and manuals reused across model years, where a new revision must match the last one everywhere it was not deliberately changed.

    A measured tuning cycle

    We judge the result on your documents, your format specification and your editors' changes — not on a generic public benchmark.

    Benchmark, train, review, repeat

    We establish a baseline on representative documents, train against approved material, then review the output with your specialists. Their corrections guide the next pass and show whether recurring edits are falling.

    • Reviewed and re-run rather than delivered once
    • Measured on your documents, not a public benchmark
    • Measure editing time and repeated corrections before and after
    Discuss model training
    Tuning cyclemeasured

    Yours alone, and reversible

    Your custom model serves your workspace only and never feeds a shared model. Training is opt-in and contracted separately, so ordinary document translation never enrols your material in model training.

    • Serves your workspace only, never a shared model
    • Opt-in and separately contracted
    • Training data is withdrawable
    Your workspaceonly

    What makes training work

    A tailored model needs a clear editorial identity and representative approved material. More content only helps when it reflects the output you want next.

    01

    It cannot invent a house style you have never written down

    The corpus is the specification. If your approved translations disagree with each other, the tuned output will disagree with itself in exactly the same proportion.

    02

    It cannot generalise beyond the material it saw

    A model trained on loan documentation will not carry that register into employment contracts, and one trained on a magazine will not write the evening bulletin. Scope each model to the document types and publications it will actually serve.

    03

    It cannot be allowed to invent to satisfy a pattern

    A model that has learned your house always opens with a day marker and a photo credit will supply one whether or not the source had it. No house-style rule licenses a fact: style is constrained to the material it was given, and anything the source did not say is left out rather than smoothed over. It is the failure mode to test for on day one.

    04

    It does not remove the reviewer

    On anything filed, signed or published, bilingual review stays. What tuning reduces is how much that reviewer changes — and that is the only number worth measuring.

    Frequently asked questions

    Which language pairs support tuning?
    All 120+. Tuning is not restricted to a subset of the language set you already translate in.
    How far back should the archive go?
    As far as you hold it. Volume helps, but consistency helps more: a clean decade beats a messy two decades, and recent work is weighted more heavily so a style that has moved does not get dragged back. We align and clean before anything is trained, and tell you what we had to discard.
    Can it match our font and type sizes?
    Yes. Type is treated as part of the specification rather than an afterthought: font family, weight and size per element, taken from your template and checked on every output. The usual failure — a table quietly reset to a smaller size so the longer target text fits — is reported as a deviation instead of shipping.
    We publish several titles with different styles. One model or several?
    Several private models, one workspace. Each title is trained on its own archive and keeps its own voice and format specification. Choose the appropriate model when you translate, or produce the same source in another title’s established style.
    Is our material used to train anything shared?
    No. A model tuned for you serves your workspace only.
    Can the model also use our approved terminology?
    Yes. Your glossary works alongside the private model, while the model learns the broader editorial and contextual patterns that individual term pairs cannot express.
    How do we know it worked?
    By how much your reviewers change. Measure editing time on a fixed set of documents before and after, not a score in a report.

    Make every approved edit work harder

    Turn the editorial and formatting decisions behind your approved work into a private model tailored to your organisation.