Latest posts Visit blog

Adding a second language market is quick these days: the machine stage translates ten thousand product texts overnight, and by morning the catalogue looks complete. Whether it also sells is another question. 76 per cent of online shoppers prefer products whose information is available in their own language (CSA Research, 2020), and 60 per cent rarely or not at all buy from English-only sites (CSA Research, 2014). Language therefore shapes revenue - but only if it holds up. This article shows how to build the translation pipeline of your online shop around clear review stages: which content needs which depth of review, how a glossary and style rules steer the machine stage, how quality becomes measurable through error classes, and what needs to be right technically so the language versions are actually found.

Why language helps decide the purchase

The link is well documented: people who cannot read a product page effortlessly buy there less often. 76 per cent of online shoppers prefer products with information in their own language (CSA Research, 2020), and 60 per cent rarely or not at all buy from English-only sites (CSA Research, 2014). This is not a question of foreign language skills but one of effort and trust: in their mother tongue, people grasp technical details faster, scan terms more reliably and enter payment data with less hesitation. Within the European Union there is the added fact that 24 official languages exist side by side (European Commission). A single English-language shop window therefore covers only part of the single market linguistically, even if every country can be supplied technically.

Legally, nobody is forced to offer the shop in every language: the Geo-blocking Regulation requires traders not to restrict access to their offering based on where the customer comes from, but it expressly does not require the website to be translated into the customer's language (European Commission). In business terms the calculation looks different: localised product descriptions, reviews and checkout processes perform measurably better in international markets than English-language variants (CSA Research 2014). Anyone planning the internationalisation of their shop should therefore not treat language as an afterthought but decide it together with prices, payment methods and shipping routes - topics we explore in our article on multi-currency support and local prices.

Language does not stop at the product page

What gets translated is usually what is visible: title, description, category text. Yet the edges also shape the purchase - filter values, delivery time statements, form error messages, payment notes and the order confirmation email. A half-translated checkout catches the customer's eye at exactly the moment they are already hesitating: just before payment.

What the machine stage delivers - and where it breaks

Machine translation is no longer a novelty but an established building block with a supplier market of its own. With structured content the technology plays to its strengths. Product catalogues, attribute lists, category pages and technical specifications follow recurring patterns and contain little irony and few allusions. Here the machine delivers consistent results across tens of thousands of records, at a unit cost that a purely manual approach hardly reaches. That is precisely why post-editing of machine output has established itself as a discipline of its own: it accounted for around 46 per cent of workflows at language service providers in 2024 (Nimdzi). The machine does not replace review, it shifts where the work sits.

Things break down where meaning does not sit in the sentence but in its surroundings. Brand voice, cultural context and domain terminology are the three areas where machine output regularly goes wrong: a component suddenly carries three different names in the target language, a negation flips, a measurement moves into the wrong unit, a sober legal text picks up a promotional tone. Anyone already using AI-generated product descriptions knows the pattern from text production - in translation it plays out across all target languages at once. One research finding is instructive here: search quality in multilingual shops depends on translation quality only up to a threshold; above that, better translation adds little for search (Amazon Science). That is a strong argument for placing review effort deliberately rather than spreading it evenly - just as it has proven useful with other AI building blocks in the shop.

Sheer volume is not a quality signal

A catalogue that appears in four languages overnight looks like progress. The cost of weak language versions typically shows up later: in returns because a measurement was misread, in support requests because a shipping condition stays unclear, and in abandoned carts because the checkout remained half German. Skipping the review does not remove the effort, it only moves it to a more expensive place.

Classify content: not every text needs the same depth of review

The most important decision is made before the first translation: what content exists, and what happens if a sentence goes wrong? The depth of review follows from that risk question. In practice three to four classes are enough; more makes the rule set unwieldy and leads to nobody following it.

  • Product data and attributes: high volume, fixed structure, little room for interpretation. Here the machine works with a glossary and protected placeholders, and review happens as a sample per batch. Well maintained source data from a PIM system noticeably reduces review effort because it limits the number of special cases.
  • Category and help texts: medium volume, explanatory in character, visible in many places. After the machine stage these texts go into post-editing, because tone and clarity already contribute to the sale here and because they often need more context in the target language than a literal rendering provides.
  • Brand and legal copy: small volume, high impact, high risk. Home page, campaign copy, right of withdrawal, warranty and safety information. Here the machine delivers a draft at best; sign-off comes from a person who knows the target market.
  • User-generated content: reviews and customer questions can be machine translated and labelled as machine translated. Reviewing everything in advance is unrealistic with a constant inflow - here a reporting route replaces sign-off, complemented by the original language version in an expandable area.

This classification is not bureaucracy but a budgeting instrument. It answers the question of what a review hour is spent on, and it counters the common pattern where all texts are treated alike and therefore all reviewed equally superficially. Anyone who also keeps an eye on the cost of their AI pipelines quickly sees that the machine is rarely the expensive part. The effort sits in the review - which is exactly why it needs to be steered.

Content classExampleMachine stageReview stage
Product data and attributesTitles, attributes, filter valuesBatch run with glossarySample per batch
Category and help textsCategory text, shipping, returnsTranslation with page contextPost-editing by an editor
Brand and legal copyHome page, campaign, withdrawalDraft for submission onlyFull expert review
User-generated contentReviews, customer questionsMachine translated and labelledReporting route instead of sign-off

The pipeline: four stages between source text and language version

A robust translation pipeline consists of four stages that are run in a fixed order - plus a return path for everything that fails. What matters is that each stage produces a clear result and that no stage is skipped because things are urgent. This is where language projects fail more often than they fail on technology.

  1. Check the source text: the source is cleaned up before translation. Inconsistent spellings, typing errors and ambiguous wording otherwise multiply across every target language. A clean source text is the cheapest quality gain in the entire pipeline.
  2. Prepare: glossary, style rules and the do-not-translate list are loaded, placeholders and markup are masked, context information is added - for example whether a short segment is a button, a heading or an attribute value.
  3. Machine stage: translation runs as a batch and results are cached so that unchanged segments do not recur in the next run. Segments where a glossary term is missing or where the length falls outside the allowed range are flagged automatically.
  4. Review stage: depending on the content class, this is a sample, post-editing or a full review. Findings are not only corrected but assigned to an error class - that is the only way to improve the rule set instead of repeating the same manual corrections in every run.
  5. Sign-off and publication: the language version goes live only after sign-off, together with its hreflang annotation and correct language markup. Until then it stays unpublished rather than half-finished online.
  6. Returns and rule maintenance: every recurring error travels back into preparation as a glossary entry or a style rule. A pipeline without this return path repeats the same errors in every run.
Machine text published without review is a search risk

Google Search documentation treats content that is machine generated or machine translated at scale and published without human review as a breach of the spam policies (Google Search Central). The difference lies not in the technology but in the process: machine translation with documented review and sign-off is expressly fine, an unchecked bulk import is not. Anyone who already has to protect rankings through a relaunch should be all the more careful not to push out unchecked language versions.

Terminology, style and placeholders: preparation decides the outcome

Most translation errors that become visible in the shop are not created inside the machine but before it. Without binding terms, every segment decides anew what a component is called; without style rules the form of address changes in the middle of the range; without protected placeholders a variable name ends up in running text. Four building blocks lift preparation to a solid level:

A glossary that binds

A maintained term list per target language defines what the central product, material and function terms are called. It works twice over: in translation and later in search and filter logic, which build on the same terms.

Style rules per language

Form of address, sentence length, handling of numbers and units, capitalisation in headings: what feels self-evident in the source language has to be decided per target language. Two pages of rules are typically enough.

Protect placeholders and markup

Variables, HTML fragments, units and article numbers are masked before translation and restored afterwards. This prevents the most common technical error class: broken placeholders and links that lead nowhere.

Supply context

A three-word segment is barely translatable without context. Details on field type, available character length and position - button, heading, attribute value - noticeably raise the hit rate of the machine stage.

These rules do not belong in a document on a shared drive but in the configuration of the pipeline. In the programming and further development of a shop, the rule set is versioned like code so it stays traceable which glossary version a language edition was created with. A rule set of this kind can stay compact per content class:

language-rules.yaml
# Rule set per content class: drives the engine, the review stage and sign-off
content_class: product_data
source: pim
target_locales: [en-GB, nl-NL, fr-FR]
glossary: glossary-technical-v7
style_guide: style-shop-factual
do_not_translate:
  - XICTRON
  - "{{sku}}"
  - "{{variant}}"
  - "{{unit}}"
review_level: sample # sample | post_edit | full_review
sample_percent: 5
approval: auto
on_glossary_miss: hold # an unknown technical term is held back
max_length_ratio: 1.35 # target language may run 35 per cent longer

---
# Brand and legal copy: the engine only delivers a draft
content_class: brand_and_legal
source: cms
target_locales: [en-GB, nl-NL, fr-FR]
glossary: glossary-brand-v3
do_not_translate: [XICTRON]
machine_output: draft_only
review_level: full_review
approval: expert

The value of this file lies in the fact that it can be checked. review_level and approval answer who signs off a language version. do_not_translate prevents brand names and variables from being translated along with everything else. max_length_ratio catches the classic case of a French button label bursting its frame. And on_glossary_miss makes sure an unknown technical term is not silently guessed but held back - the difference between a pipeline that learns and one that quietly produces errors.

Making quality measurable: error classes instead of gut feeling

As long as translation quality is negotiated as a feeling, the loudest opinion wins. It becomes measurable through error classes: every finding from the review is assigned to a category and weighted, and the total produces an error rate per thousand words. There is an international framework for grading the rework involved. The standard ISO 18587 describes the requirements for full post-editing of machine translation output and distinguishes it from light post-editing (ISO 18587). For translation services in general, ISO 17100 additionally requires a second person to revise the translation (ISO 17100). Neither standard supplies a tool, but both supply a language in which depth of review can be agreed.

Error classExample in the shopEffect
TerminologyTwo names for the same componentSearch and filters miss the target
AccuracyMeasurement or negation flippedReturns and complaints
Language correctnessWrong article, clumsy word orderLoss of trust on the product page
Style and tonePromotional tone in legal copy, shifting addressThe brand image blurs
Local conventionDate, decimal separator or unit wrongMisunderstandings during ordering
MarkupPlaceholders or HTML brokenRendering faults and dead links

Two figures are enough to start with. First, the error rate per batch, separated by content class - it shows whether the rule set is working. Second, the share of returns, meaning the segments that travel from the review stage back into preparation. If the return share falls across several runs, the glossary is doing its job. If it stays high, a term or a style rule is missing. Consistent sampling matters here: for product data, five per cent per batch is typically sufficient, provided the selection is drawn at random rather than taken from the first records - in practice those are often the cleanest ones.

Know the threshold and use it

Not every improvement pays off. For search in multilingual shops, translation quality correlates with search quality only up to a threshold; above that, additional effort adds little for search (Amazon Science). In practice that means: product data has to be terminologically correct so that search and filters work. Linguistic fine-tuning belongs instead in brand and campaign copy, where it acts directly on the purchase decision.

Shop technology: hreflang, language markup and updates

A good translation is of little use if the language version is not found or is served incorrectly. The technical foundation is manageable but error-prone, because several things have to be right at the same time - in the markup, in the URL structure and in the update run.

  • Set hreflang annotations in pairs: every language version points to all the others, and every reference is returned. If the return link is missing, the search engine does not evaluate the annotation (Google Search Central). An x-default catches visitors for whom no matching version exists.
  • Mark up the language: the page lang attribute and the markup of foreign-language passages are requirements of the accessibility guidelines (W3C). Screen readers use them to pick pronunciation - a wrongly marked-up page sounds like gibberish to blind users.
  • Decide the URL structure early: language directories, subdomains or separate country domains are a fundamental decision. Changing it later brings a full redirect plan with it, mapping every old address to its new counterpart.
  • Update runs instead of full imports: only changed segments go through the pipeline again. A comparison based on checksums saves cost and prevents texts that have already been signed off from being overwritten by a fresh run.
  • Consider mandatory information per market: some details are not translated but depend on the country - such as energy labels and product information sheets, warranty notes or information on dispute resolution.
  • Include images and alternative text: alternative text is content, not decoration. Anyone who maintains it systematically should route it through the same review pipeline as product texts.

Technology also raises the question of what happens to search. A multilingual shop needs its own word decomposition, its own synonym lists and its own stop words per language - otherwise the Dutch version finds little even though the catalogue is fully translated. That work belongs in the same planning as the search engine optimisation of the target markets, because both build on the same terms. If the glossary is kept clean, it supplies the same foundation for both tasks.

Anchor translation as an operating process

A language project does not end with the first import. Ranges grow, descriptions are reworked, legal texts change, new markets are added. Without clear ownership the language versions drift apart: the German product page names a new property, the English one does not - and nobody notices, because no run compares the two. That is why the translation pipeline belongs in day-to-day operations rather than in a one-off project budget.

Content classes and rule set

We define with you which content gets which depth of review and anchor the rule set in the pipeline under version control rather than in a file share.

Review stages with evidence

Sampling, post-editing and expert sign-off are documented - guided by the grades of rework described in ISO 18587 (ISO 18587).

Technical delivery

We set up hreflang, language markup, URL structure and the update run in your shop so the language versions are indexed cleanly (Google Search Central).

Operations and rule maintenance

Returns flow back as glossary entries and style rules so the error rate falls across runs instead of repeating itself.

The order matters: classify first, then prepare, then machine translate, then review. Turning it around and starting with the machine means reviewing texts that would not have been written that way in the first place - and correcting symptoms instead of rules. If you want to bring your shop into further languages or lift an existing language version to a solid level, we set up the pipeline: from content class through glossary and review stages to hreflang delivery. As a Shopware agency we anchor it directly in the operation of your shop - talk to us in an initial consultation about your target markets.

Sources and Studies

This article draws on data from CSA Research, Nimdzi and Amazon Science. The figures cited refer to the state of the respective publication. In addition it refers to the requirements of the standards ISO 18587 and ISO 17100, the spam policies published in the Google Search documentation, the language requirements of the W3C accessibility guidelines and information from the European Commission on the 24 official EU languages and the Geo-blocking Regulation. As at: August 2026.

For structured content such as product data, attribute lists and category pages, the machine stage usually delivers usable results, provided a glossary and protected placeholders are in place upstream. For brand copy, campaigns and legal texts it typically works only as a draft. It therefore makes sense to split content into classes with different depths of review rather than deciding for or against the machine across the board.

Everything that is legally binding or carries the brand: right of withdrawal, warranty and safety information, payment and shipping terms, home page and campaign copy. Added to that are terms that feed into search and filters, because an inconsistent technical term leads directly to false hits there. For grading the rework, ISO 18587 describes the requirements for full post-editing (ISO 18587).

The problem is not the technology but a missing process. Google Search documentation treats content that is machine translated at scale and published without human review as a breach of the spam policies (Google Search Central). A pipeline with documented review and sign-off is not affected by this. Correct hreflang annotations with reciprocal references also matter.

Light post-editing corrects errors of meaning and makes the text comprehensible, but largely leaves style and phrasing as the machine produced them. Full post-editing lifts the text to the level of a human translation, including terminology, style and audience address; its requirements are described in ISO 18587 (ISO 18587). For translation services, ISO 17100 additionally provides for revision by a second person (ISO 17100).

A few languages at good quality beat many languages at weak quality. It makes sense to start with one or two target markets where range, shipping and payment methods already work. Since 60 per cent of online shoppers rarely or not at all buy from English-only sites (CSA Research, 2014), the local language pays off where meaningful revenue is expected. An English version as a fallback language remains useful alongside it.

The largest effort sits in the first run. After that the pipeline only touches changed segments, provided a checksum comparison is in place. Budget fixed time shares for maintaining the glossary and style rules and for sample reviews of new batches. Experience shows the share of returns drops noticeably across the first few runs when findings are consistently fed back into the rule set.