Glossary
Valid from 1.9.2026

Glossary

Glossary

DeepL Glossary: Usage and Recommendations

Purpose of this document: Texts are automatically translated using the DeepL service – via a programming interface (API), i.e. from within the company’s own systems. A ‘glossary’ is a list of terms that tells DeepL how a particular term is to be translated.

Evidence levels for statements

  • [DOC] = documented by DeepL itself
  • [PRACTICE] = case studies from the translation industry
  • [STUDY] = scientifically researched
  • [OPEN] = not described anywhere; to be verified through our own tests

Definitions

The following terms are used in this document. They enable the reader to understand the rest of the text without prior technical knowledge:

  • Glossary – a list of terms provided to DeepL during translation which specifies the desired translation.
  • Entry – a line in this list: a source word and the desired translation.
  • API (Application Programming Interface) – the method through which DeepL is used automatically from one’s own systems, rather than inserting texts manually. It unlocks functions that are not available in the web interface.
  • Source language and target language – the language from which the text is translated and the language into which it is translated.
  • Ambiguous term – a word with multiple meanings, such as the English word ‘bridge’, which can mean a bridge, a footbridge or a product name.
  • Mark-up in the source text – an invisible tag surrounding a term that instructs DeepL to leave this section unchanged. Technically: ignore_tags.
  • Translation memory – a repository of complete sentences that have already been translated, which is reused. The ideal place for text blocks – unlike the glossary, which covers only individual terms.
  • Corporate terminology (term bank) – the complete, maintained directory of all technical terms. The DeepL glossary is merely a filtered extract from this.
  • Version control – a system that documents every change to a file, thereby making it traceable and revisable.

1. Key point

A DeepL glossary is not a ‘search-and-replace’ list, but rather a set of preferences for the translation system. DeepL decides anew for each sentence whether to follow these preferences. This is where most expectations fall short: the assumption that a glossary entry will always have the same effect, everywhere, is incorrect.

What this means in practice:

How DeepL behaves, as demonstrated by

DeepL adapts entries itself to the grammar of the target language. Plural and inflected forms therefore do not need to be entered separately. [DOK]

An entry also applies even if the words in the sentence are inflected or separated. The exact word order does not need to be matched. [DOK]

DeepL deliberately ignores an entry if it does not fit within the sentence. Example from DeepL itself: The entry ‘speech to text’ does not apply if these three words merely happen to appear next to one another in the sentence and do not actually belong together. [DOK]

DeepL can rearrange entire sentences to implement an entry. [DOK]

An entry also applies to compound words. Practical example: the entry ‘imager → image generator’ also produced ‘image generator module’. [PRACTICE]

First consequence: There is no guarantee that an entry will apply, nor is there any way to check this on a sentence-by-sentence basis. Where accuracy is essential, a subsequent check remains indispensable.

Second consequence: As an entry applies across word forms and compound words, an unsuitable entry can noticeably impair the quality of the translation.

2. Inclusion criteria: Which terms belong in the glossary?

Not ‘all technical terms’ – that is the wrong question to ask. The decisive factor is whether a term is unambiguous.

The decisive criterion

A term should only be included in the glossary if the desired translation applies in every single instance – without exception.

The test: Can it be confirmed without reservation that, whenever this term occurs, exactly this translation is intended? If not, it should not be included. This criterion is stricter than simply classifying a term as technical – ambiguous technical terms do not meet it. At the same time, it is broader in scope: product names and terms that are to remain untranslated are permitted.

What can be included

DeepL itself lists four groups [DOK]:

  1. The company’s own terms – corporate language, product and service names
  2. Terms used with this specific meaning only within the relevant industry
  3. Terms tailored to a specific target audience
  4. Terms that are to remain untranslated

In practice, the following also apply [PRAXIS]: technical terms with exactly one approved translation, terms agreed with customers, industry-standard abbreviations, fixed multi-word phrases, and foreign words that are intentionally to be translated.

What is excluded

What Why Supported by

Everyday words and normal colloquial language DeepL rejects obvious standard translations itself – a glossary is not a dictionary. [DOK]

Terms with multiple meanings DeepL allows only one translation per word. A term with two meanings cannot therefore be captured unambiguously – one meaning is enforced in all cases. [DOK] + [PRACTICE]

Product names that are also everyday words (see Chapter 4) The greatest risk of an entry being applied in unintended places [PRACTICE]

Very short terms and abbreviations that are also everyday words (IT, ALL, SIE) They frequently clash with general text. Whether case (upper and lower case) plays a role in this is not documented. [OPEN]

Short function words (before, and, with …) A documented case from a study: the entry ‘Before → NEXT’ transformed the beginning of the sentence ‘Before the unit is serviced …’ into a meaningless ‘NEXT, when the device is serviced …’. [STUDY]

Complete sentences These belong in a translation memory, not in the glossary. DeepL explicitly refers only to words and short phrases. [DOC]

Software controls, placeholders, code blocks These are protected by markup in the source text (see Chapter 6), not via the glossary. [DOC]

Duplicate entries, typos, old special characters such as * or – DeepL rejects duplicate entries upon upload. Old placeholder characters from previous word lists must be removed beforehand. [DOC] + [PRACTICE]

Why this criterion must be interpreted strictly

Two academic studies (MT Summit 2021 and EAMT 2023) reach the same conclusion: if a complete corporate terminology is used unfiltered as a glossary, translation quality decreases measurably – it does not merely remain unchanged, but actually deteriorates. [STUDY]

The most telling finding (English–Russian): when glossary terms were strictly enforced, the desired terms appeared in the text in 98 per cent of cases – however, the readability of the sentences fell from 9.4 to 7.5 points. With less stringent enforcement, readability was maintained, but only 76 per cent of the terms were implemented. No single approach proved to be clearly superior.

It should be noted that a high hit rate for terms is not synonymous with good translation quality. Each entry involves a trade-off between terminological benefit and a loss of readability. For general, short and ambiguous entries, this trade-off is always to the detriment of readability.

Recommended scope of the glossary

DeepL does not specify an upper limit and, in its more expensive subscription plans, advertises ‘an unlimited number of entries’ [DOK]. Industry practice is nevertheless consistent [PRAXIS]:

Compact glossaries that are limited to essential terms are more effective. Very extensive glossaries can impair translation quality.

Guideline: A realistic target is a few dozen to a few hundred entries per language pair – not the entire terminology database. By way of comparison: DeepL’s own suggestion assistant generates an average of around 20 entries from 1,000 example sentences [DOK].

The comprehensive terminology database can remain as extensive as before. The glossary for DeepL remains a filtered extract from it.

3. Importance of order

This question may relate to three different issues – each with a different answer:

a) Order of the lines in the uploaded file → no influence

A greater impact of entries appearing higher up is not documented anywhere and is highly unlikely [OPEN]. The only documented aspects are the required file structure and the fact that no header is required – the first line is already an entry [DOC].

b) Word order within a multi-part entry → no exact match required

DeepL recognises a multi-part entry even if the words in the sentence are inflected or separated [DOC]. A multi-part entry therefore does not apply only to this exact string of characters.

The reverse scenario is also documented: if the words do appear but do not belong together in terms of content, DeepL deliberately ignores the entry.

DeepL therefore compares the meaning, not the string of characters – and does so in both directions.

4. The special case: product names that are also everyday words

This is where the key limitation lies. It is not a matter of configuration – DeepL is technically unable to handle this scenario.

The problem

  • Terminologically correct would be: two terms spelled identically but with different meanings form two separate entries [PRAXIS].
  • However, DeepL only allows one translation per source term; a second entry for the same term is rejected upon upload [DOK].

An entry that leaves a product name unchanged cannot therefore be restricted to those cases where the product is actually meant. It applies without exception.

The risk depends on the source language

The difference is crucial when translating from both German and English:

Language Risk Why

German low An English product name is not a common word in German. An entry practically only matches the product name; the German word with the same meaning is a different character string and is not captured.

English high Product names in English are often also everyday words – such as ‘bridge’ in ‘suspension bridge’ (Hängebrücke), ‘bridge circuit’ (Brückenschaltung) or ‘land bridge’ (Landbrücke). An entry overrides all these uses and also continues to affect compound words.

An example of the problem, using the fictional product name ‘Bridge’ and English as the source language:

Entry Desired effect Actual problem

‘Bridge’ remains untranslated as ‘Bridge’ The product name remains unchanged. ‘Suspension bridge’ becomes ‘Suspension-Bridge’ instead of ‘Hängebrücke’ – and the error carries over into all compound words.

The same pattern occurs with all product names that are also everyday words – such as Access, Wave, Teams, Vision or Studio.

Recommendations in order of importance

1. Do not include a product name that is a common word in the source language in its basic form in the glossary – particularly not for language pairs where English is the source language.

This is the most important recommendation in this document: the loss of translation quality outweighs the terminological benefit.

2. If a glossary entry is required, then create one for each variant – and only for variants that actually exist:

Bridge L      → Bridge L
Bridge M      → Bridge M
Bridge Pro    → Bridge Pro
(no entry for ‘Bridge’ on its own)

Multi-part entries are permitted and, in this case, the preferred approach:

  • DeepL explicitly states this: an entry may contain several words or a short phrase [DOK].
  • The more specific the entry, the less likely it is to trigger in unintended contexts.
  • No word count limit is specified anywhere – only a length limit of 1,024 characters per source and target term [DOK].

Important: You cannot rely on an entry for the base name automatically covering its variants. Whether DeepL recognises a variant as a single unit or merely replaces the base name contained within it is not documented anywhere [OPEN] – there is no documented rule stating that the longer or more precise entry takes precedence. Anyone relying on this may receive a different result with every DeepL update.

Furthermore, product variants should not be entered in large numbers as a precaution: the quality of the results decreases as the number of entries increases [PRACTICE].

3. Editorial rule: always ensure the product name is recognisable in the text

For example, you should write ‘the Bridge module’, ‘the Bridge L assembly’ or ‘Bridge™’. This is the most cost-effective long-term solution: it supports both the translation system and human readers, and requires only a single editorial rule.

4. Prevention: check new product names in advance

A name that is a common everyday word in the source language will cause ongoing translation problems. This check must be carried out before the product is launched, not afterwards.

5. Ongoing maintenance of the glossary

Responsibilities

The following are required: a named individual with decision-making authority for terminology; a reviewer for each target language; and a point of contact from product management or marketing for product names. Unclear responsibilities are the most common reason why glossaries are not maintained.

Approval process for new terms

propose → review → confirm → approve → add translations in the target languages

Terms that are no longer valid are marked as ‘obsolete’ and not deleted – this ensures that the development history remains traceable. Established terminology tools provide fields designed for this purpose (such as ‘preferred / acceptable / obsolete’ as well as details on part of speech and the type of term). Specifying the part of speech is, in fact, essential for automated processing.

Publication frequency

Proposals are collected on an ongoing basis, but are only published on fixed dates (fortnightly or monthly). This allows changes to be tested in batches. A full review is recommended on a quarterly basis: which entries are appearing in unintended places, and which product names have been discontinued?

No short-term corrections should be made to a glossary in production without prior cross-checking. A single unsuitable entry directly compromises the overall result.

The crucial test: whether an entry is beneficial or harmful

An entry is only deployed in production if it demonstrably improves the test cases.

Checklist before each upload

  • no duplicate source terms (otherwise the file will be rejected)
  • no typos
  • no old placeholder characters such as * or -
  • no lower-case function words
  • no entry longer than 1,024 characters
  • Use basic forms: nouns in the singular (‘package’, not ‘packages’), verbs in the infinitive (‘to purchase’) [DOK]
  • Capitalisation according to the rules of the respective language [DOK]
  • No control characters, no spaces at the start or end, neither the source nor the target field may be empty [DOK]

Rules for the source text (maximum impact, minimum cost)

  • Formulate short, direct sentences; break up long, convoluted sentences
  • Avoid ambiguity: instead of “Press the right button”, use “Press the button on the right-hand side”
  • Use one term per concept – do not switch between “customer”, “client” and “user”
  • Write out abbreviations in full the first time they are mentioned
  • Avoid idioms
  • Correct errors in the source text first

Any ambiguity that is resolved in the source text does not need to be addressed later with a glossary entry – and cannot cause subsequent errors.

6. Overview of technical limits

Topic Value Specified by

Duplicate entries are rejected upon upload [DOK]

Languages All DeepL languages except Thai [DOK]

Language variants The target language ‘English’ applies to both British and American English; the same applies to Portuguese and Chinese [DOK]

Source language must be specified; glossaries do not work if DeepL is to recognise the language automatically [DOC]

7. Decision matrix

If the term … … then

is a technical term for which the same translation always applies: glossary entry, singular noun

is a verb or adjective for which the same translation always applies: glossary entry in the base form; permitted by DeepL – unlike in systems with rigid substitution

is a product name that is not a common word in the source language; glossary entry that leaves the term unchanged – or markup in the source text

is general, short, ambiguous or a small function word; do not include; instead, formulate the source text more precisely

is an entire sentence – Does not belong in the glossary

is a preference regarding tone or style – Does not belong in the glossary

8. Next steps

  1. Check the glossary: check the current glossary entry by entry against the criteria in Chapter 2 and mark everything that does not apply in every case.
  2. Identify product names: Which product names are also common English words? Remove these from the glossaries with English source text.
  3. Separate glossaries by source language (German separate from English) – the risks differ.
  4. Check in all translation requests whether the source language is specified – without this information, the glossary will not work. This point does not apply if the requesting system automatically includes the source language.
  5. Set up central corporate terminology, including automated export; integrate the export into version control and set fixed publication dates.

Sources

DeepL — Support

DeepL — Developer documentation

Studies

Terminology standards and industry practice