How to Manage Terminology in Large-Volume Translation Without a Glossary

Large-volume translation projects often begin before an organization has a usable glossary. In that situation, terminology consistency depends on whether the provider can identify key terms, apply approved terminology across the full content set, and verify it before delivery.

A large translation request can arrive with hundreds of pages or dozens of files but no glossary. The organization may never have managed terminology formally, yet the translation still has to use the same terms across every document.

Technical document translation, contract sets, internal administrative and legal materials, and company-wide training content all create the same problem at scale. A translation glossary can determine whether the final delivery reads like one controlled body of content or a collection of separately translated files. Yet many high-volume translation projects begin without one.

Price and delivery schedules are easy to compare. Translation quality is harder to judge before delivery. In large document sets, terminology inconsistency can be as damaging as mistranslation. If the same concept is translated one way in one file and another way elsewhere, readers may assume the terms refer to different things or question the reliability of the content. In contracts, that can affect interpretation. In training material, it can confuse learners.

This article looks at why terminology drifts in large-volume translation and what a translation provider needs in place to control it, using Hansem Global’s terminology workflow as an example.

Why Large Document Sets Create Terminology Drift

Terminology inconsistency can look like something translators should simply avoid. At scale, however, the production structure itself creates variation.

The first reason is parallel work. Large volumes and fixed deadlines often require multiple translators. Without an agreed termbase, each translator may choose a different but individually reasonable equivalent. The variation appears where their work meets.

The second reason is variation in the source. Document sets written by different departments at different times often use several expressions for the same concept. A simple function such as account registration may appear as register, sign up, enroll, or apply. In ordinary context, those differences may not confuse readers. In specialized documentation, however, terms such as inner receptacles, inner packagings, incompatible, or inspection body can carry specific domain meanings. Inconsistent handling can become a serious quality issue and, in safety-related content, can potentially affect user safety.

Even when a translation buyer understands the value of a glossary, deciding which terms to control across hundreds of pages and how to set the standard is usually outside the day-to-day scope of procurement or project coordinators. The provider therefore needs a process for establishing terminology before full production starts.

Terminology inconsistency is difficult to prevent through individual attention alone. It is a process problem. Based on the terminology management workflow Hansem Global operates with its in-house AI tools, three conditions are worth checking.

Condition 1. The Provider Should Be Able to Build Terminology Without a Glossary

Organizations commissioning large-volume translation often cannot supply a translation glossary. Sometimes a glossary exists but covers only a small portion of the content. In other cases, the lack of time or internal resources to create one is precisely why an external provider is being engaged. The first requirement, then, is the ability to identify and establish key terms from the source documents themselves.

Hansem Global uses its in-house AI platform, Hansem CALM (Computer-Assisted Linguistic Model), to extract terminology candidates from source content. Traditional automatic term extraction has often relied heavily on frequency. That can surface large numbers of common words while missing important terms that appear only once. Hansem CALM first identifies the subject of the document and evaluates candidate terms in domain context. A word may have very different importance in industrial equipment documentation than in financial terms and conditions, and that context is considered during extraction.

Extraction depth is adjusted to the project. A one-time large-volume project may benefit from quickly establishing a focused set of core terms rather than generating an exhaustive list that takes longer to review.

Subject-matter linguists then review the candidates, translate them where required, and submit them for client confirmation. AI proposes the candidates; linguists and the client make the final terminology decisions.

Extraction scope / Desired term count

For a buyer, this capability should be visible early. If a reviewable terminology list arrives near the start of the project even though no glossary was supplied, the provider has a mechanism for establishing terminology. If production begins with no terminology questions at all, key decisions may be left to individual translators.

Condition 2. Approved Terminology Must Be Applied Across the Full Translation

Approving a terminology list is one task. Making sure it is used consistently across hundreds of thousands of words is another.

Some machine translation workflows apply glossary terms by replacing words after the translation has been generated. That may work in limited cases, but post-generation substitution can disrupt grammar in languages where particles, endings, or inflection interact closely with the surrounding sentence. The correct term may appear, yet the sentence may still need human repair.

Hansem CALM incorporates the approved glossary when the AI translation is generated. The model can use the approved term within the sentence structure rather than inserting it afterward. In the Korean example below, the term base maps release to the approved Korean term “배포.” When that terminology is applied, release key can be rendered as “배포 키” rather than the transliterated “릴리스 키.” The same glossary can be used across multiple files, so consistency does not depend on a translator remembering every approved choice.

Approved glossary loaded in Hansem CALM AI translation

Condition 3. Terminology Deviations Need to Be Checked Before Delivery

Translation buyers cannot review every segment in a high-volume project. Pre-delivery QA therefore has to provide coverage that individual spot checks cannot. Traditional string-matching terminology checks, however, have blind spots.

Suppose the glossary contains log in, but the source sentence uses logged in. A strict exact-match check can fail to treat the inflected form as the same term. Expanding the search pattern can reduce misses, but in traditional rule-only workflows it can also increase false positives. Hansem CALM QA uses an AI judgment layer to distinguish valid morphological or spelling variation while keeping the broader terminology check in scope. The objective is to reduce both missed issues and unnecessary alerts.

The point is not to replace rule-based QA with AI. Rules remain useful where speed and reproducibility are clear advantages; AI judgment is used for variation and context. The two layers are combined so that each can cover failure points the other handles less well.

Approved glossary entry
Translation under review with logged in as an inflected source form
Hansem CALM AI QA result including the inflected term variant

Items found by QA are organized in a standard report for linguist review. Cases that require no subjective judgment, such as a straightforward failure to apply an approved term, can be corrected directly in the tool. Separating these paths reduces unnecessary back-and-forth before delivery.

What Should Remain After the Translation Is Complete

Finishing one large-volume translation project cleanly is not enough. The terminology decisions and standards established during the project should remain as reusable assets for the next round of work.

If the organization receives and owns the approved glossary along with the translated content, future projects can start from the same terminology baseline whether they are handled by another internal team or a different vendor. As similar work accumulates, these assets reduce repeated review and help preserve consistency across departments, document sets, and project cycles.

When evaluating a provider for large-volume translation, price and schedule are only part of the assessment. Ask how terminology will be identified, approved, applied during translation, verified before delivery, and returned as a reusable asset. Consistency across a large body of content is not sustained by individual memory; it depends on a terminology management process that spans the project.

If your organization is preparing a large-volume translation project without an established glossary, Hansem Global can help build the terminology baseline, apply it throughout translation, and verify it before delivery.