From the translation memory revolution
Translation memories have been around for more than 40 years. Originally, they were simple databases that could be queried using SQL (Structured Query Language). SQL is highly intuitive: to find the sentence “Stop the engine” in a table named table_segments, the query would simply be: SELECT target FROM table_segments WHERE source = 'Stop the engine'. If we wanted the table to return that sentence when it appears within a longer sentence, the query would be: SELECT target FROM table_segments WHERE source ILIKE '%Stop the engine%'. In this case, the query will display the sentence even if text appears before and after it, as in: “Before carrying out any maintenance, stop the engine and disconnect its power supply.”.
But you will notice that this type of memory searches for a textual formulation. If “Stop the engine” is replaced with “Shut down the engine”, the system will find nothing, even though the meaning is the same. This is clearly a major structural limitation. Overcoming it required the advent of artificial intelligence—or, more precisely, AI tools, the most important of which is vectorisation.
To the vector revolution
Vectors are fundamental to artificial intelligence because, despite what its name might suggest, AI does not understand text in the same way as a human being: it converts text into numerical representations known as vectors. Each vector consists of a set of numerical values representing characteristics of the text. You may recall the title of the famous paper published in 2017 by Ashish Vaswani and several Google researchers: “Attention Is All You Need.” The attention mechanism allows a model to assign different levels of importance to relationships between tokens according to their context. These vector representations and attention mechanisms form the common foundation of both large language models (LLMs) and the new generation of translation memories.
In these memories, source and target segments are stored in text columns. The text remains essential for retrieving and checking translations, but another column becomes predominant for search purposes: the vector column. Remember that a database table is made up of rows and columns.
The vector stored in this column is generated by an embedding model based on an architecture known as a Transformer. The source segment is first broken down by a tokenizer into tokens, which may correspond to words, parts of words or punctuation marks. Starting from the initial vector associated with each token, the model then transforms these tokens into 1,024-dimensional vectors. This is the case with the BGE-M3 model, although the number of dimensions varies between models.
The vectors are refined through the Transformer’s 24 layers, which modify them according to each token’s initial vector, its position within the segment and its relationships with the other tokens in that segment. After these 24 layers, the token vectors have been contextualised and are then aggregated to produce a single 1,024-dimensional vector representing the segment as a whole. The individual values in this final vector no longer have a directly interpretable meaning: it is their combination that represents the segment’s semantic content.
With vectors, mathematics takes over
Once we are comfortable with our precious vectors, everything becomes straightforward. The user sends a query to our system: “Find the translation of the sentence ‘Before each maintenance operation, stop the engine.’”
Our tokenisation and transformation model breaks the search sentence down into tokens and then converts it into a vector. It is this vector—not the text itself—that is sent to the vector database to perform the search.
The database then carries out a comparison operation that may remind you of your secondary-school mathematics lessons: vector calculations. Among the stored vectors, it looks for the one with the highest cosine similarity to the vector representing the search sentence.
The similarity between the search vector (A) and a stored vector (B) is calculated using the following formula:
Similarity (A, B) = (A · B) / (‖A‖ × ‖B‖)
In other words, similarity is equal to the dot product of vectors A and B divided by the product of their magnitudes. The closer this value is to 1, the greater the similarity. If the two vectors point in the same direction, the angle between them will be close to 0 and its cosine will consequently be close to 1.
Enter the realm of “thinking” memories
We have seen that the final vector assigned to each segment by the tokenizer–Transformer model takes account of the initial value of each token, its position within the segment and the values of the surrounding tokens. For any competent tokenizer–Transformer model, the segments “Stop the engine before each maintenance operation” and “Before carrying out any maintenance, you must shut down the engine” mean the same thing, and their final vectors are therefore likely to have a high degree of similarity. Whether you request the translation of the first or the second segment, there is consequently a strong chance that the same proposed translation will be returned.
This is known as vector semantic search.
The vector pipeline and machine translation model
The market has seen the emergence of effective machine translation models such as DeepL, ModernMT and others. DeepL next-gen, currently one of the most advanced of these models, can use results retrieved from translation memories as contextual information to guide its translation. It can also connect to users’ translation memories and glossaries. Existing translations can then be reused directly when a match is found; for the remaining segments, its engine generates a new translation inspired by the terminology and style provided to it.
Schematically, the following type of instruction could be submitted to a model such as DeepL next-gen: “Here are our translation memory, our glossaries and the document to be translated. Please translate the document using the terminology provided and the translations stored in our memory as guidance.” Perfect matches can be retrieved unchanged. The rest of the document is translated with reference to the glossary and the linguistic resources provided.
This is certainly an advance over earlier models, but it comes at a price: dependence on DeepL’s infrastructure, a lack of control over how the engine operates and how the memories are updated, as well as sovereignty and confidentiality concerns associated with relying on an external provider.
The adaptive semantic memory solution
The gap left by a model such as DeepL next-gen, however advanced it may be, can be filled by adaptive semantic memory technology. Vendors offering adaptive technologies include memoQ AGT and Smartling.
One of the principles behind this technology is the adaptation of fuzzy matches. Here, the prompt can be more precise: “Here are the source and target segments from an earlier translation, together with a new segment that we have found to be highly similar to the previously translated source segment. Please adapt the supplied translation to the new source segment without re-translating it in full.”
A conventional machine translation engine is not necessarily designed to receive an existing source–target pair and adapt its translation in response to detailed instructions. An LLM such as GPT, Claude or Gemini is particularly well suited to this task because it can compare all three segments, identify their differences and modify only the elements concerned..
Suppose you submit the prompt above to an LLM together with the following three pieces of information:
Previous source segment: “Change the engine oil every 300 operating hours or once a year, whichever comes first.”
Previous translated segment: “Effectuez la vidange de l’huile moteur toutes les 300 heures de fonctionnement ou une fois par an, selon la première échéance atteinte.”
New segment to be translated: “The engine oil must be changed every 300 operating hours or every six months, whichever comes first.”
The LLM would then be instructed to retain the sentence structure and terminology as far as possible and simply adapt the translation to reflect the meaning of the new segment. It could therefore propose: “Effectuez la vidange de l’huile moteur toutes les 300 heures de fonctionnement ou tous les six mois, selon la première échéance atteinte.”
And so the circle is complete: we have moved from textual search to vector semantic search, and then to semantic adaptation.
What is striking at the end of this journey is that a technology long reserved for major software vendors is now becoming widely accessible. Until recently, designing a genuinely semantic and adaptive memory required considerable research resources. Today, vectorisation, embedding models and LLMs are accessible to everyone and are often available as open-source technologies.
he different players in the sector will no longer compete over access to these models, but over how that access is provided. The major players—primarily software vendors—will continue to encourage users’ dependence on their proprietary access software and infrastructure. That is only natural. But users—particularly language service providers—will now be able to consider autonomous solutions thanks to technological progress and the wealth of open-source resources.