Skip navigation

Entry

home page help (reference guide) search download (pdf) print broken link

 

 

citation DAN Oversættelseshukommelser  SPA Memorias de traducción

 

contents
Introduction | History | The uptake of translation memory systems | Translation memory system components and how they work | Translation management system workflow | Research on translators' interaction with a translation memory system | Research potential

 

Introduction  Introduction 

Since the 1950s, scholars have been developing machine translation (MT) tools that aim to establish a semantic relationship between (at least) one source text in one language and (at least) one target text in another language on a segment or text level (Alcina 2008). The aim was to develop machines that could provide fully automatic high-quality translations. In the 1970s, after the Automatic Language Processing Advisory Committee report from 1966 had concluded that machines could not produce translations of acceptable quality, scholars, translation agencies, corporate localization departments, and individual translators started to explore translation memories (TM) as one of the pillars of so-called computer-assisted or computer-aided translation (CAT) tools. CAT tools constitute suites of tools that assist translators during the translation workflow by enabling the reuse of past translations (Garcia 2015: 69). The basic idea of a TM is to recycle translations produced by humans by storing them in a TM database as segmented and paired source and target texts to be retrieved by translators translating identical or similar segments. The first commercial CAT tools were marketed as TM systems or TM tools in the 1990s. Today, TMs are the translation tools used most frequently and widely by professional translators (n.n. 2022: 21). Presumably, the tools are referred to as a memory because the main component is a translation archive that, from a cognitive perspective, can be said to constitute a digitally extended translator memory. By storing source texts and their translations on segment level (generally one sentence long) and by recycling previously translated segments whenever the same or similar sentences occur, the translator does not have to spend time trying to remember how to translate a given word, phrase, or sentence that was translated in the past. Consequently, by storing and automatically retrieving earlier translation decisions, a TM compensates for the limitations of the human mind and helps to achieve consistency within a given translation and across translation tasks.

No other technology has revolutionized the translation industry as radically as TM technology, which for more than three decades has kept changing translation workflows and the working conditions of professional translators (Reinke 2013; Mitkov 2022). Basically, a TM system helps translators retrieve and reuse previous translations by automatically proposing a translation from the TM database as a complete or partial solution, when an identical or similar sentence is to be translated again. Nowadays, a TM system is a multifunctional tool with a range of embedded features supporting not only translation but also project management.

For many years, using TM technology has been considered an instance of computer-aided translation because it assists translators during translation, meaning that translators are in control of the translation process and can decide how and when to use the tool. However, when a machine translation (MT) feature is activated, TM systems turn into human-aided machine translation. In either case, the translator’s job is mainly to reject, accept or edit the resulting target text. Hence, in the translation industry, the bulk of the translation process is being transferred from translators to computers. Though opinions may differ about whether the translator or the machine is at the centre of the translation process when a TM system is used, scholars probably agree that various features included in the system can be said to automate the translation process to such an extent that translators using it become de-facto post-editors. In particular, neural machine translation (NMT), i.e. MT using artificial intelligence (AI), has decreased the amount of manual work required on translation platforms such as a TM system.

The taxonomy for levels of translation automation developed by Christensen, Bundgaard, Schjoldager et al. (2022) distinguishes between six levels, based on the role of the translator and/or the translation tool (illustrated in a simple form by Figure 1, below). The level of automation reflects the translators’ degree of control over the translation process. A conventional TM system (that is, TM without MT) is an example of level-1 automation (Translator Assistance), because the tool analyses the source text and retrieves TM proposals, whereas the translator evaluates these and produces the TT segment, by either accepting or editing the proposed match or translating the segment from scratch. If an MT component is activated during translation, this constitutes level-2 automation, because the system analyses the source text and comes up with a translation solution. Hence, at this level, the translator evaluates and corrects errors and inadequacies and solves any system failures. See Christensen, Bundgaard, Schjoldager et al. (2022: 31-33) for a detailed description of the taxonomy and the remaining levels of automation.

Levels of translation automation
Figure 1: Taxonomy for levels of translation automation

The overriding argument for embracing TM technology appears to be the benefits that a TM system can bring in terms of streamlining processes and projects, productivity, improved translation quality (in particular, terminological consistency), and cost savings. Worth noting is that TM databases have become an economically valuable knowledge resource for both language service providers (LSPs), translation agencies and clients, as the TM content can be recycled again and again. This might explain why LSPs and translation agencies, to a certain extent, tend to prescribe if, when and in which ways freelance translators, for instance, must use a given TM system. In most cases, translators are obliged to use a specific TM system, whereas the use of the MT feature during translation might be optional. Arguably, this indicates that translators’ right to decide for themselves during the translation process is compromised in the language industry of today. From an industrial point of view, TM systems help organize projects and streamline project workflows, allow large translation teams to work on the same project and make the sharing of work easier, thereby automating the translation process.

Besides automating the translation process, TM technology impacts on translators’ mental processes during translation in a way that makes the cognitive translation process different from typical human translation: texts are broken into segments by the TM system and processed one by one in a way that human translators do not. The segment-by-segment approach may also impose the sequence of source-language segments on the target language, resulting in structural and maybe even syntactical interference. Furthermore, when using a TM system that integrates MT, translators are offered translation proposals before they themselves have had an opportunity to consider how to translate the source text segments. This means that translators no longer handle translation problems from scratch. As post-editors, translators mainly choose between a range of translation suggestions and, if necessary, adapt them, which changes target text production dramatically. The main challenge seems to be that texts are not processed by the system as a coherent and cohesive entity, because the TM system draws on different sources.

Digital translation process
Figure 2: Model of translation competence in a digitalized translation process (Krüger 2018)

Some researchers have argued that translators using a TM risk producing a terminologically and stylistically inconsistent hybrid text (e.g. Moorkens 2011), a sentence salad (e.g. Bundgaard 2017), and that they therefore need a different set of translator competences. Krüger’s (2018: 120) model of translation competence in a digitalized translation process illustrates how technologies such as TM and MT induce shifting translation competences. Krüger’s argument is that outputs from TM and MT can relatively easily be adapted to the target text, whereas the adaptation of other types of content requires a higher cognitive effort. Therefore, he categorises TMs and MT as “digital translation data with small distance” between the source text and the target text, whereas content from sources such as termbases and web pages is categorised as “translation data with greater distance” to the target text (see Figure 2, below). Consequently, according to Krüger, the importance of text production competences is decreased in a digitalized translation process, whereas competences in text reception (including both source text reading and evaluation and the reading and evaluation of target language translation data, e.g. translation proposals), text selection, text adaptation and text optimization become more important.

This entry first outlines the historical development of TM systems and provides an overview of its uptake by translators and LSPs. Next, the basic idea and components of TM systems are explained. Then, the translation workflow with a TM system integrated into a translation management system is described. Finally, the entry reviews examples of recent research on how translators interact with a TM system that integrates MT and points to some promising research avenues.

back to top

[heading title]  History

The original idea for TM is generally attributed to Peter J. Arthern (1979), a translator for the European Commission. In a paper on the potential use of computer-based terminology systems in the European Commission, Arthern pointed out that translators wasted a lot of time retranslating (parts of) texts that had already been translated. He suggested to build a computerised storage of source and target texts that could easily be accessed by translators for reuse in current translations, allowing not only the retrieval of identical, but also similar source language segments and their translations. Arthern (1979: 94) referred to this mode of translation as “translation by text-retrieval”. Interestingly, Arthern envisioned that, at some point, it might be possible to integrate MT into a type of TM system. Martin Kay (1980), a computational linguist, is also regarded as one of the fathers of TM translation. In his paper “The proper place of men and machines in language translation”, he theorized about an electronic machine-aided tool called The Translator’s Amanuensis. His vision was to develop computerised tools that could assist human translators in their work by establishing a network of terminals linked to a mainframe computer. In other words, Kay’s basic idea was to add specific translation tools to existing text-processing technology. These tools provided various means for translators to keep track of earlier decisions and to access previous translations of the same or similar source-text passages (Hutchins 1998: 295ff). Shortly after, Melby (1981) envisioned a tool that he called The translator’s workstation, integrating various types of machine aids. His idea was to develop computer-generated bilingual concordances as a tool for translators to identify text segments with potential translation equivalents in relevant contexts. For further details about the history of TM technology, see Reinke (2013) and Mitkov (2022).

The realisation of commercial TM systems was first made possible when tools for text alignment made bilingual databases of translations possible. Four commercial TM systems appeared on the market in the early 1990s: The TranslationManager from IBM, the Transit system from Star, the Eurolang Optimizer and the Translator’s Workbench from Trados (Hutchins 1998: 303). The systems allowed companies and institutions to localize products and services into multiple languages applying consistent terminology and reusing translated content, thereby speeding up translation processes and minimising translation costs. Therefore, since the appearance of the first TM systems on the market, the dissemination of this technology has kept on growing.

In the early 1990s, TMs were translation tools stored on individual translators’ desktop computers by means of a TM database with an editor. These tools were aimed at supporting translators by taking over repetitive translation work. In the late 1990s, TM systems developed into server-based translation tools, which allowed content sharing between multiple translators. In the first decade of the 2000s, project management and collaboration features were added, resulting in the first translation management systems, allowing for highly automated workflows. Translation management systems are abbreviated to TMS, which is confusing because TM systems are also sometimes referred to as TMS. From around 2010, the systems were cloud-based and could be accessed by all users (e.g. translators) via a web browser. The first TM systems focused on language-oriented functionalities, but today they also encompass a range of business, resource planning, financial, and customer management functions (see below).

back to top

[heading title] The uptake of translation memory systems

A number of surveys have investigated the adoption of TM technology. In a 2003 study of professional translators in Germany and the UK, Wheatley (2003) found that 29% used TM daily. Based on a study conducted in 2003/2004 with responses from 391 freelance translators in the UK, Fulford and Granell-Zafra (2005) found a similar uptake with 28% reporting using TM. A few years later, based on a survey of 699 translation professionals from a wide range of countries, Lagoudaki (2006) found that as much as 82.5% reported using a TM system, indicating that TM adoption took off in the middle of the 2000s. Based on a survey conducted in May 2013 with Danish LSPs, Christensen & Schjoldager (2016) reported that 22 out of 25 LSPs used TM technology, equalling 88%. A survey with more than 7,000 translators and interpreters conducted for CSA Research found that 66% of translators use TM on most projects, making TM the most widely used technology, followed by quality checkers and terminology management (Pielmeier & O’Mara 2020). Interestingly, Georgakopoulou (2022) found that TMs are not a common feature in subtitling, since only 17% of subtitlers indicated using TM. The most recent European Language Industry Survey finds that TM is the most widely used technology among 745 independent language professionals: 53% use TM daily, 21% regularly, and 14% occasionally. Only 12% of language professionals report that they never use TM. This leads the report to conclude that CAT technology is slowly but surely implemented, also where the technology has typically been considered less useful. With an implementation rate of more than 85%, TM is also the most widely implemented technology by language companies (n.n. 2022: 23). Consequently, while MT has become omnipresent, it is clear that TM technology is still the preferred tool in the profession. The report also mentions that the top three list of TM systems used by language companies are Trados, MemoQ and Memsource, now Phrase, in decreasing order (n.n. 2022: 25).

back to top

[heading title] Translation system memory components and how they work

Although commercial TM systems vary somewhat in interfaces and internal processes, they share the basic function of deploying existing translation resources in a new translation project. In this section, we describe traditional TM system components and the main features integrated into today’s systems. According to Esselink (2020: 119), these are the translation editor, the TM database, MT, terminology management and translation quality assurance (QA). We use the CAT tool Phrase as our case in point. This is the unified brand of Memsource and the localization company Phrase, which merged in 2022. As stated by Mitkov (2022: 370), the web-based CAT tool Memsource (now Phrase TMS) is rapidly gaining in popularity due to the relatively simple interface that incorporates all main features. We will refer to this tool as Phrase TMS. Phrase TMS is subject to payment; however, Phrase and several other TM providers offer academics free access to their systems. Other TM systems such as OmegaT can be used for free by everyone. OmegaT offers CAT features included in standard TM systems but does not resemble a translation management system (see below). For a detailed review of components of TM systems, see Garcia (2015) and Reinke (2013).

The translation memory database

The TM database is the main component of a TM system. It can be said to provide the key functionality of a translation tool, because it helps establish the semantic relationship between a source and a target text on a segment level by offering translation proposals based on matches. In most conventional TM systems, texts are divided into segments, applying so-called segmentation rules, based mainly on punctuation and formatting marks: for instance, full stops (periods), question marks, exclamation marks, and (optionally) colons and semi-colons as well as end-of-paragraph markers. In most TM systems, segments equate to sentences (and sometimes entire paragraphs) and sentence-like units such as headings and lists. A pair of source and target text segments stored in the database is referred to as a translation unit (Melby & Wright 2015: 662). Typically, they are stored together with a set of metadata, such as subject field, client, date of creation and name of the translator. The metadata can be valuable for translators wishing to validate the usefulness of TM proposals during translation. Metadata can also be used to filter smaller TMs for specific translation jobs. Applying a so-called fuzzy-logic algorithm, the TM system can retrieve a matching translation, not only for identical, but also for similar segments in a new text. Today, many segment-based TMs integrate solutions that can construct the original context of translation units stored in the database. Other systems apply aligned bitexts, meaning that source and target texts are not segmented into smaller units, but stored in their entirety (see Mitkov 2022: 367).

When a new text is to be translated, the TM system compares the character chains of each new source text segment with character chains in segments stored in the TM database. This is referred to as matching the source segment against the database. The use of character chains for match retrieval is considered a shortcoming of TM systems, because the identification is an entirely form-based process that does not consider semantic, pragmatic, or contextual aspects. Thus, a TM system does not understand the meaning of texts, and it does not perform any kind of language processing, such as morphological analyses. This means, for instance, that it does not allow inflectional variants of a word to match. The fuzzy-matching algorithm of most commercial TM systems is simply calculated on the edit distance (also referred to as the Levenhstein distance) between character strings in a source text segment in the database and a new segment to be translated, that is, the difference in the number of letters (Mitkov 2022: 371). Recent research has tested methods that can perform semantic matching and retrieve matches for sentences with the same meaning, even though the meaning is represented in different syntactic structures (e.g. paraphrases) or expressed through synonyms. However, so far, no commercial TM system offers a matching method that can retrieve matches that are semantically equivalent, but syntactically different from the segment to be translated. See Mitkov (2022: 371-376), for an overview of recent research that aims at developing intelligent TM systems by exploring different semantic matching approaches that can help increase the usefulness of TM proposals and translators’ productivity.

Based on similarity between segments stored in the database and a new segment that is to be translated, a segment-based TM system provides translators with translation suggestions retrieved from a database (matches). In the editor (see below), the TM system will highlight differences between source segments in the database and the segment to be translated. The default threshold of many TM systems is 70%, but this may be changed. The key is to define a threshold where retrieved matches are useful instead of wasting translators’ time. The usefulness of a TM database for a particular translation assignment obviously depends on the number of segments in the database.

A TM database can be created by translating texts from scratch, meaning that the number of paired segments in the database grows continuously. Another possibility is to align previous translations, applying an alignment feature that links corresponding source and target text files on segment level and import them into a TM database. It is also possible to import (and export) other databases created with the same TM system or databases from other TM systems available in the Translation Memory eXchange format (TMX), an open, standard text-based format representing structured information that is supported by all commercial systems (Garcia 2015: 71). However, the usefulness of the TM also depends on the thematic (and semantic) interrelatedness of the content of the database and the text to be translated. Hence, it does not necessarily make sense to create all-in-one TMs (also referred to as Big mama TMs), such as TMs covering different subject fields, which may generate irrelevant match retrieval. To obtain internal and external consistency, TMs tend to be kept segregated for, say, a particular topic or client.

The surface matching algorithm of most commercial TM systems has operated with four match types for many years: exact matches, context matches, fuzzy matches, and no matches. An exact match (100% match) is a segment that is completely identical to an existing TM segment. A context match (sometimes referred to as a 101% match) is an exact match where the segment in question is preceded and/or followed by exactly the same match, i.e. occurs in the same context. This is achieved by storing the relevant context segments together with the actual translation units and sometimes by considering information from style sheets and document templates, for instance. Database-oriented TM systems sometimes apply a reference text approach to retrieve translation units by using bilingual files from previous translation projects and combining these with TM databases (Reinke 2013: 33). Typically, it is assumed that exact and context matches can be recycled without adjustment, and this seems to be the reason why many freelance translators are not paid to edit these matches.

If the source segment is not identical, but to some extent similar to a segment in the TM, the TM retrieves a fuzzy match. The degree of similarity for a fuzzy match can range from one to 99%. If the TM does not contain a similar segment, we talk about a no match. If the threshold is set at 70%, all match values below 70% are treated as no matches. In traditional TM systems, the translator would have to translate such segments from scratch, but ever since TM systems have integrated MT features, translators (or agencies) can decide to let the MT engine translate no matches. These are referred to as MT matches or AT (automatic translation) matches.  

Translators can engage with a TM system using a so-called interactive or pre-translation mode. Using the interactive mode, translation matches are presented to the translator segment by segment, i.e. one at a time, giving the translator the option to accept, revise or reject the translation proposals. When using the pre-translation mode, the translator (or the LSP) pre-translates the source text using TM and/or MT. Pre-translation is also referred to as batch-processing, which typically takes place during project preparation. Here, a complete source text is compared with segments in a TM database, and retrieved matches are automatically inserted into the editor’s workspace. If MT is also used, segments below the TM threshold are machine-translated and the target text field is pre-filled with MT output. In the interactive mode, translators can activate the MT engine whenever the TM database cannot retrieve a match. In both scenarios, when translators accept or edit the suggestions, translation units are stored in the database for future reuse.

If MT is combined with TM, translators no longer translate sentences from scratch. Instead, they simply act as post-editors who correct errors and other inadequacies to obtain a target text that fulfils the aim of the translation. According to Esselink (2020: 121), the content that professional translators translate from scratch is decreasing to such an extent that the nature of the translation profession is changing. For instance, TM systems can now recognise content that clients or translation agencies do not want translators to translate. Such content is referred to as non-translatables. They contain characters, symbols and words that often do not need to be translated, e.g. numbers, formulas, codes, email addresses, people and product names and currency. If this feature is enabled, these non-translatables can be confirmed and locked during the pre-translation process.

The translation editor

Some early commercial TM tools, like Trados and Wordfast, integrated their editor into a word processor, such as Microsoft Word; but nowadays most TM systems offer word-processing-like functionalities in a dedicated translation environment for working with content to be translated, either online or offline. It is the translation editor or the front-end of the system that shows the source text, terms stored in the integrated termbase and TM matches. The translator uses the editor to open source texts in various file formats. In this way, the editor basically supports reading and writing texts. Figure 3, below, illustrates how the editor of Phrase TMS displays source segments together with a workspace into which retrieved matches from the database are automatically imported. Hence, the editor can be said to be the TM system’s workbench. Like most systems, Phrase TMS applies a horizontal or tabular presentation model with the target language workspace of the translator appearing beside the currently active segment. Other systems apply a vertical presentation. The editor is also where translators accept, reject, and edit translation proposals retrieved by the database as well as translate source text segments from scratch. The source segment and its translation are saved into one or more TM databases when a segment is confirmed by the translator in the editor.

Example of translator editor
Figure 3: The TM editor from Phrase TMS (screenshot provided by Phrase)

Below the editor, the translator can see a real-time preview of the source text or the target text. When the translator clicks on a segment in the editor, the corresponding text is highlighted in the preview window. In this way, the feature allows the translator to see the content of the active segment within a broader context. Clicking on text within the preview makes the translation grid indicate the corresponding segment for editing. The translator can also switch from the “preview” mode to a so-called “context note” if the file format contains additional information of relevance for translation.

In the Phrase TMS editor (see Figure 4, below), translation suggestions coming from the TM database, termbase and MT providers can populate the translation editor directly under the “Target” tab in the editor’s right column or be displayed in the far-right window, called the CAT results interface (cf. Figure 3). Suggestions are displayed together with a match value (a coloured score), which in the case of TM matches indicates the similarity between the source text character string and segments included in the database. In contrast to TM match values, those corresponding to MT matches are underlined. These values are quality estimations calculated by means of a Machine Translation Quality Estimation (MTQE) feature. Based on previous results, MTQE uses a machine-learning engine to predict the usefulness of the raw MT translation for post-editing purposes. If a termbase is activated and the source segment contains a stored term, it will be displayed in the CAT results interface and highlighted in the editor source segment. Here, results will also appear if the translator runs a concordance search in the “Search” pane at the bottom right (cf. Figure 3). This feature allows translators to search for and retrieve text fragments below sentence level from the TM database. After activating the “Search” pane, translators can type or paste a word or a string of words in the appearing window. Translators are then presented with a list of translation units from the TM, with the searched-for word or string of words highlighted. If a suggestion seems useful to the translator, the translator can copy-paste it into the relevant target segment.

Exmple 2 of editor
Figure 4: Types of translation suggestions in Phrase TMS (screenshot provided by Phrase)

Machine translation

We have already explained how the fundamental idea of a TM system is to store and retrieve translation proposals from a database that contains paired and aligned segments from previously translated texts. Thus, the TM proposals are derived from a human translation or, as is mostly the case nowadays, from a translation that has been confirmed by a human translator. In contrast to translation using a TM, the idea of MT is to translate texts from one natural language into another using computers without human involvement, though human translations are used to train MT engines and humans may need to post-edit the raw machine translated text. Both TM and MT technologies are data-driven technologies that retrieve translations from other texts. When the MT technology was first invented in the 1960s and until the 1980s, the data approach was typically rule-based, while the approach in the 1990s was mainly statistical, drawing on multilingual corpora. From around 2014, these approaches were displaced by NMT or hybrid systems. NMT is based on so-called artificial neural networks, i.e. AI, aiming to reproduce the learning processes of the human brain. NMT systems are trained using all sorts of material, including TM databases, termbases, and multilingual and monolingual corpora (e.g. Christensen, Bundgaard, Schjoldager et al. 2022). The technology makes use of algorithms based on linguistic patterns to locate pre-existing candidate translations. Nowadays, commercial TM systems offer interfaces to a variety of MT engines. Arguably, the integration of MT features into TM systems has blurred the boundaries between TM and MT. As a consequence of this integration, translators are now presented with TM proposals produced by humans as well as output generated by a machine. Lines have also been blurred because MT engines integrated into a TM system are often trained using in-house human-produced TM databases and termbases to increase MT quality.

The TM system from Phrase currently connects to 30+ generic and custom-built MT engines from AI-powered NMT tool providers such as Amazon, DeepL, Google, and Microsoft. The MT features are embedded into Phrase TMS and, based on past performances of the MT engines, taking into account domain, language pair and historical edit distance, the tool automatically selects the most suitable engine for each translation job. This feature is called Phrase Translate. Finally, Phrase Translate now offers a generic NMT engine called NextMT supporting sub-segment matching. When a fuzzy match is found in the TM database, NextMT uses the matching parts and then machine translates the non-matching parts. The tool supports formatting and placeholder tags, and, if a termbase is activated, it provides all terms in the termbase with correct morphological inflections. Currently (January 2023), NextMT supports four language combinations, all including English.

Terminology management

A termbase is a bilingual glossary of paired terms stored with metadata such as definitions, grammar and status. It is a terminology management program that stores, retrieves and maintains key terms and their translations, for instance, customer-specific terminology. When a translation segment is activated in the editor, a term recognition feature compares the segment against the termbase. If a source term is found for which the termbase includes a target language match, the target language term is displayed. This applies to all match types. Either before, during or after the translation process, a term match may be imported to the active segment in the editor, which means that terms are added manually. Terminology can also be inserted automatically into the target text segment. To speed up the process of expanding the population of terms, a terminology extraction feature can be used to extract candidate term lists based on term frequencies in machine-readable source texts and target texts. Terms may be extracted monolingually from source texts or bilingually from bilingual texts (typically, TMs). Regardless of the extraction method, term candidates must be evaluated and eventually imported into the termbase by a human.

Translation quality assurance

Before completing a translation job, a translator and/or a project manager can activate a QA (quality assurance) feature. This is an automated check, which searches for basic errors and quality issues. The automated QA is limited to detecting issues with spelling, numbers and dates, untranslated segments, untypically short or long segments, inconsistent tagging in source text segments and corresponding target text segments, the use of terminology and forbidden terms stored in the database as well as unedited TM matches and MT matches. In Phrase TMS, a list of general warnings is displayed when the translator runs a quality check that is included in the interface of the tool. Figure 5, below, illustrates the QA interface from Phrase TMS.

Translation quality
Figure 5: Quality assurance interface from Phrase TMS (screenshot provided by Phrase)

An automatic QA does not resemble a prototypical revision or review process carried out by humans, as it is limited to issues that it controls and can identify; it cannot evaluate the accuracy or usefulness of a translation for a given purpose.

back to top

[heading title] Translation management system workflow

As stated above, today many TM systems are an integral part of a Translation Management system (TMS). Taken from Esselink (2020: 112), Figure 6, below, shows a basic setup of a TMS and some workflow steps.

A TMS aims “to streamline, automate and manage the entire translation process, from handing off the source content to be translated, through translation and review, to ultimately publishing the final translated content” (Esselink 2020: 110). It is basically a platform for creating, forwarding, and finishing translation projects and for connecting all parties in the translation supply chain (customers, project managers at LSPs, and translators) in a single automated workflow. Today, they are typically cloud-based systems that can be accessed by all users via a web browser, allowing for collaboration between multiple translators in real time. When translators connect to the platform via their browser, they can access the integrated TM editor. Project managers use the platform to carry out pre-analyses of the translation assignment, to keep track of translation assignments and to communicate with translators and customers. Therefore, a TMS has a broader scope than a TM system, but their functionalities overlap as they both include functionalities such as TM, MT, and terminology management features.

TMS workflow

Figure 6: TMS environment and workflow (Esselink 2020: 112)

A TMS can be managed by customers (end-users of translations) or by LSPs. Customers that invest in a TMS typically “want to keep a tighter grip on their translation processes, workflows, language assets and selection of translation suppliers or language resources” (Esselink 2020: 111). Even if customers do not want to invest in the translation technology themselves, many still expect their LSP to use a TMS to increase productivity and cost savings. Based on Esselink (2020: 111-112), a prototypical process when using a TMS includes the following steps:

  • Customers’ selection of a text that needs to be translated from a content platform (e.g. a content management system).
  • Automatic transfer of content to the LSP.
  • Automatic project creation and conversion of text into a TMS-readable format
  • Analysis of a source text against the TM database and generation of a word count report, which is submitted to the customer, outlining the number of words to be translated, volume of reuse from the TM, costs and expected delivery time.
  • Approval of a translation quote by the customer.
  • Pre-translation of content using TM and, in cases where there is no match in the TM, content is pre-translated by means of MT.
  • Submission of the translation job to translators who can accept or reject the task.
  • In the translation editor, the translator is offered TM and/or MT matches and terminology from the termbase, which the translator can reject, edit or accept. To complete the translation process, the translator and/or the project manager typically runs a QA.
  • Submission of the translation to a reviewer, e.g. another translator and/or the customer for check.
  • Submission of edits, comments and general feedback to the original translator, who goes through and eventually implements them.
  • Delivery of the final translation to customer.

For an example of a workflow in a specific LSP, see Bundgaard (2017: 95).

To sum up, when machines translate and translators are mainly left to correct errors, translators become peripheral and are no longer at the centre of the translation process: they become agents delivering input to a complex process defined by technologies and other agents, such as LSPs and customers. This peripheral role of translators has been problematized by LeBlanc (2017), concluding that, faced with the inevitable TM implementation and enforced recycling of TM matches, translators generally experience a loss of professional autonomy and a decline in professional satisfaction. In this light, the widespread uptake of TM technology certainly calls for more research on all aspects of translators’ interactions with TM systems, among others as part of a TMS.

back to top

[heading title] Research on translators' interaction with a TM system?

Though TM technology definitely pervades modern-day translation practice, it has received little attention within translation studies (e.g. O’Hagan 2013: 506ff). In a mapping study of translation technology research published between 2006 and 2016, Christensen, Flanagan & Schjoldager (2017: 14) found that translation workplace studies were scarce and so were studies on translators’ interactions with translation tools and the effects of this on translators’ minds and work processes. Instead, research had primarily focused on technical aspects and quality assessment, the implementation of technology in the language industry and the impact of technology on the profession and on translator training. Recent research seems to focus either on TM (e.g. LeBlanc 2017; Sannholm 2021) or MT (e.g. Daems, Vandepitte, Hartsuiker et al. 2017; Ortiz-Boix and Matamala 2017), and lately, it seems that more research attention is paid to MT than to TM, probably due to the remarkable development of MT in recent years. Considering the fact that TM is the preferred tool of the profession and that it is typically combined with MT, we will now describe some recent empirical studies that focus on translators’ interaction with TM and MT in combination. For this, we will draw on our own research within this particular field as well as a systematic search for studies published in central translation-studies journals between 2017 and 2022. As in Christensen, Flanagan & Schjoldager (2017), a limited number of studies was found.

Bundgaard (2017) conducted a workplace study of MT-assisted TM at a large Danish LSP. In an ethnographically inspired experiment, she studied eight professional translators’ interaction with MT-assisted TM in SDL Trados Studio, including the translators’ acceptance, rejection and revision of various match types, time spent on editing TM and MT matches and the translators’ attitudes to MT-assisted TM. The study showed that the MT-assisted TM process involves highly complex interactions between the translator and the tool. In this process, translators manage different translation suggestions and draw on different resources and functionalities while considering client preferences and the situational context of the target text. For instance, the study showed that, in a technical text, 9% of MT matches were accepted without changes. In terms of time spent editing TM and MT matches, also in a technical text, results showed that the average editing speed for TM matches with match values between 70 and 74% and for MT matches was similar, whereas the editing speed was higher for TM matches with higher match values. Although they identified many negative aspects of the process, particularly related to MT, translators expressed a pragmatic attitude towards MT-assisted TM.

Like Bundgaard (2017), Teixeira & O’Brien (2017) conducted a workplace study. For this, keystroke logging, screen recording, eye-tracking and retrospective interviews were used to study cognitive and ergonomic aspects of ten translators’ work when editing TM and MT matches and translating from scratch in the CAT tool memoQ. The study showed that the main tool for searching for terms and expressions was the concordance feature in memoQ and that there was great variation in the way the translators used the two screens at their disposal: some translators only used one screen during the entire task, whereas others switched frequently between the screens, but all translators spent more than 90% of the time on the screen where the memoQ main window was visible. The eye tracking data revealed that roughly half of the time was spent looking at the target segment and that less than 20% of the time was spent looking at the source segment. The authors also found that consultation of resources is one of the workplace activities that requires the highest effort.

Bundgaard & Christensen (2019) also reflected on translators’ interaction with translation resources. Drawing on data from Bundgaard (2017), they studied translators’ use of resources during the post-editing of TM and MT matches. The study showed that concordance searches accounted for approximately 74% of translators’ resource consultations, and that the concordance feature is typically the translators’ first choice of resource. Interestingly, concordance searches were more frequent in the post-editing of MT matches than of TM matches, which indicates that the concordance feature is a central resource for the post-editing of MT. The results also indicated that translators verify the MT output against the TM, even when there is no perceived translation problem. Considering the results of Hvelplund (2017), who found that bilingual dictionaries are the preferred resource in non-aided translation, Bundgaard & Christensen (2019) suggest that using a CAT tool changes translators’ search patterns.

Using an agency theory lens in a focus group study of human factors that determine professional translators’ (non-)adoption of MT, Cadwell, O’Brien & Teixeira (2018: 314-315) point out that MT and TM are clearly intertwined in practice and that translators typically trust outputs from TMs more than outputs from MT systems. Thus, for instance, focus-group participants argued that, while both TM and MT segments can contain errors, “TM segments are at least consistent in the type of errors they can contain, and that differences in TM matches are highlighted, but that MT output errors are unpredictable, inconsistent and foster distrust” (Cadwell, O’Brien & Teixeira 2018: 314). It is not mentioned whether or not the translators have experience with NMT, but, judging from the date of the data collection, their results are probably based on statistical MT.

The only study that we know of that explores literary translation in combination with TM and MT and other technologies is a study by Slessor (2020), who provides much-needed insights into the practices of literary translators. Based on an online questionnaire with 40 literary translators, Slessor’s results show that a large majority of literary translators never use technologies such as TM and MT. More specifically, 77.5% said that they never use TM, 7.5% rarely use it, 10% use it occasionally, while only 5% often do so. In terms of MT, 70% reported never using it, 20% rarely use it, 10% occasionally use it, while no one used it often. Interestingly, some translators reported that they use TM for maintaining formatting or for conducting collocation research in bitexts, which indicates that concordance searching can also be a helpful resource for literary translators. According to Slessor, cost and lack of familiarity with CAT tools are barriers to the adoption of specialised translation technology.

back to top

 Research potential Research potential

Professional translators are now so accustomed to working with a TM system that they may feel lost in translation without it, not least because it helps them increase their productivity in a highly competitive market. However, the productivity increase goes hand in hand with lower translation fees and some would say a more peripheral role in the translation process. Also, the translation workflow is often controlled by LSPs or customers, who own the TM databases. In any case, today professional translation is without a doubt an instance of human-computer interaction, and we argue that it is of utmost importance to further investigate this relationship.

While translation studies has been rather reluctant to consider translation technology as an area of interest, other disciplines like computational linguistics and computer science have been keen to research the development of MT. However, this research has mostly focused on the software rather than the human agents and the combination of MT with other technologies. We therefore argue that translation studies needs to catch up with recent technological developments and foster more interest in translation technology research, taking the human agent into account. This leads us to conclude this entry by a list of suggestions for further TM research focusing on translator-computer interaction:

  • We need to know more about how translators interact with NMT-assisted TM.
  • We need more studies on how translators interact with and perceive various recent software features, such as the sub-segment integration of TM and MT.
  • We need more investigations into professional translation workflows, including collaborative interactions between translators, project managers, clients, TM systems and other translation tools.
  • We need to know more about how the technology impacts on job satisfaction and perceived status in society.
  • We need more investigations into the use and potential benefits of MT-assisted TM in other contexts than translation for special purposes, e.g. within literary translation.
  • Though TM technology has gradually made its way into translator training, we still need to research how TM technology is taught and which competences should be taught at both undergraduate and graduate levels. It would also be interesting to explore the potentials of integrating TM or MT-assisted TM into computer-assisted language learning in general. 

back to top

Creado con eXeLearning (Ventana nueva)