Tool Box Logo

 A computer journal for translation professionals


Issue 19-9-304
(the three hundred fourth edition)  
Contents
1. Local Machine Translation Engines
2. Changing Focus
3. The Tech-Savvy Interpreter - 100 Years of Conference Interpreting ... and Technology
4. Tips and Tricks: Hidden features in SDL Trados Studio (Guest column by Christine Bruckner)
5. Eyeing Success
6. New Password for the Tool Box Archive
The Last Word on the Tool Box
Dictionary Exchange

Like many of you, I'm looking forward to this year's ATA conference in Palm Springs. I'm looking forward to presenting about the exciting Translation Insights and Perspectives tool that I've been working on for a number of years, and I also can't wait for a repeat of last year's dictionary exchange. For the first time last year it was possible to bring your old and shelved dictionaries (maybe you have a newer or a digital edition, maybe you don't work in that particular language or subject matter anymore, or maybe you are -- hurray! -- retired) and let colleagues take them home to put them to the good use those poor dictionaries deserve and may not get to see in your office anymore.

If you're among those who think dictionaries are kind of old-fashioned and no longer useful, you should have seen the hordes of starry-eyed translators around the dictionary exchange table last year in New Orleans, many eventually running off with a couple of dictionaries clutched under their arms (see here for a recap of last year's event).

The part I liked best about the generous extravaganza was that it seemed to make both givers and takers equally happy.

So, if you're planning to come to the ATA this year, be sure to fill up your suitcase with those otherwise dusty dictionaries. And if you can't make it to the ATA but would still like to lighten the load on your bookshelves, please let me know and I will give you an address to which you can send the dictionaries. I can promise that I will lovingly place them on the table -- though I can't promise that they will stay there for very long.  

Dictionary Exchange
1. Local Machine Translation Engines

At the outstanding APTIF9 FIT Asia conference, I listened to a very dynamic talk by Lucy Park, an engineer for the Korean MT engine Naver Papago (Naver is the dominant search engine in Korea and Papago -- "parrot" in Esperanto -- is its MT engine). While Google's and Microsoft's search engines are often talked about as the dominant players in the market of easily accessible, generic MT engines, I find it intriguing that there certainly are a good number of generic local engines. Some European languages have seen the rise of DeepL, Russian has Yandex, Chinese has Baidu (and others), and Korean has Naver Papago. Korean is particular challenging because of its honorific system, but I was assured by many translators at the conference that this and other difficulties are handled much more seamlessly by Papago than by its global competitors. So I asked Lucy whether she would be willing to be interviewed. She was, and here's the result:

 

Jost: Naver Papago is a neural machine translation between Korean and 14 other languages (English, S+T Chinese, French, Portuguese, Thai, German, Italian, Indonesian, Japanese, Spanish, Russian, Vietnamese, and Hindi) that is very popular in Korea and -- according to my impression at the APTIF 9 in Seoul earlier this year -- widely used by many professional translators. It seems to me that outside of Korea it's not particularly widely used (correct me if I' m wrong). I had a number of Korean translators tell me that they like Papago because it captures the subtleties of Korean much better than your big competitors, such as Google and Microsoft. Is that your impression also and can you tell me why that is?

Lucy: The Papago machine translation team mainly focuses on the translations of Korean, English, and other Asian languages such as Chinese and Japanese. So it is quite flattering to hear that your fellow professional translators are complimenting Papago when translating the Korean language, because that is actually what we're trying to do best.

We put in a lot of effort to enhance the quality for Korean translations by acquiring as much bilingual and monolingual data as possible and cleaning the data thoroughly. There are also efforts on the modeling side. For example, last January we launched a new feature where the user can control the honorific level of Korean when it is the target language (see here). We also do various experiments that leverage the characteristics of languages or their corresponding scripts.

Jost: What about other language combinations without Korean? I (very unscientifically) looked at some translations between English and German, and the quality was not up to the standard of Google or DeepL. The first question is: Are those pivot translations with Korean as the pivot language? And the second question is (or maybe it' s more like an assumption): Is your goal to really focus on Korean in combination with other languages and leave the other language combinations up to the "big boys (and girls)"?

Lucy: When we pivot languages, we pivot with the best model available to our team, and it does not necessarily have to be Korean.

We focus mainly on Korean, because Papago was first created in Korea and therefore has many Korean users. That's why we currently focus most on CJK and English. If our users start telling us that they need English to German translation to work better, we'll try our best to enhance performance for those language pairs as well.

Jost: I noticed that you are not offering any ready-made app for using Papago' s machine translation engine within CAT tools. Why? Is the market of Korean translators so small? Or is the professional translation market not your focus?

Lucy: We have been monitoring the professional translation market, and we think it is an appealing market. However, we currently don't have plans to approach the market yet, due to other priorities. This situation can easily change in the future.

Jost: Your big competitors are committing themselves not to use any text that is submitted for translation if one uses the paid API (which professional translators typically do). Do you have anything like that? Or do you process the data that is being committed for further training purposes?

Lucy: We take user privacy seriously. We currently are not using API logs for analysis and/or model improvement, but if we do decide to use any data in the future, it will always be with permissions granted from the user.

Jost: According to your experience with Papago (this brings us back to the first question): Do you think that there is a market for language-specific, generic machine translation engines?

Lucy: Yes. According to my experience, the global translation market is full of diverse needs. Some need fast translation with okay-ish quality (as opposed to slow and high-quality), some others need domain-specific translation (for medicine, shopping, etc.), all in different situations and for different translation requirements. It is difficult for one vendor to excel in all areas and fulfil all those needs.

Likewise, if the demand is large enough for translation between several languages, there also lie opportunities if you can attract that market towards yourself.

Jost: Anything else you would like to share about Papago's plans in the near future?

Lucy: We are planning to launch offline translation models soon, and many more features to come! Please keep an eye on us.

ADVERTISEMENT

Catch the Sapphire and Discover the New Features in Across v7

Across version 7 is entering its next round. The Gilded Sapphire update comes with great additional features, such as new QM criteria, filter options for machine-translated paragraphs, and job offers from crossMarket directly in the Across Translator Edition.

Get the new update for Across v7

2. Changing Focus

Early versions of Word (up to 2003) had a DOS/WordPerfect emulation mode that allowed you to change the screen to blue and the font to white (under Tools> Options> General> Blue background, white text). Many used this for proofreading purposes because it offered a new perspective on the text and seemed to illuminate typos. For some reason it was dropped from Word 2007 on. Word 2013 silently re-introduced something similar (select View> Read Mode, and within the Read mode select View> Page Color> Inverse), and just now it has been made more prominently available in two options called "Focus" and "Learning Tools" (both accessible under View and Immersive).

These options have actually been available since the 2016 version of Word for Mac but, as I said, have just been introduced for Word 365 (since the end of July of this year) as well as Office 2019 for the Windows environment

The Focus mode is great for just focusing on your text (duh!). It's a full-screen mode (the ribbon bar only appears -- in black -- when you place your cursor at the very top of your screen), and unlike the (previous and current) Read mode, it allows you to write as well as read. On the status bar of Word there is even a link to quickly change into that mode.

Learning Tools comes with its own menu:  

Learning Tools

Different folks will have different preferences on how to focus in on a text to possibly catch things like errors. I like the Page Color options (it also changes the font color, but only for the purpose of reading the text without actually changing it for good). But the same might be more easily accessible for some by changing the Column Width (also only temporarily), focusing on only a few lines at a time (Line Focus), or of course by having the text read out aloud (with every word highlighted as it is read).

Naturally not all languages are supported for all options. Text spacing, for instance, is not available for languages with complex or connected scripts, such as Arabic. Syllabification is naturally not available for languages without syllables, such as Chinese, but it is accessible for three dozen European languages.

I was really excited to see these new options, and while I know that many of us don't proofread, edit, or translate in Word, it's good to have these options when we do.

ADVERTISEMENT

The new world economy is driven by content that's:

  • Flowing continuously instead of coming in separate orders.
  • Urgent - expected to be done in hours, not days.
  • Large in overall volume, but tiny in individual pieces.

Make sure your business isn't left behind. Find out more here.  

3. The Tech-Savvy Interpreter - 100 Years of Conference Interpreting...and Technology (Column by Barry Slaughter Olsen)

In the summer of 1919, World War I came to an end with the signing of the Treaty of Versailles. The treaty also put English and French on an equal footing as the languages of diplomacy, and with that, the stage was set for the creation and development of the modern interpreting profession. Multilateral diplomacy had a new home in the ill-fated League of Nations and the still relevant International Labor Organization (ILO). It was at these two pioneers of modern multilateral diplomacy that conference interpreting found its feet and where new technology began to disrupt how and where professional interpreters worked.

It was first at the League of Nations and later at the ILO where the "Filene-Finlay simultaneous translator" was used to attempt simultaneous interpretation. The reviews were mixed. World War II ensued before this mode of interpreting and the technology had a chance to take root, and it would not be until the Nuremberg War Crimes Trials in 1945 that the equipment dreamt up by Edward Filene and the simultaneous interpretation it made possible would be dusted off for its debut on the world stage. And the rest, as they say, is history.

Fast forward to 2019. Where do we stand 100 years later? Simultaneous and consecutive interpreting are both still very much alive. But in the last 10 years, there has been a lot of hand wringing about the future of interpreting. First, the quantum leaps in communication technologies at the beginning of the 21st century have given way to various forms of remote interpreting that are slowly but surely changing how and where interpreters work. And in the last five years, advances in artificial intelligence (AI), particularly in the fields of natural language processing (NLP) and neural machine translation (NMT), have led to a proliferation of speech-to-speech translation (S2ST) applications, some proudly trumpeting the replacement of human interpreters in business meetings, conferences, and elsewhere. Professional interpreters continue to do this demanding work while remote interpreting technology providers continue to move interpreters further from the speaker's side at a lectern or a on stage. In the case of S2ST, the technologies seek to replace interpreters altogether. Although, that is still a very long way off, in my opinion.

In 2019, just like in 1919, interpreters are nervous. They wonder what the future may hold. One thing is certain. It has been an amazing ride so far, and that needs to be celebrated. On October 3-4, 2019, the University of Geneva's Faculty of Translation and Interpreting and the ILO are co-hosting a celebration of the centenary of conference interpreting. 100 Years of Conference Interpreting: Looking Back, Looking Forward. The two-day conference will focus on interpreting research, training and practice and includes a veritable "who's who" of the conference interpreting world. I'm honored to have been invited to moderate a panel discussion on 100 years of conference interpreting practice on day one of the event.

A lot has happened in interpreting in the last 100 years, and it is important to celebrate this milestone. New technologies have been with us since the genesis of our profession creating both opportunity and controversy. And when you look at it from that angle, not much has changed since conference interpreting began in 1919. We conference interpreters are still here, and so is technology. Our relationship with it is just as rocky now as it was back then. I hope you can join me in Geneva. If not, be sure to follow me (@ProfessorOlsen) and the conference account (@conf1nt100) on Twitter leading up to and during the conference.

May the next 100 years be the best our profession has ever seen, and kudos to the University of Geneva and the ILO for not letting this anniversary slip past without due recognition and celebration.

Do you have a question about a specific technology? Or would you like to learn more about a specific interpreting platform, interpreter console or supporting technology? Send us an email at inquiry@interpretamerica.com.

ADVERTISEMENT

Memsource Mobile v2: Join the Mobile Translation Revolution

Manage projects and carry out segment-by-segment translations from the palm of your hand. The new version of the Memsource Mobile app now includes a mobile translation tool -- Memsource Editor for Mobile.

Learn more

4. Tips and Tricks: Hidden features in SDL Trados Studio (Guest column by Christine Bruckner)

This is the third in a series on tips and tricks with hidden features in common translation-related tools. Let me know if you're interested in writing a guest article for another tool.

 

The following tips and tricks are taken from my presentation "Tips & Tricks for SDL Trados 2015/2017" which I gave at the European Trados User Group Conference in summer 2017. I have now updated them for SDL Trados Studio 2019 SR-2 and spiced them up with some more of the "hidden" features.

I owe a lot of inspiration and helpful information to Paul Filkin and his MultiFarious blog. So if after reading my article you still want to know and test more hidden features in Studio, check out Paul's conglomeration of Studio wisdom -- and the plenty of useful plugins and extensions that are available on the SDL AppStore.

Fragment Matching against Whole Translation Units (TUs)

SDL Trados Studio 2017 introduced fragment matching. During TM lookup, fragment matching now also retrieves and displays smaller chunks such as

  • whole translation units
  • fragments of translation units
Trados 1

Option a), matching against whole translation units (TUs), means that complete translation units with a default minimum length of two words contained in the TM will be automatically found within longer new segments and displayed in the Fragment Matches window and via AutoSuggest.For example, if you translate a software manual and have the already translated (shorter) software strings stored in one of your TMs, you will see the software strings referenced in sentences as fragment matches even if they are too short to be recognized as fuzzy matches.   

While I am a bit skeptical about the usefulness of retrieving and displaying auto-aligned TU fragments at the sub-segment level (option b), displaying whole translation units (TUs) can often be really helpful and avoid (some) manual concordance searches.

However, the Fragment Matches window is only displayed in the foreground when there are no results available in the Translation Results window -- so useful whole TU matches could be hidden by 100% or fuzzy matches from your TM or by MT proposals. To avoid this, undock the Fragment Matches window from its position next to the Translation Results tab and move it to a separate window position so that it will be always visible in the Studio editor:

Trados 2

(If you do not manage to set up the Fragment Matches window as shown above -- don't worry: You can return to the previous layout at any time via the Reset Window Layout button in the View ribbon.)

Moreover, whole TU matches can be a valuable source for terminology: You can easily mark them in the Fragment Matches window and send them as terms to your MultiTerm termbase.

You might also want to do it the other way round: Convert a large termbase like the Microsoft Terminology into a reference TM in order to make it available for whole TU fragment matching. You can do this via the Glossary Converter or the Bilingual Excel file type (see below).

When you run a Studio analysis, you can see the number of whole TU matches in the corresponding Fragment Words (whole TU) column of the analysis report, so you might get a feeling for how much you could benefit from such fragment matches during your translation:

Trados 3

The "All Content" View

The Display Filter in the Review ribbon of SDL Trados Studio -- and its "cousins," the Advanced Display Filter and the Community Advanced Display Filter -- are very useful for hiding certain segments based on the selected filter criteria.

There is one option in the Display Filters which can actually reveal additional information: the All Content option. It allows you to display additional content that might be contained between the segments. Although this could result in a cluttered view with lots of tags, you might discover relevant context or even translation-relevant information, such as XML comments or isolated variables, graphic symbols, tagged strings, etc. (This mainly affects XML and text-based files, but also InDesign and MIF.)

Trados 5

More Context Information -- via the DSI Column and Viewer

Many users complain about missing context information when they work in the Studio Editor, especially when it comes to XML or text-based files for which there is no preview available.
When you take a closer look at the right-hand column of the Studio Editor (see also the XML sample file in the All Content screenshot above), you will see colored boxes with alphanumeric strings which represent document structure information (DSI). When you hover over or click them, they might give you helpful hints about the type, position, or context of the corresponding segments.

With the new DSI Viewer app from the SDL AppStore you can make this document structure informationpermanently visible; this is especially useful when there are several information elements available (indicated by the + sign in the DSI column; see screenshot below):

Trados 5

(Of course, you can also move the DSI Viewer window to a different position, in the same way as the Fragment Matches window.)

The Bilingual Excel File Type

The Bilingual Excel file type in SDL Trados Studio is my favorite file type because it has some very nice features that make it useful for various purposes:

  1. more easily translate bilingual Excel files, especially when you simultaneously need to review existing target content and translate missing target segments
  2.  show context information from additional Excel columns and integrate length checks in Studio
  3. use it as a workaround for adding new translation units to a TM (which unfortunately is not possible in the Studio Translation Memory view) or, more generally speaking, for converting a bilingual table into a TM

Let us now create a TM from a partially translated Excel file, and at the same time prepare this file for translation in Studio so that we can later review and fully translate it, taking advantage of context information from additional columns and the prepared TM. The following procedure might sound more complicated than it actually is -- just give it a try with a simple, bilingual table which you might want to import into a TM:

  • Check or possibly prepare your Excel file: source content in one column, target content in a different column; optionally context information, comments, and length restrictions in separate additional columns.
  • Make sure to save your file in the .xlsx format (not as .xls or .csv)
  • Create a new Studio project and, before adding the Excel file, click on the File Type Identifier. Enable the Bilingual Excelfile type by moving it to the first position among the Excel file types (or by unchecking all other Excel file types).
  • In the Common options of the Excel Bilingual file type, map the individual columns from your Excel file to the text boxes in the Columns section (and optionally in the Context Information and Comment Information section):
Trados 6
  • Now add the Excel file to your Studio project and prepare the project in the usual Studio way.
  • Open the resulting sdlxliff file in the Studio Editor in order to check whether the Excel columns have been mapped correctly.
  • Run the batch task Update main TM: All confirmed bilingual segments will be updated into your current TM. You might want to check the results in the Update Main TM report or the Translation Memory view.
Trados 7

In the screenshot above, I have enhanced the process by converting non-translatable strings into inline tags via the Embedded Content section of the Bilingual Excel file type.

Note that the bilingual Excel file type is cell-based, i.e., it does not further apply segmentation to the content of the Excel cells, so you might end up with some multi-sentence segments in your TM. If this is problematic for your subsequent translations:

  • You could manually split multi-sentence segments in the Studio Editor before running the update TM task (you would need to activate the Allow source editing option in the project first), or
  • you could use a dedicated TM with paragraph-based segmentation, both for updating your legacy bilingual content and for your subsequent Excel translations

And if you have a partially pre-translated Excel file, you can now just re-open the sdlxliff file and review the existing translations (using the usual QA features in Studio) and fill in the new translations, hopefully supported by many fuzzy, fragment, and concordance matches from your prepared TM.
There is, however, a disadvantage with the Bilingual Excel file type (not the standard monolingual one): When the original Excel file contains formatting (like bold, coloured text, etc.), this will most likely get lost in the target file.

How to Avoid Unpleasant Localization Surprises -- the Pseudo-translation Batch Task

Speaking of avoiding unpleasant surprises like lost formatting after translation:

When you have MS Office documents with complex tables or footnotes, or you would like to find out in advance about potential problems with restricted length translation formats, newly created XML, or text-based file types, it is always wise to run the pseudo-translation task sequence in a temporary Studio project before actually starting the translation.

I usually work with the following settings for pseudo-translations:

Trados 8
  • Use [ ] to mark segment boundaries -- so it is easier to see the start and the end of the Studio segmentation in the pseudo-target file(s).
  • Select the length modification factor according to your language combination: the default value 1,3 usually fits for translating a shorter language (like English) into a longer one (like German).
  • Choose the $ sign to represent characters in the target file (numbers will remain unchanged). This allows you to more easily discover untranslated content in the pseudo-target files.

But be aware that the $ signs will lead to problems with Excel files when worksheet names are set to translatable. In such cases, simply look for segments with DSI code WSN+ in the Studio Editor and change its $ characters into "harmless" strings (like "x") before you generate the pseudo-target file(s).

Trados 9

 

About the author

Christine Bruckner holds university degrees in translation and in computational linguistics. She was one of the early adopters of CAT tools in her freelance translator's life in the 1990s. Between 2001 and 2017, she worked for several German corporate and government language services and an LSP, where she was responsible for the introduction, administration, localization engineering, and user training and support of CAT, terminology tools, and translation management systems (mainly SDL Trados products), as well as machine translation solutions. Since 2018 Christine has worked freelance as a translation technology consultant, mainly for corporate customers.

ADVERTISEMENT

Translation software that is trusted by over 250,000 translators!

Download your free 30-day trial of SDL Trados Studio 2019 

Read more about why SDL Trados Studio is the market-leading CAT tool

5. Eyeing Success

Translator and industry commentator Valerij Tomarenko (see here for his blog) just published a substantial 300-page book in the BDÜ Fachverlag called "Through the Client's Eyes: How to Make Your Translations Visible." He sent me an early copy and ... I really liked it!

Here is Valerij's central argument: "A poorly formatted translation bears no visual proof of an outstanding piece of content work, nor any sign of your signature packaging. A translator's failure to produce a consistent, visually attractive document will not only negatively affect how it's perceived by the client, but will make the result close to unusable." (p. 202) The "packaging" he mentions here is one of Valerij's "translator's four Ps" (professionalism, project-oriented approach, personality, and packaging), which he sees as the four pillars to a successful life as a translator. The book specifically focusses on the packaging (i.e., how you present a translated product to the client), which in turn expresses at least two of the other three Ps.

While all those Ps may be a little confusing, the book is super-practical in its guidance for ensuring that the translation product is superiorly packaged. Valerij gives very helpful tutorials on fundamentals of design and typography, as well as how to use tools like Word, PowerPoint, InDesign, and Photoshop and how to deal with pesky PDFs and graphics, among many other things.

For me it was a great reminder that while most of my effort as a translator naturally goes into the translation part of my job (dealing with the meaning of text in and between two languages), in my client's eyes the packaging of my deliverable is probably much easier to evaluate, and that packaging might very well be the clincher that awards me their next job as well.

6. New Password for the Tool Box Archive
As a subscriber to the Premium version of this journal you have access to an archive of Premium journals going back to 2007.
You can access the archive right here. This month the user name is toolbox and the password is ata60DictionaryExchange.
New user names and passwords will be announced in future journals.
The Last Word on the Tool Box Journal
If you would like to promote this journal by placing a link on your website, I will in turn mention your website in a future edition of the Tool Box Journal. Just paste the code you find here  into the HTML code of your webpage, and the little icon that is displayed on that page with a link to my website will be displayed.
If you are subscribed to this journal  with more than one email address, it would be great if you could unsubscribe redundant addresses through the links Constant Contact offers below.
Should you be interested in reprinting one of the articles in this journal for promotional purposes, please contact me for information about pricing.
© 2019 International Writers' Group