[Please note that some of the entries are very old.*]
Showing posts with label termbase. Show all posts
Showing posts with label termbase. Show all posts

Friday, 10 July 2015

CafeTran Espresso 2015's Resources, an Overview

Historically, CAT tools only offered the possibility of re-using your translations. The document that needed to be translated, was split up into parts called “segments” (usually sentences), and the translated segments were stored in a memory, the Translation Memory (TM), which was referred to during the translation. In due course, more and more resources have been added. The following is an overview of the resources CafeTran Espresso 2015 currently offers. The number of resources you can connect to is unlimited, so you can have any number of TMs, Termbases, and web resources simultaneously.


  • The TM. The translation memory for segments. It’s still there, and it’s the only resource you really can’t do without. It’s a TMX mark-up file that offers far more features than the original TMs, including fuzzy matches and easy but extensive maintenance from within CafeTran.
  • The Termbase, a memory for words and phrases, very similar to the TM, with the same benefits.
  • The Glossary (also called Lexicon or tab del) for terms and phrases, a plain, tab delimited text file that’s past its expiry date. It doesn’t offer the goodies the Termbase offers, and it’s a left-over of the days you could only “connect” to one Termbase.
  • The Database. In CafeTran, a simple relational database (SQLite by default), that allows for very fast searching gigabytes of data (think of the DGT) because it’s indexed. It can also create a (TMX) file by “recalling” data from the database.
  • Web resources. You can search websites from within CafeTran, and the results are shown in the CafeTran interface.


All resources can “float” so you can give them a more prominent place on you screen, or even move them to a second monitor. Click to enlarge.

  • Machine Translation (MT), also integrated in the user interface. Results can be triggered by going from a segment to the next one, or manually. The free MyMemory is the default MT, Bing and other MTs are included in the Options.
  • TM-Town. A very recent development. I’m afraid I haven’t tried it yet. You search your own resources, but it offers extras like tokenisation.
You may also want to have a look here for my take on the settings of TMX resources.

Friday, 13 March 2015

TMX Files, an Approach

All TMX files are standard translation memories, so they are equal. But if you treat them equally, it’s unlikely you’ll benefit the most of them.

For this approach, I’ll assume you use Auto-Assemble (AA). This may not be the best approach in all situations, nor for all language pairs. However, I think AA is very useful for most situations, in fact I consider it to be the core feature of a decent CAT tool.
AA comes up with suggestions based on the CAT tool's algorithms and your (priority) settings of the connected memories for segments (TM), and memories for words and phrases (termbase). If you get the wrong "hits," you can change those settings, and/or add the correct term in the TM/TB with the highest priority, so next time, it'll show up correctly.
In Menu | Edit | Options | Auto-Assembling, you can select if you want to use the Auto-Assembling Panel (a pop-up panel), or the Automatic Insertion of Matches. I prefer the latter, however, I can imagine automatic insertion of those results can be counterproductive, especially if the word order of the target language differs from the one in the source language. That doesn't make the actual results less useful, though. Besides, when you arrive at a new segment, the AA results have been selected, so you can delete them with your very first keystroke if you don’t like the result. Memorise the results you do like, though.

You can use TMX files for memories for segments (TM), for memories for words and phrases (termbase), or for both. Whenever you create or add a new memory, you’ll have to indicate how you want to use it in the New Memory dialogue. It’s the first entry, under Memory Type. Or you can select it in the Dashboard, using the gear wheel at the bottom. Beware, however, that in the Dashboard, you cannot use different settings for your various TMX files. For the Dashboard, all TMX files are equal… In my approach, using a TMX file for both segments and words and phrases isn’t very useful.

TMs
For me, different TMs require different settings. It is of course possible to use only one TM, but it’s far more likely you’ll end up having several. I distinguish between:
  • ProjectTM. A TM for segments. It’s the TM CT wants you to (de)select first in the Dashboard. I strongly suggest you select it, also because it can play a huge role in Auto-Completion. Since it’s usually not a big file - though it’ll “grow” during the project - you can use it to automatically save your work very frequently (I set Autosave Project to after two segments in Options | Workflow, all other TMs I set to 5). It allows you to save tags, and this is the only TM for which I think this is useful. It’s also very suitable to check consistency within the project (see QA). The latter means that you should set the ProjectTM to Keep Newer Duplicates when you create or open it. And since this is your current job, you should assign the highest priority to the ProjectTM. Settings: In short, in the New Memory dialogue, you should select (from top to bottom): [Memory Type] Translation Memory, Processing Tags, Terms Consistency Check, [Options] High Priority, Automatic, Fuzzy and Hits, Keep Newer Duplicates.
  • Any memories for segments provided by the client. Ideally, they are very important, and should be used for high-priority hits and consistency check. You don’t want to “pollute” them with your own translation, so they should be set to Read-Only. The settings: [Memory Type] Translation Memory, Terms Consistency Check, Read-Only, [Options] High Priority, Automatic, Fuzzy and Hits. Since they are Read-Only, you don’t have to worry about the duplicates. You may have to review those settings, as som client provided TMs are pure faeces.
  • A general memory for segments (Big Mama). This is the TM in which you keep all your translations for the language pair. It’s optional of course, but results from your Big Mama may surprise you positively. It may become too big to use for AA, so you may have to set it to Manual workflow integration. Exclude the Big Mama from consistency checks when doing the QA at the end of the project is of the essence. The settings: [Memory Type] Translation Memory, optional: Pretranslate Only, [Options] Low Priority, Manual, Fuzzy (Fuzzy and Hits will take much longer to assemble, this goes for all TMs, of course), Keep All Duplicates.
  • Huge third-party subject specific memories for segments, like the DGT for EU jobs. The settings: [Memory Type] Translation Memory, optional: Pretranslate Only, [Options] Medium Priority, Manual or Pretranslate, Fuzzy (Fuzzy and Hits will take much longer to assemble, this goes for all TMs, of course).
  • Other subject specific memories for segments: You’ll have to decide the settings based on the situation. Since I use a Big Mama, I don’t have much experience with them. If those TMs are from other sources than your client or your own jobs, be very careful.
  • Memories for terms play a major role in my approach. I don’t use a project specific termbase, because the project terms will show up in my high-priority project specific TM anyway. However, I do use a
  • Big Papa, the equivalent of the Big Mama for words and phrases. Add to it as many general words and phrases as you can, it will pay you back generously. The settings: [Memory Type] Termbase, [Options] Low Priority, Automatic (unless it gets too big, which is less likely than in the case of your Big Mama), Fuzzy, Keep All Duplicates.
  • Any memories for terms provided by the client. See memories for segments provided by the client. The settings: [Memory Type] Termbase, Terms Consistency Check, Read-Only, [Options] High Priority, Automatic, Fuzzy. Since they are Read-Only, you don’t have to worry about the duplicates. You may have to review those settings, as some client provided TMs are pure faeces.
  • Subject specific memories for terms (rather than a client specific ones). You may want to use more than one. They are used next to the Big Papa, and should overrule it. Arguably the most important TMX memories I can think of. The settings: [Memory Type] Termbase, [Options] High Priority, Automatic (unless it gets too big), Fuzzy, Keep Newer Duplicates.
  • Huge third-party subject specific memories for words an phrases, like the IATE for EU jobs. The settings: [Memory Type] Termbase, optional: Pretranslate Only, [Options] Medium Priority, Manual or Pretranslate, Fuzzy.
UPDATE: Since the introduction of Total Recall, the "pretranslate" function seems to have become redundant. Unless the TM, resulting trom Total Recall, is too large to be processed the regular way.


CafeTran 2015 For Switchers

CafeTran 2015
for Switchers



In this introduction, I take it you are familiar with the basics of CAT tools.

  • Run CafeTran
  • The user interface will appear, showing the Dashboard, with from left to right a column for Translation Memories (for segments, “TM”, *.tmx), TMX Termbases & Glossaries (*.tmx files for terms and phrases, “TB” -- tab delimited text files for terms and phrases, “Glossaries”), Resources (integrated Internet resources, e.g. IATE), and your Project.
  • For a number of reasons, it’s a good idea to enable the Project Memory, and adjust the settings for it by clicking the Settings icon. Select your language pair for this job.
  • Drop your job on the Dashboard. The job can be a Document (monolingual files, various formats, a Project (bilingual files in various formats: XLF, XLIFF, SDLXLF, SDLPPX, TXML, and TTX), a TMX file to edit, and more. Check if the language settings correspond with the setting you selected in the previous step. They should. Please also check the Segmentation field. For most source languages (SL), Automatic Segmentation will be the obvious choice.
  • Click OK.
  • The project will be loaded into the top-left pane, your editing space is the pane next to it, whereas below you’ll see your enabled resources in a tabbed pane.
  • You could start working now, but I think it’d be wise to add and check a few things first:

  1. Go to Menu | Edit | Options, and check at least if the settings in General, Workflow, and Memory are set according to your way of working. In Memory, increase the Java heap if you work with large TM files.
  2. Add any tmx TMs, in Menu | Memory | Open Memory, and fine-tune their settings.
  3. Add any tmx TBs, in Menu | Memory | Open Memory, and fine-tune their settings.
  4. Add any Internet resources if you didn’t add them in the Dashboard earlier, Menu | Tools
  5. Add any Glossaries if you didn’t add them in the Dashboard earlier, Menu | Glossary | Add Glossary, and import them in a tmx TB to make use of the more advanced features of TMs.
  6. Add any resources the client sent you by importing them in Menu | Memory | Import. You can import various file formats, including SDLTM and TBX. You can later save them as TMX files. It’s probably a good idea to set those resources to Read-Only so you can’t add segments or terms and phrases to them.
  7. Add a DBMS file or import a TMX file or a database in plain text. (Advanced)
  8. Check if you’re happy with the lay-out of the GUI. CafeTran’s UI is highly customisable.
  • Click the coffee cup with the green arrow and Next next to it to start working.
  • Tags in CafeTran are placeholders, the numbering always starts with 1 in every segment, and entering them is as easy as typing the number and hitting the Esc key.
  • CafeTran offers spell checking on-the-fly. You may have to add a free Hunspell dictionary for your target language.
  • If you enabled Auto-Assembly and Automatic insertion of matches, perfect TM matches and the most likely candidates for your TL will appear in the TL pane. Other ways to add words and phrases to the TL pane are typing, typing using Auto-Completion, selecting terms or phrases in your resources (selecting only will do the trick, no deed to copy/paste).
  • Click the coffee cup with the green arrow again to save the segment and go to the next one. Warning: the translated segment pair will be added to all TMs (for segments) that are not Read-Only. Mind your Settings! They are all-important.
  • Add words and phrases to your TBs by selecting both the SL and the TL word(s) and clicking the coffee cup with the tiny plus sign. Unless you disabled it, a pop-up window will show up where you can make any corrections or add information. You can also uncheck any TBs you don’t want to add terms to. Mind your Settings! They are all-important.
  • Continue till the last segment, after which - in the case of the translation of a document - a pop-up window will appear to ask you if you want to perform a QA or want to export the translated document. You can find your translation in a folder CafeTran created in the location of your choice, with the language code of your TL added to it. In that folder, you will also find a copy of the SL document (again with the language code of your SL added to it), and any TMs (ProjectTM, for example) you saved to that location.   If you translated a Project (bilingual), you will have to choose QA in the Menu.