[Please note that some of the entries are very old.*]
Showing posts with label Big Mama. Show all posts
Showing posts with label Big Mama. Show all posts

Saturday, 11 April 2015

The Big Mama, an Approach


A Big Mama (BM) is a general memory for segments. It is the TM in which you keep all your translations for the language pair. It’s optional of course, but results from your Big Mama may surprise you.

Some people doubt the use of a Big Mama, but at the end of this entry, I will list its advantages.

As I said, the BM is where you save all your translations. You can of course add TMs from other sources, but you’ll have to make sure they are of at least the same quality (and unfortunately this isn’t always the case), and not excessively specific. If you never do chemical stuff, it’s not recommended to add a very specific chemical TM to your BM, since it will be used for Auto-Assemble (AA) and also for Auto-Complete (please see the relevant entries in the Wiki) and is therefore probably counter-productive.


When you start using a BM, I take it it’s rather small. This allows you to benefit most of the TM settings. The recommended settings:




  1. Memory Type. It’s clear we’re dealing with a Translation Memory (a memory for segments). Leave all the other options in this section disabled (though some people think it’s a good idea to enable “Processing Tags”, I have my doubts).
  2. Priority. Set it to “Low Priority”. As a rule, all other attached TM are more important.
  3. Workflow Integration. With “Automatic” you will benefit optimally from Auto-Assemble. CT will search your BM real-time.
  4. Matching Type. “Fuzzy and Hits” will offer you most results.

If you use your relatively little Big Mama at this stage, automatic concordance search and AA will show up instantly when you move to the next segment.

Unfortunately, as your Big Mama grows, the delay will increase.
When you notice that, the first thing you should do is increase the RAM assigned to CT in Menu, Edit | Options | Memory | Java Memory Size. Do not assign more to CT than the physical memory of your computer. Increasing the RAM improves the speed of concordance search and AA without affecting the results you’ll get.
If the delays increase again, the next step would be to set the Matching Type (3) from Fuzzy & Hits to Fuzzy. This will have some impact on your results, but I bet you can still live with it.
Switch to the "Pretranslation" mode from the default "Automatic" in Memory Setting in Project Info and check the Translation in Review box. When you start the translation, you will se a progress bar that indicates the background pretranslation. There's no need to wait till the pretranslation is ready, just start working. However, it’s recommended to use all attached TMs for the pretranslation (together with the Big Mama), since later changes will not be included in it. This also almost requires you to use a Project TM with the highest Priority, and set for use in the QA. 
And when your Mama really gets out of hand, you should set the Workflow Integration to Manual: For use with the Search function only.

There are two more ways to still use an oversized BM, but they (seemingly) defy the Big Mama concept, saving all your translations in one file:
You can set your BM to Read-Only. It will use considerable less RAM, making the process faster. I don’t know if this has any consequences for the results, since in this setting, the segments are stripped from their internal XML codes. After finishing the translation, you should use your ProjectTM to update your BM. More on the ProjectTM later.
Import your BM into the external database: Menu, Edit | Options | Total Recall | CAT Tools Exchange | Load from TMX Memory… Since the table will be indexed, concordance searches are blistering fast, but you won’t benefit from AA. Unless you use the Recall functionality, Menu, Edit | Options | Total Recall | Recall Segments to Memory… This kind of reverses the process, CT will create a memory that consists of only the relevant segments. More on external databases later. After finishing the translation, you should use your ProjectTM to update your BM, and to update the table by simply loading it to it.
Both solutions may present problems for the elderly users who tend to forget things. I suggest running a script to remind you of it.

Maintaining a BM is a good idea for a.o. the following reasons:
  • You kind of automatically keep track of everything you ever translated.
  • The content is yours, so you can trust it (we hope).
  • You’ll be surprised how much “repetition” you will stumble upon, often from texts you did years ago.*
  • Auto-Complete will take suggestions from your BM.

* Earlier this year, I did a 20,000 words financial report. I’d set my BM to manual workflow, but I soon found out that a search in it resulted in a lot of hits. I tried “Insert All Exact Matches” from the Translation menu, and even though the workflow was manual, I got heaps of matches, some of them from 2006, possibly earlier. It reduced the job to some 2,000 words.


Friday, 13 March 2015

TMX Files, an Approach

All TMX files are standard translation memories, so they are equal. But if you treat them equally, it’s unlikely you’ll benefit the most of them.

For this approach, I’ll assume you use Auto-Assemble (AA). This may not be the best approach in all situations, nor for all language pairs. However, I think AA is very useful for most situations, in fact I consider it to be the core feature of a decent CAT tool.
AA comes up with suggestions based on the CAT tool's algorithms and your (priority) settings of the connected memories for segments (TM), and memories for words and phrases (termbase). If you get the wrong "hits," you can change those settings, and/or add the correct term in the TM/TB with the highest priority, so next time, it'll show up correctly.
In Menu | Edit | Options | Auto-Assembling, you can select if you want to use the Auto-Assembling Panel (a pop-up panel), or the Automatic Insertion of Matches. I prefer the latter, however, I can imagine automatic insertion of those results can be counterproductive, especially if the word order of the target language differs from the one in the source language. That doesn't make the actual results less useful, though. Besides, when you arrive at a new segment, the AA results have been selected, so you can delete them with your very first keystroke if you don’t like the result. Memorise the results you do like, though.

You can use TMX files for memories for segments (TM), for memories for words and phrases (termbase), or for both. Whenever you create or add a new memory, you’ll have to indicate how you want to use it in the New Memory dialogue. It’s the first entry, under Memory Type. Or you can select it in the Dashboard, using the gear wheel at the bottom. Beware, however, that in the Dashboard, you cannot use different settings for your various TMX files. For the Dashboard, all TMX files are equal… In my approach, using a TMX file for both segments and words and phrases isn’t very useful.

TMs
For me, different TMs require different settings. It is of course possible to use only one TM, but it’s far more likely you’ll end up having several. I distinguish between:
  • ProjectTM. A TM for segments. It’s the TM CT wants you to (de)select first in the Dashboard. I strongly suggest you select it, also because it can play a huge role in Auto-Completion. Since it’s usually not a big file - though it’ll “grow” during the project - you can use it to automatically save your work very frequently (I set Autosave Project to after two segments in Options | Workflow, all other TMs I set to 5). It allows you to save tags, and this is the only TM for which I think this is useful. It’s also very suitable to check consistency within the project (see QA). The latter means that you should set the ProjectTM to Keep Newer Duplicates when you create or open it. And since this is your current job, you should assign the highest priority to the ProjectTM. Settings: In short, in the New Memory dialogue, you should select (from top to bottom): [Memory Type] Translation Memory, Processing Tags, Terms Consistency Check, [Options] High Priority, Automatic, Fuzzy and Hits, Keep Newer Duplicates.
  • Any memories for segments provided by the client. Ideally, they are very important, and should be used for high-priority hits and consistency check. You don’t want to “pollute” them with your own translation, so they should be set to Read-Only. The settings: [Memory Type] Translation Memory, Terms Consistency Check, Read-Only, [Options] High Priority, Automatic, Fuzzy and Hits. Since they are Read-Only, you don’t have to worry about the duplicates. You may have to review those settings, as som client provided TMs are pure faeces.
  • A general memory for segments (Big Mama). This is the TM in which you keep all your translations for the language pair. It’s optional of course, but results from your Big Mama may surprise you positively. It may become too big to use for AA, so you may have to set it to Manual workflow integration. Exclude the Big Mama from consistency checks when doing the QA at the end of the project is of the essence. The settings: [Memory Type] Translation Memory, optional: Pretranslate Only, [Options] Low Priority, Manual, Fuzzy (Fuzzy and Hits will take much longer to assemble, this goes for all TMs, of course), Keep All Duplicates.
  • Huge third-party subject specific memories for segments, like the DGT for EU jobs. The settings: [Memory Type] Translation Memory, optional: Pretranslate Only, [Options] Medium Priority, Manual or Pretranslate, Fuzzy (Fuzzy and Hits will take much longer to assemble, this goes for all TMs, of course).
  • Other subject specific memories for segments: You’ll have to decide the settings based on the situation. Since I use a Big Mama, I don’t have much experience with them. If those TMs are from other sources than your client or your own jobs, be very careful.
  • Memories for terms play a major role in my approach. I don’t use a project specific termbase, because the project terms will show up in my high-priority project specific TM anyway. However, I do use a
  • Big Papa, the equivalent of the Big Mama for words and phrases. Add to it as many general words and phrases as you can, it will pay you back generously. The settings: [Memory Type] Termbase, [Options] Low Priority, Automatic (unless it gets too big, which is less likely than in the case of your Big Mama), Fuzzy, Keep All Duplicates.
  • Any memories for terms provided by the client. See memories for segments provided by the client. The settings: [Memory Type] Termbase, Terms Consistency Check, Read-Only, [Options] High Priority, Automatic, Fuzzy. Since they are Read-Only, you don’t have to worry about the duplicates. You may have to review those settings, as some client provided TMs are pure faeces.
  • Subject specific memories for terms (rather than a client specific ones). You may want to use more than one. They are used next to the Big Papa, and should overrule it. Arguably the most important TMX memories I can think of. The settings: [Memory Type] Termbase, [Options] High Priority, Automatic (unless it gets too big), Fuzzy, Keep Newer Duplicates.
  • Huge third-party subject specific memories for words an phrases, like the IATE for EU jobs. The settings: [Memory Type] Termbase, optional: Pretranslate Only, [Options] Medium Priority, Manual or Pretranslate, Fuzzy.
UPDATE: Since the introduction of Total Recall, the "pretranslate" function seems to have become redundant. Unless the TM, resulting trom Total Recall, is too large to be processed the regular way.


Saturday, 5 March 2011

The CafeTran UI: The Project File Simplified (3)

Below, I have tried to summarise the minimum settings required to configure a translation project for four situations. The following ignores most options (including useful ones), reviewing, billing, etc. The sole purpose is to get you started. You can always adjust your settings in the Workflow UI: Project>Project Info.

I. Configure a new translation project without existing databases

This is typically the situation for translators who are new to CAT tools, or who don't have memory or terminology databases at their disposal.
Click to enlarge
  • Click the Document (1) button and browse to the file you want to translate
  • Select or check your source and target languages (3-4)
  • Select or check the file extension (5) of the file you imported in 1
  • Check the New Project Memory box (8)
  • Click OK

Done. All other settings are either not important at this stage, or their default values will suit your purpose for the moment.

II. Configure a new translation project making use of your existing databases

This is typically the situation for translators who are familiar with CAT tools and/or have database resources of their own, or who received databases from the client.

  • Click the Document (1) button and browse to the file you want to translate
  • Select or check your source and target languages (3-4)
  • Select or check the file extension (5) of the file you imported in 1
  • Check the Load Database Memory (9) box if you have a terminology database at your disposal (a Big Papa)
  • Check the Load Memory File (10) box, and locate your memory database (Big Mama) after clicking the TMX button
  • Click OK


III. Resume a translation project

Click the button for your current project (NEW!, a new CT feature...), go to Recent Projects in the Menu Bar and select one of your recent projects, or click Project (13), and browse to it. Click OK.
Click to enlarge


IV. Start a CT supported translation project

CT supports a number of projects, including projects set up in other CAT tools. They typically have the extensions .xlf, .ttx, .tmx, etc. Click Project (13), and browse to the project of your bread and butter. Click OK.

Friday, 4 March 2011

The CafeTran UI: The Project File 2

Next to the Document Settings (2-7), you'll find two more tabs:


The Memories Settings Tab
Click to enlarge
18. Workflow integration. Manual integration - searching the memory for the matches of the current segment takes place only when click the Translate button in the target window toolbar. Automatic integration (default) - searching the memory for the matches of the current segment take place automatically as you take the next segment. The searching may take some time when you work with very large memories. Pretranslation - the memory is being searched for all document's segments in one step in the background. You can translate all your segments during this process and the results are shown instantly as you take a segment. This method is preferable while working with large (big mamas and papas) memories.
19. Matching type. Fuzzy matches. The typical CAT memory matching for the current segment with shown percentage. Subsegment matches. These are 100% matches for the subsegments of the current segment as well as the approximate (statistical) matches for the subsegments of the current segments. Detailed matching. This is character based matching rather than word based matching (for languages where there are no word separator such as a space). By selecting here you decide what kind of matching you want to have during translation. The default setting allows for both Fuzzy matching and Subsegment matching taking place.
20. Character encoding. UTF 8 is default. However, Trados tmx memories use UTF 16 encoding. So if you plan using the TMX in Trados you can change the default setting. I heard that Trados also recognize UTF 8 encoding correctly.
21. Memory for segments. Permits adding segments to the memory.
22. Read only. The memory can be queried but cannot be saved. This option is also very important for tuning the memory performance. The memory uses less RAM, and matching is faster. So it should be used with large memories set for reference only.
23. Memory for terms. Permits adding terms to the memory.


The Segment Properties Tab
Click to enlarge
24. Project. Enter the name of the project the segments come from. Based on this property, you can later filter the memory to extract only the segments with the given property.
25. Field. Enter the subject matter of the project (e.g. Medicine, IT, Legal, etc.) you add to the memory.
26 and 27. Define your own property names.


The Menu Bar
On the very top of the UI, you'll find the Menu Bar as usual. I should probably have started with it, so let's call it a proteron husteron (which in itself is a husteron proteron, of course).
Click to enlarge
Recent Projects speaks for itself. Header, well..., I hope Igor can help me out on that one. 

Sunday, 27 February 2011

The CafeTran UI: The Project File 1

[This entry replaces two older ones. You can still find them here, and here]
Igor has simplified the Project File UI. I will discuss it in three parts, starting with the Document Settings.
To start or set up a project, run Start.jar (and not Cafetran.jar). You will see the following interface:


Click to enlarge

  1. The Document button. Click this button to browse to import a document or a folder with documents into CT. The Document button is very prominent in this UI, but you probably use it only once: When you configure your project. You can use it to add documents later, but you can also do that in the workflow UI. You shouldn't use it to resume a project, unless you want to add documents. Instead, to resume your translation without adding documents, just hit Open Project... (25).
  2. The bar with the Document Settings, Memory Settings, and Segment Properties.
  3. Set your Source Language. No comment needed, too obvious.
  4. Set your Target Language.
  5. File Type. Choose the type of file you want to import. Usually, the file type field adjusts itself to the format of the document chosen in 1. CT offers you a multitude of file formats - which of course is great - but as far as I know, you cannot translate multiple formats in one project. Apart from that, if you check the available formats, it may come as a shock that there's no ttx, sdlxliff or other (Trados) file format available. Now I'm not much of a Trados fan (to put it mildly), but I have to admit that translating those files contribute a lot to my miserable income. Luckily, CT can handle those files, but it treats them - correctly - as project files, not as documents. So instead of entering all the information in 1 - 3, you just click Open Project... (25), and browse location of the project concerned. You can find more on how to handle these project files here.
  6. Tool-id. You can choose between CafeTran and CafeTran-OpenOffice. Beats me. No doubt, Igor can explain this, and the following:
  7. Phase. Choose between Translation, Review, and End.
  8. New Project Memory. Check this box only if you do not use a general memory database (Big Mama), or if you do use a Big Mama, but want to create another database specifically for this project.
  9. Load Database Memory. Check this box if you want to use an existing (general) terminology database (Big Papa).
  10. Load Memory File. Check this box if you want to use an existing memory database (usually your Big Mama).
  11. The path to the memory database mentioned in above.
  12. Billing. Let's worry about that later. Translate first, bill later. If only the other way around were possible...
  13. Project. Click this button to open an existing project, including projects of other CAT tools (ttx, sdlxliff, etc.)
  14. Translation Mode. Segmentation mode - the text is segmented "as you go forward". You create a segment and translate this segment at the same step, then next segment and and so on. Autosegmentation mode - the whole text is segmented first, and only then you start translating. That is how most CAT tools work, I suppose (Igor says). Clipboard mode - a special translation process which automates as much as possible importing from, translating and sending back the segments straight to external applications. The segments import and export takes place via system clipboard. There is a chapter in the Handbook explaining this mode in detail. Image mode - I also wanted CafeTran to be a sort of editor for translation of scanned images or paper documents such as short certificates, forms, etc. Here, you only create target language segments in the target window (without having any source segments) and export them all to the text format. You can open the scanned image in CT interface and query manually your memories and other resources. Of course, you can also OCR and convert the image documents to a word processor as an  alternative.
  15. Translation in Review. Checking this box allows you to make automatic searches in memories even when you already have the target segments created, for example, when you review the text.
  16. Documents Alignment. Check this box to align two documents (source and target) in order to put their corresponding segments to memory or extract phrases to a terminology base.
  17. Recent Resources. Check this box to use recently used databases.

Wednesday, 16 February 2011

The CafeTran UI: The Project File (intermezzo)

Igor has done it again: He simplified the Project interface after reading my previous entry on this blog.
Click to enlarge
However, he's cheating a bit - a few bytes even - because all 28 items are still there, hidden behind tabs (highlighted in yellow). This UI is far less `frightening' than the old one, so it's an improvement anyway.

In my wish list, I suggested to skip this interface all-together, and go straight to the workflow UI. Igor does not agree, and argues this `sounds good only if you want to continue the project.'

My arguments in favour of skipping the Project UI:

You can find the very same Project UI in the workflow UI under Project>Project Info. So why not take the `risk'? If you don't continue a project, set up a new one in the workflow UI.

Igor argues that it would take more steps: `first you would need to close all the memories, glossaries from the previous project and next go to the Project UI to configure them anew, which might be totally different than in the last project.' But is that true? I mean, `totally different'? I asked a question on Rosetta: How many language pairs do you work in, and how often do you switch between them? I don't have the results of this poll yet, but I bet the number of pairs is very limited. And if you work with Big Mamas and Papas for each pair, switching will not take much time. No time at all, if Igor manages to link the Mamas and the Papas (all the leaves are brown) to the various language combinations. That would only leave the project specific glossaries to select, and you'll have to do that anyway.

IMPORTANT NOTE for MAC users: For the time being, close CafeTran via the Project menu instead of using the keyboard shortcut. Igor will fix this soon. SOLVED

Sunday, 13 February 2011

The CafeTran UI: The Project File (1-7)

To start or set up a project, run Start.jar (and not Cafetran.jar). You will see the following interface:
Click to enlarge

Now people say that CafeTran has a steep learning curve, and if you look at the Project File above, you would probably agree. I hope Igor will do something about it, because it may frighten new users. Setting up a project and start translating is not that difficult at all. But a start-up screen with 28 items (sic!), some of them with a drop-down menu, is not very encouraging. Limiting this interface to the bare essentials would be a good idea, an optional set-up wizard may even be a better idea, since you can change or adjust most settings after you configured the project anyway.
Since this UI is extensive, I will discuss the various items it consists of in a number of blog entries, this one being the first and probably most (only?) really important one.
Click to enlarge
  1. The Document button. Click this button to browse to import a document or a folder with documents into CT. The Document button is very prominent in this UI, but you probably use it only once: When you configure your project. You can use it to add documents later, but you can also do that in the workflow UI. You shouldn't use it to resume a project, unless you want to add documents. Instead, to resume your translation without adding documents, just hit Open Project... (25).
  2. Set your source language. No comment needed, too obvious.
  3. Set your target language.
  4. Choose the type of file you want to import. Obvious, although I'd say you would choose this before you choose a document to import (1). [Correction: It seems the file type field adjusts itself to the format of the document chosen in 1. Wonderful, but then why is the field there anyway? To check only?] CT offers you a multitude of file formats - which of course is great - but as far as I know, you cannot translate multiple formats in one project. Apart from that, if you check the available formats, it may come as a shock that there's no ttx, sdlxliff or other (Trados) file format available. Now I'm not much of a Trados fan (to put it mildly), but I have to admit that translating those files contribute a lot to  my miserable income. Luckily, CT can handle those files, but it treats them - correctly - as project files, not as documents. So instead of entering all the information in 1 - 3, you just click Open Project... (25), and browse to its location. You can find more on how to handle these projects here.
  5. The date is entered automatically. I don't know if you can change it, and I don't care anyway.
  6. Tool-id. You can choose between CafeTran and CafeTran- OpenOffice. Beats me. No doubt, Igor can explain this, and the following:
  7. Phase. Choose between Translation, Review, and End. Why?

Monday, 31 January 2011

File Formats: InDesign IDML

There isn't really much to say about translating InDesign .idml files using Cafetran. Just import them like any regular file, and start working. Since I don't have InDesign, I had no idea of the results. They turned out to be more than acceptable to excellent. Below you see a screenshot of a part of the translation before any editing or DTP-ing, in other words, straight from CT.
Click to enlarge

Even most of the tables were OK:
Click to enlarge

Sunday, 23 January 2011

File Formats: Trados TTX

Trados TTX files may be legacy files, unfortunately they still abound. The good thing is that you can process them in CafeTran. However, you should keep a few things in mind:

CT correctly considers the files to be "projects" rather than documents. This came as a complete surprise to me when I first wanted to change the settings in the Project File. Instead of the usual interface, I got this:
Click to enlarge
No way I could make any changes.

The next thing is that you should ask for "presegmented" TTX files, or you should presegment them yourself. The reason for that is not clear to me, since the files in the project have already been presegmented. However, if you don't take that extra step, you may end up with untranslated segments  that you are not aware of, but the client will be. You can presegment the files yourself using a demo version of Trados (under Windows, I'm afraid).
Click to enlarge
Nelson Laterman describes the process very clearly on his site, so there's no need for me to repeat it.

The problem with formats like TTX is, that there are no empty target segments. If there's a perfect match, you will see it in the target segment, if not, you will see the source text in the target segment. Very annoying, also because your memories won't make any suggestions.
I was going to mention this on my Wish List, but I made the mistake of mentioning the problem to Igor first. He instantly implemented a solution for it. For reference only, I will mention it on the Wish List anyway. Igor added the option to remove the target segments 
(see the first screenshot above). A great option, depending on the number of perfect matches in the file and of whether or not the client sent you a TM. In the latter case, you can use the option without having to worry about a thing. The perfect matches will show up, although you will have to insert any tags. If you don't have a TM at your disposal, you're in trouble. You will have to decide whether or not to use the new option. I was lucky. Most files were very short, so I went through them, and added the perfect matches to the memory before removing all target segments.

I'm afraid I didn't even try it, but since TTX files are project files, I don't think you can import them all in one go. I don't even think you import them, you "open" them. Please correct me if I'm wrong. When I started, I wrote down which TTX files I processed. I found out soon enough that that isn't necessary.
Click to enlarge
As you can see, the processed/translated files are shown as with different icon. 


Monday, 10 January 2011

JUMP!

One of the items in my Wish List was "jump to the next segment".  The Wish List is getting a bit messy because Igor complies with my wishes just about instantly, so I mention it here: DONE, that is, for numbers anyway. Not only can you jump to the next segment that isn't a number, in the process one of my other wishes is fulfilled: The comma/dot conversion. Great for Annual Reports and things!

You will need to change some settings, though. In the Project File, select Auto-Segmentation
Click to enlarge
Then go to Translation>Jump Over, and check Numbers
Click to enlarge
What a timesaver!

Tuesday, 4 January 2011

More on Spelling

So I finally managed to have my spelling checked on the go. But there is more to it. CafeTran not only shows you the word you misspelled, it also offers you alternatives. Select the misspelled word, and hit CTRL+ALT+Space, and if necessary, repeat the process by pressing CTRL+Space till you see the word you want. It will be inserted automatically.
Click to enlarge

For several languages, there is also a thesaurus available for OpenOffice. Hold the CTRL key while clicking with the right mouse button on a word. The list of meanings and their synonyms for the word will pop up to check or replace. The keyboard shortcut for this action is CTRL+SHIFT+Space.

Go to the OOo extensions website for an extensive list of dictionaries. 

TIP: The OpenOffice dictionaries will only work for OOo. You can't blame OOo, and they are great for working with CT. For as far as I know all native Mac applications and most third party apps, you can try Aspell, a free and Open Source spell checker. It's cross-platform, so it will also work under Linux and Windows, although to my knowledge, support for Windows was terminated ages ago. For the Mac, Aspell comes in handy, especially for "unusual" languages. The list of dictionaries boasts almost 100 supported languages, including Zula, Kinyarwanda, Tetum, and several other languages I never heard of.
Click to enlarge
Since this tip has little or nothing to do with CafeTran, I will not mention it in my TIPS entry.
OK then, one more thing (as they say) on spelling: On the Mac, you can even activate the spell checker in Skype, including Aspell. Just right-click in the text field, and choose Spelling and Grammar>Show Spelling and Grammar

Monday, 3 January 2011

CafeTran and Office Suites

In my entry Configuring a New Project, I mentioned that you can use one of the OpenSource Office Suites to convert a .doc file to a .docx or xml file. CafeTran is XLIFF based, so this conversion is necessary, and although other apps can do the trick, these free Office Suites offer you very useful other benefits: They allow you to review your translation (only file formats supported by OO, of course), and to check your spelling in CafeTran on the go.

It seems most people don't have problems `integrating' an Office Suite into CT, but it took me ages, and without the help of Igor, I probably couldn't have done it, or - more likely - would have given up trying. Finally, I managed to do it, but I still haven't the foggiest of what actually did the trick.

I installed NeoOffice first, but that didn't work. Then I tried the latest version of OpenOffice, 3.2.1, and got the error message that JRE (Java Runtime Environment) didn't work. So I downloaded and installed version 2.4.1, and that was worse: It prompted me to enter something unspecified in the Terminal.


Time for Igor.

Igor suggested to try OOo 3.2.1 again, and to set the path to OOo in CT. Go to Edit>Options>General. In the field next to the Editor button, choose OOWriter from the drop down list. Click the Editor button itself, and in the dialog box that pops up, click the Classpath button. In most cases, you then move to your Applications folder, and you click OpenOffice.
Click to enlarge

You should the see the following encouraging pop-up screen:
Click to enlarge

Click OK and save. In Options, also click OK, and restart CT. In the menu item Project, you should now see OpenOffice.
Click to enlarge

Encouraging, but that didn't mean that the spell checker was enabled. OOo is not a 64-bit application, so I was told to change the settings for Java to 32-bit. Go to Applications>Utilities>Java Preferences, and in the General tab, change the order so 32-bit is on top. Strangely enough, you cannot uncheck 64-bit, but changing the order is good enough.
Click to enlarge

Or so I hoped. Nope. Still no spell check in CT. Only when I restarted the Mac, everything was working as it was supposed to do:
Click to enlarge.

Wednesday, 15 December 2010

Auto-Completion

A particularly useful feature of CafeTran is Auto-Completion. Start typing a word, and CT suggests its completion, highlighting it.
Auto-Completion is based on the words in your TM(s), the Google and Bing databases, and the words you typed before.
You can accept the suggestion by hitting Return. If you are not happy with it, just continue typing, or press Esc or any of the Arrow keys.
You can change the settings by going to Edit>Options >Workflow.
Click to enlarge
You can even uncheck this timesaving functionality. Not recommended.

Saturday, 11 December 2010

The Wish List

It's the time of the year, I'm afraid, so I would like to present my Wish List. I'm still struggling with my InDesign file - as was to be expected. Although the problems I ran into seem to be inherent to the file format rather than caused by CafeTran, the job inspired me to suggest a few changes in CT I see as improvements. Please add your comments and suggestions for other improvements here, because it is not possible to do so at the Wish List page.

Tuesday, 7 December 2010

For Real (cont.)

Going for the real thing turned out to be a good idea. It taught me already a good many things:

  • The client wanted “proof” CafeTran could handle Indesign files, but I haven't finished the translation yet. Now in the project folder, there’s a file with the same name as the source file, but with the abbrevation of the target language in it. In my case: Katalog2011_52_104_nl.idml. So I sent the client that file, thinking it would do the trick. It didn’t. The file only contained the source language. It will only show the target language after exporting the project (but see the TIPS). Things turned out to be far easier than I thought. In CT, just go to Translation>Preview Document, and there’s your .idml file with (part of) the target text, again in the same folder.


  • “Add Current Segment to Memory” will take you to the next segment. But I don’t want to add all the numbers of a table to the Memory Database, so in that case I use Control+Option+Right Arrow to go to the next segment. That will take me there without adding the previous one to the database. However, that will not bring up the pre-translation of that next segment. So I use Control+Option+I to insert the next number, followed by Control+Option+Right Arrow again. Until I forgot to insert the number, and went to the next segment anyway. I realised I forgot to insert it, so I went back to correct it. Not necessary. CT inserted the source segment automatically. That saved me a lot of work.


  • I wanted to use superscript. I didn’t know how to do it, so I asked Igor. He claimed that only Word based CAT tools could do that. I checked with DV3, and he seems to be right. No idea how to solve that problem. It reads  “1.” “Tag” in the source text (two segments), and I want “1e dag” in Dutch. Temporary solution: I joined the segments, so I got “1. Tag” in the source language, and translated it with “Dag 1”. Acceptable, but not up to my standards.

Sunday, 5 December 2010

For Real

I was freewheeling nicely I thought, and all of a sudden, I found myself working in CafeTran. Not testing, working. For the first time in all those years, I got a job I couldn't possibly handle in DV3. The file format is .idxlm - InDesign - and the job is probably ideal for a CT debutant like me: Heaps of repetitions, lots of time to do the job.
I did some 2,000 words today, and that's not bad for a start. That doesn't mean there aren't any problems. A look at the picture below will tell you that there's a lot of nasty tables, and I have some problems with the inevitable tags that come with it.
For more on tags, see Igor's remark in the Tips section.

Thursday, 2 December 2010

Question Time

Everything seems to be working fine now, but I'm afraid I still have some questions on using databases:

Q1: If you include Google/Bing, everybody works with at least 2 databases. Can CafeTran assemble a translation by taking terms from these databases in a particular order? In my case, look for perfect matches first, then for matches in the Lex, then in the TDB?


A1: The assemble order is the following:
1. Perfect (100%) matches within the project itself.
2. Matches in the glossary (the Lexicon in your case).
3. Matches in the Terms memory which is joined to the main memory (menu Memory | Join Terms memory).
4. Matches in the main memory (the TDB) or other memories.
(Igor)

Q2: And in continuation of Q1, can CT do so automatically (AutoAssemble in DéjaVu)? In other words, you go to the next segment, and all the terms from the various databases are already there, assembled in the preferred order.


A2: It depends how you integrate you memories in the workflow.
The default Automatic mode does the assembling while you take the next segment. If you notice an unacceptable pause waiting for the results, you should switch to Pretranslation mode.
The Pretranslation mode performs the assembling of all segments in one run. You may continue taking next segments while the program is autotranslating in the background. The matches are available instantly. (Igor)


Q3: In the Project File, I set the path for both my MDB and the TDB to the folder in which they are located (.../cafetran/resources/memories/ENG>DUT), and I checked the boxes for "Memory For Segments" and "Memory For Terms." However, in the Project itself, I end up with one memory database only, ENG>DUT. Did CT combine the two? Merged them?

Click to enlarge


A3: All memories in this folder are merged resulting in one memory. If you do not wish to merge them so, you need to open them separately through the Memory menu. Then, they will open in its own tabs. Alternatively, you may join a Terms memory to the main memory to get their matches in one tab and still keep the memories separate (menu Memory | Join terms memory). (Igor)


Q4: Can I open and edit my databases in CT? If so, how, and if not, what (Mac) software should I use? Clicking to open a database now triggers AppleTrans, and learning one CAT tool at a time is already more than I can handle...


A4: Yes, you can. There are two ways to accomplish this:
1. Open a TMX memory like a regular project (menu Project | Open project). Then, edit and save its segments the same way you would review the project segments. This method has your system memory RAM limitation because the program loads all the segments to RAM.
2. The other way involves the true Database approach, which means creating a Database table and importing the memory segments there for edition and searching (menu Database). Here, you do not have any RAM memory limitation.  
(Igor)