[Please note that some of the entries are very old.*]
Showing posts with label CafeTran. Show all posts
Showing posts with label CafeTran. Show all posts

Friday, 10 July 2015

CafeTran Espresso 2015's Resources, an Overview

Historically, CAT tools only offered the possibility of re-using your translations. The document that needed to be translated, was split up into parts called “segments” (usually sentences), and the translated segments were stored in a memory, the Translation Memory (TM), which was referred to during the translation. In due course, more and more resources have been added. The following is an overview of the resources CafeTran Espresso 2015 currently offers. The number of resources you can connect to is unlimited, so you can have any number of TMs, Termbases, and web resources simultaneously.


  • The TM. The translation memory for segments. It’s still there, and it’s the only resource you really can’t do without. It’s a TMX mark-up file that offers far more features than the original TMs, including fuzzy matches and easy but extensive maintenance from within CafeTran.
  • The Termbase, a memory for words and phrases, very similar to the TM, with the same benefits.
  • The Glossary (also called Lexicon or tab del) for terms and phrases, a plain, tab delimited text file that’s past its expiry date. It doesn’t offer the goodies the Termbase offers, and it’s a left-over of the days you could only “connect” to one Termbase.
  • The Database. In CafeTran, a simple relational database (SQLite by default), that allows for very fast searching gigabytes of data (think of the DGT) because it’s indexed. It can also create a (TMX) file by “recalling” data from the database.
  • Web resources. You can search websites from within CafeTran, and the results are shown in the CafeTran interface.


All resources can “float” so you can give them a more prominent place on you screen, or even move them to a second monitor. Click to enlarge.

  • Machine Translation (MT), also integrated in the user interface. Results can be triggered by going from a segment to the next one, or manually. The free MyMemory is the default MT, Bing and other MTs are included in the Options.
  • TM-Town. A very recent development. I’m afraid I haven’t tried it yet. You search your own resources, but it offers extras like tokenisation.
You may also want to have a look here for my take on the settings of TMX resources.

Monday, 22 October 2012

Ubuntu and CafeTran

Just for kicks…


VirtualBox: 4.1.20
Ubuntu: 12.04

Click to enlarge

Tuesday, 3 July 2012

CafeTran July Update

It's just about impossible to keep up with Igor's updates, and usually, I don't. I mention them in "Latest" (top right of this page), and that's it.

The July update requires  some more attention, I think.

In earlier versions, CafeTran's spell-checker - Hunspell - required running OpenOffice, LibreOffice, or NeoOffice for "check spelling while typing". Those office suites opened automatically when you started CT. No problem. The only restriction was, that you couldn't increase the RAM assigned to Java to more than 2 GB (unless you run the suite under a Linux distro, the only version in 64 bit so far). So in the case of very large TMs, you had to choose: Either use the spell-checker and limit the RAM to 2 GB, or increase the the RAM an do without spell checking while you type. This latest update changes that. Spell-checking now works independently from an office suite, so you can have your cake and eat it too.
Please pay attention to the following:
The path to install the dictionaries is a bit longer than you may think it is. You should install in the "resources" rather than the "Resources" folder.

Click to enlarge

Please check here if your dictionary is available. The list is quite long, and you would think most languages are covered, but Turkish turned out to be left out. Very strange, it's not exactly a minority language (see UPDATE below).

Other improvements:
  • Set priorities for multiple TMs, both for TMX files and tab delimited glossaries.
  • Save tags in the TM. Extremely useful. The tags in the TM won't interfere with matches in case the source segment does not contain tags. Wonderful.
  • Typing a tag number after a number. Minor problem solved.
  • Go to the last edited segment in a TTX project. Solved.
  • Add fields (like subject, client, etc.) to glossaries made possible (useless, I think, but if you are that way inclined, go ahead).
  • The language codes in TMs are no longer case sensitive. This makes life a lot easier when you import TMs.
  • Autocompletion (AutoWrite) hints for two words.


"Every disadvantage has its advantage", Johan Cruijff used to say. This renders quite a few blog entries completely and utterly useless (thank you, Igor…).

UPDATE: Igor already solved the "Turkish problem".
After renaming both files from tr to tr_TR, now spell check works, says Selcuk (and I'm strongly inclined to believe him). Turkish users can download Hunspell dictionaries here.



Wednesday, 4 April 2012

Language Codes

After installing the new build of CafeTran a month or so ago, I noticed that importing a TM (in TMX format) in a project was rather slow. And what was worse, lots of source language entries just didn't show up.
I suspected a problem with the language codes, and Igor confirmed it. My main language pair is English - Dutch, and opening the TM in a text editor showed that the source language code occurring in the file was both American English (EN-US) and British English (EN-GB). I used TextWrangler to Find/Replace all instances of EN-GB into EN-US, but any text editor should be able to do the job.

Click to enlarge
This time, all source and target entries showed up in CT, but loading the memory into the project was still slow.
The answer to a question of another user in the CafeTran support forum solved the problem: The code consists of two parts, the language code (English in this case) and the region code (US). And it is case sensitive. Again I had to use TextWrangler, this time enabling the Case Sensitive box. And I repeated this action for nl-NL - my target language - and my other TMs.

Click to enlarge
Problem solved.

Tuesday, 3 January 2012

Espresso 2012 UPDATE

Espresso 2012 is only a week young or so, and there is already an update. No, no bug fixes, a genuine update.


Says Igor:


Dear Cafetranslators,

I have uploaded the new build of CafeTran. The update concerns the logic of automatic insertion of translation into the target segment.
The current logic is in the following order:

1. Exact match.
2. Fuzzy match based on the Fuzzy much percentage level set in Options, and if not present:
3. Autotranslation (also called Autoassembly) based on its homogeneous accuracy level set in Options.

The trial version is available for download at http://www.cafetran.com/download.html.
Notice that CafeTran has a new address http://www.cafetran.com now.
The previous address will point to this site.

I will start sending this version to CT users this week. The current version is: CafeTran Espresso 2012010202.

Friday, 23 December 2011

Espresso 2012

Igor lets us know:


"It is my pleasure to announce officially the new version of CafeTran called CafeTran Espresso 2012. It features a lot changes in the user interface to make it much more intuitive and user friendly.


There are also numerous enhancements to existing functions as well as new important features such as optional prefix matching (stemming), QA, filtering and
sorting segments, implementation of the new Google Translate API v. 2.

The current manual is still outdated but YouTube video tutorials on various workflows will follow next. I hope the first one will appear by the end of this year."

Monday, 14 November 2011

New Project Info User Interface

CafeTran now has a new Project Info UI. It should make it much easier to set up a new project, especially for beginners.

Click to enlarge

Wednesday, 19 October 2011

New CafeTran Blog?

To my surprise, I stumbled upon another CafeTran blog. And a very good looking blog at that. At first, I thought Igor wrote it, but I doubt if he would write a very Dutch URL like:
Click to enlarge
I now suspect another Hans, a well-known Dutch translator who usually goes by the name Hans List.


Another blog would definitely be a good idea. I haven't been able to add any entries for weeks. That's not out of laziness, it's simply because I was busy using CT, and didn't have any issues. I almost exclusively worked on TTX files, and there's no reason to write about that format again. I admit that I also did a few TXT and HTML files, but they didn't present any problems. CT just works...

Thursday, 28 July 2011

Connect My Mac

I admit that it's all a bit experimental at this stage, but I can connect from my iPad to CafeTran on my Mac using, er, Connect My Mac. The basic app is a free download from the App Store, developed by Hana Mobile. Time will learn how useful this is, but the idea of being able to manipulate your Mac on your iPad is fascinating.

Click to enlarge

Thursday, 7 July 2011

Dictionaries and Dictionaries

"CafeTran offers a flexible interface to access and update your dictionaries in the workflow", it says in the CafeTran Handbook. It continues to distinguish between dictionaries and glossaries, and how to integrate a dictionary into CT.

All very well, but when I tried to integrate my main dictionary like I integrated an Internet resource, it didn't work. It turns out that the term "dictionary" applies to wordlists, very similar to glossaries, but not to software dictionaries. Apparently, there are dictionaries and dictionaries.

Mac OS comes to the rescue. With Services, you can easily link your dictionaries to CT. In the menu bar, go to CafeTran>Services>Services Preferences...
Click to enlarge
Choose or add your dictionary. In my case, that's Van Dale. Assign a keyboard shortcut to the dictionary, Ctrl+v, in this case.
Click to enlarge
Next time you want to look up a word in CT, just double-click it, hit the shortcut, and magically and revolutionarily, your preferred dictionary shows up with the word you selected.
Click to enlarge
Pretty neat.

Tuesday, 5 July 2011

Proofreading in CafeTran (cont.)

I was aware that I could export/save the project as HTML, but that seemed to me to be pretty useless format for proofreading.
Click to enlarge
It is, until you see the light, and open the HTML file in a word processor.
Click to enlarge
Get rid of the column with the source text (and convert Table to Text, if you like), and you will end up with this:
Click to enlarge
This is a format suitable for reviewing, especially if you print it. Be sure to make changes in the Project so your translation memory is up-to-date.

Wednesday, 29 June 2011

New Build

Igor released a new CafeTran build. New features:

  • Gluing segments of many documents
    CafeTran allows you to work with many documents in one project by selecting the folder that contains them when creating a new project. They are available for translation the menu Project>Documents ... . During the translation, you can switch between them or glue them in one view for convenience. Of course, the segments must be first created so it is best to choose  Automatic segmentation when setting up a new project with several documents.
  • Filtering and sorting segments
    A very handy feature that allows you to hook into the workflow only selected segments. Filtering can be done by a word, phrase, note, checked or unchecked segments. In addition, the program lets you sort segments alphabetically. The described features are available in the Find menu.
  • Saving project segments to the HTML fileMenu Project>Save as..., next choose File of type - html. This is useful for displaying segments in the browser, e.g. to review them outside the program or open in the table of Ms Word or OpenOffice.
  • Highlighting spelling mistakes in red instead of underlining in black them when the program is connected to the OpenOffice spellchecker.

  • Bug fix
    Skip numbers now doesn't skip whole segments that start with a number.

    Friday, 6 May 2011

    Internet Resources

    In my previous entry, I mentioned youalign. The developers of that wonderful tool asked me to check out their new Internet resource, webitext. I did, of course, and then I thought of trying to integrate it into CafeTran, one of the many CT features I haven't even looked at.

    It's easy. In the workflow UI, you go to Library>New Resource and choose Internet>OK. You will see the following dialogue box.
    Click to enlarge
    I'm afraid I only entered the webitext URL because I had (and have) no clue about Address Start and Address End. Never mind, click OK and go back to the workflow UI. To activate webitext, go to Library>Internet>webitext.


    Click to enlarge
    And there you are. Some things don't work as I would have expected, and it's a bit slow, but then again, this is my very first try.
    Click to enlarge

    UPDATE: Please do enter the Address Start and End fields (see comments below). If you do, the whole lot works faster, plus you can use the CT Search Field to enter terms (by double clicking or selecting them in the source pane) to get results from webitext.

    Wednesday, 4 May 2011

    Aligning

    When I was introduced to CAT tools about 14 years ago, I thought the alignment feature would be great. Imagine, all your pre-CAT tool translations instantly retrievable, ready to use for your current project! I was wrong. I tried the alignment process once, and gave up after a page or so. Time consuming, boring, and especially for general subjects, not worthwhile the effort.
    However, to prepare for an upcoming project, I changed my mind. The reference files mentioned to me by the customer contain lots of very specific information, a conditio sine qua non for the project. So I gave it another try. In CafeTran.
    The alignment workflow is pretty straightforward and - for once - well described in the Handbook. Just don't forget to check the Documents Alignment box (as I did the first time).
    Click to enlarge
    So I followed the instructions, started aligning, and... gave up after a few segments. L'histoire se répète. I simply can't do it. I'll have to live with it. And then I remembered there was a free alignment service on the Web. I typed "align" in the address line of Safari, and immediately got the URL of that service. I was smart enough at the time to bookmark it, not smart enough now to remember I did.
    To use youalign, you'll have to register. After confirming your registration, you can upload the documents you want aligned (with a few restrictions on format and size), you wait a minute or so, and you'll see a preview and the link to download the alignment file as a TMX file. 
    Click to enlarge
    I checked the result in CT
    Click to enlarge
    and it was wonderful. Better still, it was perfect. I wonder how they do it. Pretty smart algorithms they must use.
    I uploaded the other reference files, and the final step was merging them into one TMX file by saving the individual files in one folder, opening the folder in the CT Project File, and then saving the resulting memory as one TMX.

    I feel a bit sorry for Igor. Aligning in CT is easy, and the explanation of the workflow in the Handbook is excellent. But you can't beat youalign...

    Sunday, 17 April 2011

    Proofreading in CafeTran

    Usually, I proofread my translations in the final document. The good thing about that, is that you see the text in its final version, the bad thing is that you either have to make the changes in your CAT tool - CafeTran - or end up with databases that don't reflect the final version. In other words, your mistakes will probably show up again in a next translation.
    In this case, I had to do the proofing in CT because the final document didn't show all the text. Lots of it was hidden, a known MS Word .docx issue: All .xlm files are equal but MS doesn't stick to the rules, as usual.

    So I proofed in CT for the first time. I thought it was a good idea to check the translation in the left-hand panel, rather than in the usual translation pane. So I expanded that part of the UI.
    Click to enlarge
    If you come across a segment that needs changes, just click on the blue segment number, and you go to the relevant segment and you can make the changes in the translation pane. What surprised me, was that CT apparently automatically propagates the changes. Wonderful!
    What I didn't realise, is that the spellchecker doesn't work in the lefthand panel. Not so wonderful, but understandable.
    The ideal solution would be to proof in the Preview so you could see all of the context, pictures, diagrams, and so on, whereas the changes made in the Preview are propagated in the Project File and the memory or memories. It may even work that way, but I didn't have the time to give it a try.

    On Rosetta, somebody expressed the wish that his CAT tool would offer the possibility to mouse-click terms to re-arrange them. According to him, an extremely useful feature for Japanese/Chinese and similarly constructed languages where Auto-Assemble doesn't really work because of completely different syntactical structures. CT can do it. Even better, if most of your terms are in the database, just double-click them so they appear in the right order. I made a screencast of it, but don't get too excited, it's my first attempt. Next time, I'll try QuickTime instead of Jing. Both methods are free, by the way.

    Saturday, 5 March 2011

    The CafeTran UI: The Project File Simplified (3)

    Below, I have tried to summarise the minimum settings required to configure a translation project for four situations. The following ignores most options (including useful ones), reviewing, billing, etc. The sole purpose is to get you started. You can always adjust your settings in the Workflow UI: Project>Project Info.

    I. Configure a new translation project without existing databases

    This is typically the situation for translators who are new to CAT tools, or who don't have memory or terminology databases at their disposal.
    Click to enlarge
    • Click the Document (1) button and browse to the file you want to translate
    • Select or check your source and target languages (3-4)
    • Select or check the file extension (5) of the file you imported in 1
    • Check the New Project Memory box (8)
    • Click OK

    Done. All other settings are either not important at this stage, or their default values will suit your purpose for the moment.

    II. Configure a new translation project making use of your existing databases

    This is typically the situation for translators who are familiar with CAT tools and/or have database resources of their own, or who received databases from the client.

    • Click the Document (1) button and browse to the file you want to translate
    • Select or check your source and target languages (3-4)
    • Select or check the file extension (5) of the file you imported in 1
    • Check the Load Database Memory (9) box if you have a terminology database at your disposal (a Big Papa)
    • Check the Load Memory File (10) box, and locate your memory database (Big Mama) after clicking the TMX button
    • Click OK


    III. Resume a translation project

    Click the button for your current project (NEW!, a new CT feature...), go to Recent Projects in the Menu Bar and select one of your recent projects, or click Project (13), and browse to it. Click OK.
    Click to enlarge


    IV. Start a CT supported translation project

    CT supports a number of projects, including projects set up in other CAT tools. They typically have the extensions .xlf, .ttx, .tmx, etc. Click Project (13), and browse to the project of your bread and butter. Click OK.

    Friday, 4 March 2011

    The CafeTran UI: The Project File 2

    Next to the Document Settings (2-7), you'll find two more tabs:


    The Memories Settings Tab
    Click to enlarge
    18. Workflow integration. Manual integration - searching the memory for the matches of the current segment takes place only when click the Translate button in the target window toolbar. Automatic integration (default) - searching the memory for the matches of the current segment take place automatically as you take the next segment. The searching may take some time when you work with very large memories. Pretranslation - the memory is being searched for all document's segments in one step in the background. You can translate all your segments during this process and the results are shown instantly as you take a segment. This method is preferable while working with large (big mamas and papas) memories.
    19. Matching type. Fuzzy matches. The typical CAT memory matching for the current segment with shown percentage. Subsegment matches. These are 100% matches for the subsegments of the current segment as well as the approximate (statistical) matches for the subsegments of the current segments. Detailed matching. This is character based matching rather than word based matching (for languages where there are no word separator such as a space). By selecting here you decide what kind of matching you want to have during translation. The default setting allows for both Fuzzy matching and Subsegment matching taking place.
    20. Character encoding. UTF 8 is default. However, Trados tmx memories use UTF 16 encoding. So if you plan using the TMX in Trados you can change the default setting. I heard that Trados also recognize UTF 8 encoding correctly.
    21. Memory for segments. Permits adding segments to the memory.
    22. Read only. The memory can be queried but cannot be saved. This option is also very important for tuning the memory performance. The memory uses less RAM, and matching is faster. So it should be used with large memories set for reference only.
    23. Memory for terms. Permits adding terms to the memory.


    The Segment Properties Tab
    Click to enlarge
    24. Project. Enter the name of the project the segments come from. Based on this property, you can later filter the memory to extract only the segments with the given property.
    25. Field. Enter the subject matter of the project (e.g. Medicine, IT, Legal, etc.) you add to the memory.
    26 and 27. Define your own property names.


    The Menu Bar
    On the very top of the UI, you'll find the Menu Bar as usual. I should probably have started with it, so let's call it a proteron husteron (which in itself is a husteron proteron, of course).
    Click to enlarge
    Recent Projects speaks for itself. Header, well..., I hope Igor can help me out on that one. 

    Sunday, 27 February 2011

    The CafeTran UI: The Project File 1

    [This entry replaces two older ones. You can still find them here, and here]
    Igor has simplified the Project File UI. I will discuss it in three parts, starting with the Document Settings.
    To start or set up a project, run Start.jar (and not Cafetran.jar). You will see the following interface:


    Click to enlarge

    1. The Document button. Click this button to browse to import a document or a folder with documents into CT. The Document button is very prominent in this UI, but you probably use it only once: When you configure your project. You can use it to add documents later, but you can also do that in the workflow UI. You shouldn't use it to resume a project, unless you want to add documents. Instead, to resume your translation without adding documents, just hit Open Project... (25).
    2. The bar with the Document Settings, Memory Settings, and Segment Properties.
    3. Set your Source Language. No comment needed, too obvious.
    4. Set your Target Language.
    5. File Type. Choose the type of file you want to import. Usually, the file type field adjusts itself to the format of the document chosen in 1. CT offers you a multitude of file formats - which of course is great - but as far as I know, you cannot translate multiple formats in one project. Apart from that, if you check the available formats, it may come as a shock that there's no ttx, sdlxliff or other (Trados) file format available. Now I'm not much of a Trados fan (to put it mildly), but I have to admit that translating those files contribute a lot to my miserable income. Luckily, CT can handle those files, but it treats them - correctly - as project files, not as documents. So instead of entering all the information in 1 - 3, you just click Open Project... (25), and browse location of the project concerned. You can find more on how to handle these project files here.
    6. Tool-id. You can choose between CafeTran and CafeTran-OpenOffice. Beats me. No doubt, Igor can explain this, and the following:
    7. Phase. Choose between Translation, Review, and End.
    8. New Project Memory. Check this box only if you do not use a general memory database (Big Mama), or if you do use a Big Mama, but want to create another database specifically for this project.
    9. Load Database Memory. Check this box if you want to use an existing (general) terminology database (Big Papa).
    10. Load Memory File. Check this box if you want to use an existing memory database (usually your Big Mama).
    11. The path to the memory database mentioned in above.
    12. Billing. Let's worry about that later. Translate first, bill later. If only the other way around were possible...
    13. Project. Click this button to open an existing project, including projects of other CAT tools (ttx, sdlxliff, etc.)
    14. Translation Mode. Segmentation mode - the text is segmented "as you go forward". You create a segment and translate this segment at the same step, then next segment and and so on. Autosegmentation mode - the whole text is segmented first, and only then you start translating. That is how most CAT tools work, I suppose (Igor says). Clipboard mode - a special translation process which automates as much as possible importing from, translating and sending back the segments straight to external applications. The segments import and export takes place via system clipboard. There is a chapter in the Handbook explaining this mode in detail. Image mode - I also wanted CafeTran to be a sort of editor for translation of scanned images or paper documents such as short certificates, forms, etc. Here, you only create target language segments in the target window (without having any source segments) and export them all to the text format. You can open the scanned image in CT interface and query manually your memories and other resources. Of course, you can also OCR and convert the image documents to a word processor as an  alternative.
    15. Translation in Review. Checking this box allows you to make automatic searches in memories even when you already have the target segments created, for example, when you review the text.
    16. Documents Alignment. Check this box to align two documents (source and target) in order to put their corresponding segments to memory or extract phrases to a terminology base.
    17. Recent Resources. Check this box to use recently used databases.

    Wednesday, 16 February 2011

    The CafeTran UI: The Project File (intermezzo)

    Igor has done it again: He simplified the Project interface after reading my previous entry on this blog.
    Click to enlarge
    However, he's cheating a bit - a few bytes even - because all 28 items are still there, hidden behind tabs (highlighted in yellow). This UI is far less `frightening' than the old one, so it's an improvement anyway.

    In my wish list, I suggested to skip this interface all-together, and go straight to the workflow UI. Igor does not agree, and argues this `sounds good only if you want to continue the project.'

    My arguments in favour of skipping the Project UI:

    You can find the very same Project UI in the workflow UI under Project>Project Info. So why not take the `risk'? If you don't continue a project, set up a new one in the workflow UI.

    Igor argues that it would take more steps: `first you would need to close all the memories, glossaries from the previous project and next go to the Project UI to configure them anew, which might be totally different than in the last project.' But is that true? I mean, `totally different'? I asked a question on Rosetta: How many language pairs do you work in, and how often do you switch between them? I don't have the results of this poll yet, but I bet the number of pairs is very limited. And if you work with Big Mamas and Papas for each pair, switching will not take much time. No time at all, if Igor manages to link the Mamas and the Papas (all the leaves are brown) to the various language combinations. That would only leave the project specific glossaries to select, and you'll have to do that anyway.

    IMPORTANT NOTE for MAC users: For the time being, close CafeTran via the Project menu instead of using the keyboard shortcut. Igor will fix this soon. SOLVED