Showing posts with label Google Translate. Show all posts
Showing posts with label Google Translate. Show all posts

Sunday, May 29, 2011

Google Translate API Deprecation Causes Commotion

In a post in the Google Code blog, among news of new APIs and other updates, Adam Feldman (APIs Product Manager) announced that the Google Translate API would be shut down according to its deprecation policy.

By clicking to the API page link we learn that "The Google Translate API has been officially deprecated as of May 26, 2011. Due to the substantial economic burden caused by extensive abuse, the number of requests you may make per day will be limited and the API will be shut off completely on December 1, 2011. For website translations, we encourage you to use the Google Translate Element."

The first reactions from the developer community were negative, as the tone and quantity of comments to the announcement indicate. On the language camp, the reactions fell in two groups: "I told you so"  and "Don't be evil, my eye!" (from the people that were skeptical about Google's good intentions of honest decision-making that disassociates the company from any and all cheating.)

I reached out to my contacts at Google to try to get an official position, but they declined to comment.

First, let's make it clear that Google Translate is not going away! The announcement is only about the API, and will affect programs that have incorporated it, like Trados, Wordfast, and DéjàVu, plus hundreds of smartphone apps that were developed on this platform. I will particularly miss the My-Translator plugin for Firefox.

What does this announcement mean to the language industry?
  • MT price will go up. The value of MT solutions like AsiaOnline and Systran will go up as developers will not have access to the free solution provided by Google (unless they resort to web scraping.)
  • Migration to Bing. Microsoft's MT solution doesn't cover as many languages and is not as good in as many domains as Google Translate, but it does the basic job well, specially for IT-related content. 
  • Google Translator Toolkit continues to be a good alternative to use translation memories in combination with MT. My guess is that the functionality of this tool will continue to improve, since this is the environment Google uses to localize its own applications.
  • Naggers will be empowered. The traditional arguments about confidentiality issues, quality of translation, misuse, working for free for a commercial entity will remain unchanged in the language industry. Now, the argument that Google can't be trusted will become part of the portfolio of reasons not to use Google Translate. 
I feel bad particularly for non-profit and practical integrations of the API that will be lost. I think that Google could just set up a price for the API to solve the problem of "abuse," even though I have a feeling that this is just a lame excuse.

As for me, I will continue to use it to read texts in languages that I don't understand.

Tuesday, May 03, 2011

Duolingo: Crowdsourcing at its Best for the Translation Industry

For the last month I have been reading tweets and notes about Duolingo as the place where "you learn a language and simultaneously translate the Web," but I kept postponing getting more information about it. As a good procrastinator, I figured that if this was really important, it would eventually make its way to me. Well... it did!

Ultan Ó Broin mentioned my name in his Blogos entry "The Future of Web Translation: Haters Gonna Hate" and I felt compelled to watch the video by Luis von Ahn, the inventor of Recaptcha, at a recent TEDx event at Carnegie Mellon University.

Duolingo does for translation what Flickr did for photography and what Wikipedia did for encyclopedias. It brings the knowledge of amateurs to do some work that only professionals could do. The advantage of Duolingo is that − unlike Facebook or Hootsuite, who also use community translation − the user learns a language in the process.

Crowdsourcing − the approach used by Duolingo − is the act of outsourcing tasks, traditionally performed by an employee or contractor, to an undefined, large group of people or community (a "crowd"), through an open call. The goal for Duolingo is to get 100 million people to translate the web into every major language for free.

The trade-off here, according to Luis, is that there are 1.2 billion people in the world learning a second language and they have to pay for it. With Duolingo, they will learn a language for free and translate the web in return. A really revolutionary and innovative concept.


According to Luis, using this approach, Wikipedia could be translated into Spanish in five weeks with 100,000 people or in 80 hours with one million individuals.

Who does this approach benefit? Everybody.

Who does it hurt?
  • Insecure translators who like to complain about things they can't control.
  • Rosetta Stone, Livemocha, Fluenz and other software-based language learning software.
  • Machine translation providers like AsiaOnline, PROMT and SDL, because Duolingo could be a faster/better solution.
Duolingo is a welcome addition to the arsenal of language solutions around the world. It is clearly a solution for making knowledge and information that would never be professionally translated available, especially in languages where the translator pool is insufficient for the amount of content that is available for translation. Watch out Google Translate!

Sunday, February 13, 2011

Google Translate App for iPhone Censors Swear Words

WARNING: This post contains graphic language. Don't read it if you are too sensitive.

Since I bought an iPad for our home, my five- and eight-year old children have learned every imaginable swear word on YouTube by watching Lego videos and Justin Bieber parodies. After the initial surge in interest and the inevitable uncomfortable situations in public, they seem to have gotten over the potty mouth phase.

Last week, I installed the new Google Translate app on my iPhone and was impressed by the voice recognition in 15 languages and the accuracy of the translations. I actually thought that it would have come handy in a couple of situations in Korea and China, where there were no foreign language speakers around.

This weekend, one of my brothers from Brazil came to visit and we started to play with the Google Translate app. Just like when we were kids moving to a new country, the first phrases we used to test the functionality of our new toy were swear words. And this is where the big surprise came: Google Translate doesn't print swear words in English. But ONLY in English.

As you can see in the screen capture, every time we used the word fuck, it was replaced by ####.

Out of curiosity, and using a very scientific approach to the process, we went about testing offensive sentences expressed in other languages and translated into English. Interestingly, the app had no problem translating and printing the word fuck from other languages into English, as you can see from the screen captures below.

This makes me wonder what is behind this policy and who makes the decision to enforce it.
  • App Store policy against offensive language? If so, does this also happen in the Android version of the app?
  • Hypersensitivity of Americans to four-letter words? If so, does this also happen in the UK or Australian versions of the app?
  • Why does the "censorship" only applies when you speak the swear words? In fact, you can actually type them and the text will be translated fine.
In any case, this is just a funny thing. The tool itself is excellent and very practical. I have been able to dictate relatively long sentences with acceptable accuracy, especially for a free tool. No wonder it is already the number one download in the App Store.

If you have an iPhone, download the app today and test it with your language pairs.

The screen captures below were the result of speaking a sentence in a foreign language and having it translated into English using the voice recognition feature.






Thursday, January 13, 2011

Google Introduces New Type of Telephone Interpretation

As announced yesterday on Google's blog, next month Google will launch what it is calling the Conversation Mode in Google Translate for Android. You can see a preview of it here.

It is a basic process of Voice Recognition, followed by Machine Translation that is converted back to voice using text-to-speech. The service will start to be offered in February 2011 between English and Spanish, but other languages will follow soon.

Google alerts that this is still an experimental feature that is in its early stages and that it cannot handle accents, background noise or rapid speech.

Is this the so awaited Universal Translator that we saw in Star Trek? Will this replace telephone interpretation or even human interpretation?

Not yet. In fact, I have seen demos of voice-based MT systems several times. Language companies used it as a technique to impress investors and get some venture capital. One of the first ones I was from Lernout & Hauspie that translated between English and Chinese. More recently, I was very impressed by how Speaklike was able to create a functioning demo just using off-the-shelf or free software.

Just like Google Translate, the Conversation Mode will help in situations where an interpreter would never be called before, like the shoe store case presented in the preview mentioned above. The applications are limited and the accuracy is not consistent. And just like Google Translate, the Conversation Mode will probably help increase the awareness of the importance of professional interpretation. Or would you go to court in foreign country using your Android phone as your translator?

Sunday, December 12, 2010

What I expect to see in 2011

This is the time of the year when people start making predictions for the next year. Well, as I have already been asked several times what I see in my crystal ball, let me share it with you.
  • Content. Let me start with a quote from futurist Ray Kurzweill in a recent interview for Time Magazine: "Our intuition about the future is linear. But the reality of information technology is exponential, and that makes a profound difference. If I take 30 steps linearly, I get to 30. If I take 30 steps exponentially, I get to a billion."

    So content is growing exponentially and that's not news, but for the language industry there will be two trends that will accelerate in 2011. First is the atomization or chunking of content, i.e., translation projects will come in smaller sizes (in line with the trend in the software industry to move to apps). Second is velocity of content, i.e. clients will want these projects faster. These two trends will drive increased demand for productivity gains.
  • Voice. I believe that there is going to be an increase in demand for voice translation. Not only on-site and over-the-phone interpretation, but also dubbing and subtitling. Everybody talks about the ascendance of video, but video means very little for the translation industry; what needs to be translated is what people say, hence the increase in video will lead to an increase in the demand for voice-based translations. (Note to translators: Learn interpretation skills).
  • Languages. Be prepared for increased demand for Indonesian (Indonesia is right after the U.S. in numbers of Facebook users), Vietnamese, and African languages. I also expect increased demand for Brazilian Portuguese as the predictions for growth in the Brazilian economy are very positive.
  • Business. Acquisitions will happen. Expect several announcements and some consolidation at the top. The main discussion will be once again the fair valuation of companies. Naturally Welocalize will lead the charge, but I expect to hear from SDL, Moravia, CLS, HiSoft, and the Scandinavian companies like Semantix, AAC, and LanguageWire. Either as aquirers or targets of acquisition.
  • Pricing. It is true. Price pressure is really a fact now. Mature clients are shopping around for better prices in order to translate more with the same budget. For many years I have said that prices had been stable in the industry, but I believe that in 2011 companies will succumb to the haggling of the big buyers. The only way out of this is to dramatically increase productivity using technology at levels never seen before. This will be especially important for Single Language Vendors. Freelance translators should think about measuring their income per hour or per month, instead of their price per word.
  • The year of interoperability in the cloud. All this talk about privacy and how Google Translate breaches confidentiality clauses will disappear. Translation memories will be shared in the cloud and the chatter of the last two years will become just that; chatter. The big winners in technology will be  the MT solution providers and Kilgray, with its MemoQ technology (that works very well with files generated by their competition and thus achieves de facto interoperability). It is not surprise to me that MemoQ only has raving fans. Asia Online stands a good chance of growing a lot this year as the last stalwart of independent MT. I predict SDL will still grow out of pure momentum, not because of its "innovative" solutions.
What I don't expect to be news in 2011, even though there is going to be a lot of talk about it still, is the adoption of Machine Translation and the impact of Social Media as a source of more translation and localization.

In my opinion, MT crossed the chasm in 2010, and Social Media content is generated almost exclusively in local languages, with very little impact on the demand for translation and localization. Social Media might be a driver, but not demand generator in itself. However, I wouldn't be surprised if a few startups come up with the idea of creating companies focused on localizing Facebook pages and Twitter feeds.

Now I need to catch a plane....

Wednesday, December 01, 2010

Industry Reacts Negatively to EPO and Google Deal. But Should They?

The EPO (European Patent Office) announced today that it signed a Memorandum of Understanding with Google Translate to automatically translate patents into the languages of the 38 countries that it serves.

The immediate response in the social media forums was quite negative, as most reactions involving Google and machine translation of late. But looking closely at the press-release, we see that "the collaboration aims to offer faster and cheaper fit-for-purpose translations of patents for companies, inventors and scientists in Europe."(our emphasis)

What this means is that the goal is to provide good enough translations, not perfect translations. Further in the release, the EPO states that "the partnership with Google to create machine translation tools for patents will help inventors, engineers and R+D teams to retrieve relevant documents efficiently - in their own language - from our wealth of published patent information." This means that the purpose is to allow people to search the EPO database for patents that have been already published, something that falls exactly into Google's expertise: Searchability.And probably something that those professionals already do on their own by copying and pasting information into Google Translate.

As discussed in a previous post, the EU is trying to promote the adoption of a single EU patent system, which is facing some resistance in the European Court of Justice because the court's Advocate General believes that a centralized patent is "incompatible with the treaties" that created the EU.

For Google, this is a bonanza that will provide them with a vast database of quality translations of approximately 1.5 million documents, a number that grows by more than 50,000 new patent grants every year. This means that the quality of Google Translate should improve in several scientific knowledge domains. By definition, patents facilitate and encourage disclosure of innovations into the public domain for the common good. Protection of inventions is achieved by making the information public and not secret. To me, this is a perfect match with Google's stated mission to organize the world’s information and make it universally accessible and useful.

Is this the end of patent translations for LSPs? I don't think so.

Patents are serious business and are worth a lot of money for their holders. Pharmaceutical patents are easily worth billions of dollars over their terms of protection and many lawsuits have been filed because of the interpretation of specific terms. Large organization that file hundreds or thousands of patents every year will not be penny pinching on translations and will still prefer to use the services of specialty LSPs like RWS for filing purposes.

Add to this, the fact that there are other language pairs that are not covered by the EPO/Google deal, and that patents still need to be filed in Japan, China, Brazil, and other countries. Deals like the one announced earlier this year between Asia Online and Lexis-Nexis Univentio can still proliferate, since their goal is to achieve publication quality.

From my point of view, this is a good announcement that proves the growing maturity of Google Translate, which will become a better tool for all its users due to the addition of good content to its database. But Google Translate will continue to be generic tool with good enough results.

Thursday, July 15, 2010

SDL Acquires Language Weaver. First Reactions.

Other experts like Kirti Vashee and my former colleagues from Common Sense Advisory will certainly post more detailed analyses of this news, but I wanted to document my initial reactions to what was announced today by SDL and LanguageWeaver, the first developer of a commercial Statistical Machine Translation software.

In recent months, I have been thanking SDL for the great job that they are doing at alienating their technology customers by providing sub-par customer service and support. Clients contact us at Milengo looking for alternative solutions, which we are happy to recommend. SDL has been very successful at irritating translators, LSPs, and final buyers with their technology approach.

LanguageWeaver, a pioneer in SMT for commercial purposes, has struggled to sell a product profitably, when it has to compete with free solutions like Google Translate and MOSES. It's main client is the U.S. government and the main language pair is Arabic-English. In fact, the announcement points out that in 2009, the company had a loss of $1 million for revenues of $12.2 million.

So why is SDL paying $42.5 million (or 3.5 times revenue) for a company that loses money?

I believe that -- whether it works or not, and whether it is deployed or not -- acquiring a software company is something that investors at the London Stock Exchange put a very high value on. This is a good story that will boost SDL's stock, just as the IBM Websphere MT deal boosted Lionbridge's stock to the levels that it is today (from one dollar to $5.28). This is a good story that helps SDL to further position itself as a software company instead of a service company.

The second benefit for SDL, is opening a door into the U.S. government R&D funds through DARPA. LanguageWeaver has advanced mostly because of the availability of such funds.

I don't see the technology itself as a major game changer for SDL. SDL had already acquired Transparent Language, a Rules-based Machine Translation developer, and not much has been heard about that technology since. After a little time, LanguageWeaver might take the same route as Idiom's Worldserver, which was growing fast and was virtually discontinued by SDL.

If the patterns of previous acquisitions prevail, SDL will get very excited with LanguageWeaver, but after the excitement wears off, the product will be abandoned to its own fate. So LanguageWeaver clients who already work with Trados, TMS, and other of their products already know the level of service provided by SDL, and should maybe run for the hills when they come offering LanguageWeaver solutions.

Finally, for competitors -- unless SDL gets its act together -- they have nothing to fear. Just keep providing excellent customer service. That's what Milengo does.

Sunday, March 14, 2010

Machine Translation in the News Again

Google Translate and Google Translator Toolkit made the rounds of the big U.S. media last week, with major stories in The New York Times and the Los Angeles Times. These were tweeted, retweeted, facebooked, LinkedIned, and forwarded by e-mail ad nauseam. How are people reacting? I identify two major groups:
  • It's the end of the world for translators. This doomsday approach stems from fear of the unknown and amazement with the quality of the translation that Google has been generating for some language pairs. 
  • MT is never going to reach perfection. So, no worries. This nonchalant attitude comes from those who only see the defects in the tools and feel safe in their current positions.
Nem tanto ao mar, nem tanto à terra, is a Portuguese expression (don't try to google-translate it, it's not going to work) that literally translates as "not so much to the sea, not so much to the land" but means that the neither extreme is right and the truth is probably in the middle. If you follow my postings or presentations, you should know by now that I believe that translators should use MT to improve their productivity and it is only useful if the user knows the language into which the text is being translated. 

I agree with Ben Sargent from Common Sense Advisory, when he says in the Global Watchtower that "...machine translation could remove the cloak of invisibility from translators, giving them greater recognition and status. As 99.99 percent of translation is done by the machine, two things may happen: 1) The volume of human translation could increase; 2) the perceived value of human translation could increase."

Nabil Frej and John Yunker have posted on their blogs the preliminary results from the “Which Engine Translates Best?” challenge organized by Gabble On which asks volunteers to evaluate Google Translate, Microsoft Bing, and Yahoo Babel Fish translations (if you haven't done it yet, I strongly suggest you spend 10 minutes doing it). And it looks as if Google is doing a better job than the other two, but with some exceptions.

From a translation business perspective, I am adopting a pragmatic approach. At Milengo, we are running a few pilot projects with some of our clients to evaluate seven language pairs using the Asia Online technology. We have also used the API for Google Translator Toolkit to connect it with Milengo's Translation Management System and we are currently running some test projects with it. Our goal with these efforts is not to replace human translation, but to increase productivity and to allow our clients to translate content that would otherwise never be translated because of cost and deadlines.


The situation reminds me of a story that my friend João Roque Dias, from Portugal, told me about how government officials in Portugal would fend off requests in the late 70s by saying that outcomes were unpredictable because the country was in a PREC (Processo Revolucionário em Curso or Revolutionary Process In Progress), which eventually became synonymous with "a mess."  Language technology for me is in a PREC: Any outcome is possible, so I am hedging my bets!

Saturday, December 12, 2009

Video of My Presentation in Bangkok

Here is a video of my presentation in Bangkok. If you missed the presentation at the ATA in New York, you will see that this one is very similar. The video lasts 36 minutes.

Monday, June 15, 2009

Google Translator Toolkit: A New Player in Translation Technology

This week, Google launched its new platform for translation projects, the Google Translator Toolkit. The tool is designed for translators and is similar to translation memory (TM) tools available in the market -- such as Across, Déjà Vu, Trados, and Wordfast -- and integrates Google Translate's statistical machine translation.



As we have been discussing in Common Sense Advisory's research, and in recent industry gatherings, this is the long-needed revolution in an industry that has been trying to "out-Trados" Trados, or trying to increase the productivity of processes and pump up technology that is old and cumbersome. Google Translator Toolkit incorporates all the collaboration features of current technology in an elegant way and enables translators to regain control of the process.
Even though it is still a bare bones solution, it will attract early adopters. Hardcore TM users, on the other hand, will likely shun the new technology.

It is still early to predict the impact of this launch, but we expect that the following will happen:
  • TM tools will develop interfaces that will read/write Google TMs and Google MT if they want to stay in the market.

  • Pre-translation and post-editing will become standard practices, even for the most recalcitrant translators.

  • Discussions about intellectual property of translation memories will become irrelevant, with negative impact for efforts like TM Market Place and the TAUS TDA initiative.

From the Google Translator Toolkit website, we also learn that:

  • It supports 47 languages.
  • Translations and glossaries each have a maximum size of 1MB.
  • Documents can be uploaded in most common file formats.
  • Translation memories have a maximum size of 50MB per upload.
  • Google Translator Toolkit is free, but in the future, Google plans to charge users whose translations exceed high-volume thresholds.

Google Translator Toolkit is not perfect. There are valid concerns about using it, along with the predictable resistance to change by those tied to the existing model. However, Google has already changed our behavior in the way we look for information. Now, it is launching a platform that has the potential to revolutionize the translation process, especially if combined with Google Wave, which is expected to be launched soon.

The role of the language services industry is to evolve from this stage. Alea jacta est!

Tuesday, October 30, 2007

Google MT and dotSUB - an atomic combination

I wanted to do something cool for my presentation at the ATA this weekend -- Quality Still Doesn't Matter, which focuses on topics that should really matter to translators and LSPs, such as productivity, new technologies, sales, etc. -- and decided to start with a video subtitled in dotSUB.

dotSUB is a cool site that allows you to upload your video, transcribe and subtitle very fast. I actually wrote about it in the Global Watchtower almost one year ago.

So... this week I read the news about Google abandoning SYSTRAN and starting to use its own Statistical Machine Translation Engine instead. I played with it a little bit by writing some text in Portuguese and having it translated into English. I was very surprised with the quality of the translation into English. That's when I decided to really play with Google Translate and dotSUB.

Here is what I did:

1) Wrote a script in Portuguese.
2) Had it translated into English with Google Translate.
3) Read the script to my webcam.
4) Pasted the Portuguese text into dotSUB.
5) Pasted the English transaltion into dotSUB.

All of this took me no more than 10 minutes to do.

Then I decided to have some fun. I edited the English translation a little bit and used Google Translate to go from English into Arabic, Spanish, French, Italian, Japanese, Chinese, and Russian.

I found that the quality of the translation was much better than I expected. I can judge Spanish, French, and Italian. I asked someone here in the office to check the Russian, but I have no idea of how the Japanese, Chinese, and Arabic came out. I don't even know if my visual pasting of the subtitles didn't break any words in the middle.

Check out for yourself. Try changing the languages using the small arrows in the bottom right corner of the video.




What I am going to say in my presentation is essentially that translators that are not using Google to pre-process their jobs, are doing too much work. MT is here to stay... as I had predicted a couple of years ago, this can be the disruptive player in the market.

Look out for mash-ups of Translation Memory technologies with Google from Elanex, XML-Intl, Proz.com, and other players. I can now see huge projects incorporating all these new technologies: a Ning portal for discussion and training, Google Translate for preprocessing translations, a shared translation memory repository from LingoTek, a wiki in wikidot.com for editing the translations in a collaborative way. These are all free technologies that would allow a company to manage a huge project in a much more efficient way than using the tools of today. All of this could be managed in ]Project Open[ and the sales process might have been tracked in FreeCRM or SugarCRM.