← the writing notes 10 min

Machine Translation + Humanization: The Pipeline That Ranks

Raw machine translation doesn't rank. It barely reads. Here's the exact pipeline I use at Seahawk to turn DeepL drafts into localised content that actually earns organic traffic in competitive non-English markets.

Open vintage dictionary on a wooden desk with golden afternoon light and shallow depth of field

A client rang me in early 2023 in a genuine panic. They'd paid an agency (not Seahawk, for the record) about £14,000 to localise their e-commerce site into French, German, and Spanish. Eight months later, zero organic impressions from any of those three locales. The content existed. Google had indexed it. Nobody was clicking, nobody was ranking, and nobody could figure out why.

I pulled up one of the German product pages. Three sentences in, I spotted the problem. The copy read like a Google Translate screenshot from 2011. Technically accurate. Utterly unnatural. No German speaker would write "Das Produkt ist sehr gut für Ihre Bedürfnisse geeignet." Not unless they were a robot transcribing a PowerPoint deck.

That's the machine translation trap. And I see it constantly, across the 12,000+ sites we've touched at Seahawk. Raw MT output doesn't rank because Google's quality signals aren't fooled by technically correct foreign text. Neither are users. And when users bounce in 4 seconds, the ranking signal loop collapses fast.

Here's the pipeline that actually works.

---

Why Raw MT Fails Search (And It's Not What You Think)

Most people assume MT fails because of grammar errors. That's the wrong diagnosis.

Modern tools like DeepL are genuinely impressive. DeepL's translations, especially for European languages, are often grammatically clean. The problem is something subtler: cultural register and search intent mismatch.

The Search Intent Problem Across Languages

When a French user types "chaussures de course femme pas cher" into Google, they're signalling a very specific intent. A literal English-to-French translation of your "cheap women's running shoes" page will probably produce that phrase. Fine. But the body copy, the headings, the FAQ structure, the way you address the reader, all of it will carry the cadence of English-language content wearing a French costume.

French readers notice. French crawlers (i.e., Googlebot rendering pages for the fr locale) measure engagement. Short sessions, high bounces, low scroll depth. Those signals push you down.

I ran an A/B test on a Seahawk fashion client's French site in mid-2023. Version A was pure DeepL output, proofread for grammar. Version B was the same DeepL base, run through a native French editor for 25 minutes per page. Version B had a 34% lower bounce rate within six weeks. Same URLs, same backlinks, same technical setup. Only the humanisation layer changed.

---

The Tools I Actually Use (In Order)

Look, I've tried a lot of things. Here's what's in the current Seahawk stack for MT-plus-human work.

  1. DeepL Pro for the base translation. Not Google Translate. Not ChatGPT set to "translate this." DeepL's neural model is trained on high-quality bilingual corpora and the output is significantly cleaner for most European languages. For Japanese and Korean, I lean more on a custom GPT-4o prompt we've built internally, because DeepL's Asian language quality is patchier.
  2. Phrase (formerly Memsource) for translation memory and glossary management. If a client has brand terms, product names, or specific taglines that must not be translated, Phrase keeps that consistent across every page, every sprint, every editor. Without TM tooling, you get "Shopping Cart" on one page and "Panier" on another and "Panier d'achat" on a third. Consistency matters for crawlers.
  3. A native human editor. This is non-negotiable. I use a network of freelancers sourced primarily through Workana and ProZ. Not professional translators doing a fresh pass, that's expensive and slow. I use them as editors: 20-30 minutes per 800-word page, focused on three things only: natural phrasing, cultural reference swaps, and CTA language.
  4. Surfer SEO (or Clearscope for some clients) for on-page optimisation in the target language. You cannot use your English keyword data and assume it maps cleanly. I've seen clients obsessively optimise for a French keyword that gets 40 searches per month when the actual high-volume phrase is something slightly different. Surfer's Content Editor works in French, German, Spanish, Portuguese, and several others.
  5. Google Search Console split by country property. Set this up before you launch a single localised page. You need per-locale data from day one.

---

The Humanisation Brief: What I Tell Editors

This is where most pipelines leak. You can't just say "make it sound natural" and send a DeepL doc to a freelancer. I've done that. It produces inconsistent results and wastes everyone's time.

My standard humanisation brief covers four things:

  • Register: Is this formal or informal? German "du" vs "Sie" is a business decision, not a grammar question. Get it wrong and you signal the wrong brand positioning immediately.
  • Idiom swaps: Flag any English idioms in the source and ask the editor to replace them with native equivalents, not translate them literally. "Hit the ground running" is not an idiom in French. It's confusion.
  • CTA rewrite: Every call-to-action gets a fresh write, not a translation. "Buy Now" in German becomes "Jetzt kaufen" grammatically, but a native editor might know that "Gleich bestellen" converts better in their market. Let them decide.
  • Length tolerance: Some languages expand significantly. German is notorious. A 250-word English section often becomes 310+ words in German. Factor this into your layout templates or you'll have broken designs on every device.

Seahawk had a SaaS client in 2022 where the German version of their pricing page completely broke the card layout because nobody had budgeted for text expansion. We caught it in QA but it cost us two days. Now it's in every brief we send.

---

Keyword Research in the Target Language: Do It Again From Scratch

I cannot stress this enough. Your English keyword research is a starting point, not a map.

Back in 2020, I worked on a travel site targeting the Italian market. The client had an English page optimised for "budget hotels Rome" ranking solidly. We translated it for Italian, optimised for "hotel economici Roma" (the literal translation), and it limped along at position 18 for three months.

Then a native Italian editor pointed out that Italian users more commonly search "hotel a basso costo Roma" or even "hotel conveniente Roma." We swapped the primary phrase and rewrote the meta. Within eight weeks, position 6. Same domain authority, same backlinks, same page structure.

The Google Keyword Planner lets you switch the language and location settings. That's your starting point. Then cross-reference with Ahrefs or Semrush filtered to the target country. Look at what your actual local competitors are ranking for, not what you assume they'd target.

One more thing: search volume thresholds are different across markets. A French keyword with 800 monthly searches might be worth going after. In the UK, I'd probably skip it. Market size context matters.

---

Technical Setup: hreflang, Subfolders, and the Mistakes I Keep Seeing

The content pipeline is only half of it. The technical scaffolding either amplifies your work or quietly cancels it out.

hreflang Tags

Get them right or don't bother. Google's hreflang documentation is clear on this: every localised page needs a bidirectional hreflang implementation. Your English page points to your French page, your French page points back to your English page and to every other locale. Miss one link in the chain and the whole signal degrades.

I audit hreflang implementations on probably 40-50 sites a year through Seahawk. I'd say 60% of them have at least one broken or missing reciprocal tag. Use Screaming Frog's hreflang validator after every deploy.

Subfolders vs Subdomains

My default recommendation is subfolders (/fr/, /de/) over subdomains ( fr., de.). The domain authority consolidation is worth it, and subfolders are easier to manage in most CMS setups. Subdomains make sense in very specific situations, usually where the international version needs a separate CMS or has significantly different technical architecture.

Canonical Tags

Watch for self-referential canonicals that point to the English page. This is a WordPress plugin problem. Some multilingual plugins (looking at you, older WPML configurations) used to set the canonical on translated pages to the original English URL. That single misconfiguration will tank your entire localisation effort. Check every locale's canonical on launch day.

---

Measuring the Pipeline: What Numbers I Watch

After launch, here's what I track and roughly when I expect to see movement:

  • Weeks 1-4: Indexation rate per locale in GSC. If Google isn't indexing your localised pages within two weeks, you have a crawl or sitemap problem.
  • Weeks 4-10: Impressions by country in GSC. Impressions before clicks. If impressions are growing, the content is entering consideration sets. If impressions are flat, the keywords or content quality need revisiting.
  • Weeks 8-16: Avg. position movement on target phrases. This is where you'd expect to see the humanisation work paying off. Pages that bounced users fast won't climb here regardless of how good your hreflang is.
  • Month 4+: Organic sessions and on-site engagement per locale. Bounce rate, session duration, pages per session. These are the lagging indicators that confirm whether the pipeline produced genuinely useful content or dressed-up MT slop.

Honest answer: a well-executed MT-plus-humanisation page, with clean technical setup and sensible keyword targeting, typically shows meaningful ranking movement between weeks 10 and 20 in moderately competitive locales. Faster in low-competition markets. Slower in highly competitive ones like Germany or France where local publishers are strong.

---

What This Pipeline Actually Costs

Let me give you real numbers because vague ranges are useless.

For a standard 800-word page into one language:

  • DeepL Pro: roughly £0.04 per word at current API pricing. Call it £32 per page.
  • Human editor: a good B2-level native editor on Workana charges between £18-35 per page for a 25-minute pass. I budget £25.
  • Surfer SEO content optimisation pass (done in-house): 30 minutes at your internal rate.

So the hard cost per localised page is around £55-60 before your time. Compare that to a full human translation at £120-180 per page. You're saving 50-60% and, in my experience, the quality ceiling is close enough that Google can't tell the difference, if the humanisation is done properly.

Where you should still use full human translation: legal content, medical content, anything where a mistranslation creates liability. Don't run a terms-of-service page through DeepL and call it done.

---

FAQ

Does Google penalise machine-translated content?

Not automatically, no. Google's spam policies mention auto-generated content created "with the primary purpose of manipulating search rankings" as a violation. The operative phrase is "primary purpose." MT content that genuinely serves users in their language isn't the target. Low-quality, unedited MT spam is. The humanisation layer is your protection here, and it also happens to produce better content. Both things are true.

How many languages should I launch at once?

One, maybe two. I've seen people try to launch seven locales simultaneously and produce mediocre content in all of them. Pick your highest-opportunity market based on traffic data and existing demand signals, do it properly, learn from it, then expand. Seahawk typically recommends a phased rollout: one locale per quarter for the first year.

Can I use ChatGPT instead of DeepL?

Yes, with caveats. A well-engineered GPT-4o prompt with persona, register, and tone instructions can produce output competitive with DeepL for many language pairs. It's often better for Asian languages. The trade-off is consistency: without TM tooling sitting above it, you'll get variation in terminology across sessions. Use DeepL for European languages where it excels. Use GPT-4o for languages where DeepL's quality is weaker, with a carefully maintained glossary in the prompt.

What's the single biggest mistake in localisation projects?

Treating it as a translation project instead of a content project. Translation is about accuracy. Content is about performance. The moment your brief only asks "did they translate this correctly?" instead of "does this page earn traffic and convert visitors?", the project is already pointed in the wrong direction.

---

The pipeline isn't magic. It's a sequence: strong MT base, focused human edit, native keyword research, clean technical implementation, patient measurement. Each step is skippable. But every step you skip shows up in your GSC data about three months later, and by then you've already spent the budget.

Do it right the first time. It's cheaper.

Need this done, not just read?

start a project book 30 minutes