Humanizing AI Text in Spanish, French, and German
I ran the same AI-written paragraph through a detector and then a humanizer in three languages. All three started at 100 percent AI and finished at 0 percent, which sounds like a clean win until you count the words. The Spanish came out slightly longer. The French lost 28 percent of its content. The German lost 29 percent and mangled a compound noun on the way through.
That gap is the whole story of working outside English. Clearing the detector is the easy part, and in some languages there is no detector to clear. Keeping the text worth publishing is where it gets hard, and no tool will tell you when it has failed.
First, check whether a detector even runs on your language
Most advice on this topic assumes detection works the same everywhere. It does not, and the gap is bigger than people expect.
So how many languages does Turnitin's AI detection actually cover? Three: English, Spanish, Japanese. The documentation is blunt about everything else, saying a paper in an unsupported language will not be processed. Submit in French or German and you get no AI report. Not a low score. Nothing. Spanish arrived September 2024, Japanese April 2025, and reading the release notes right through mid-2026, a French or German model never shows up. The paraphrasing layer is English-only besides, so Spanish misses that one too.
I am spelling this out because a number of articles claim Turnitin covers seven languages including French and German. That is wrong, and you can check it three ways on Turnitin's own site: the capabilities FAQ, the file requirements page, and the release notes.
Self-serve tools cast a wider net. GPTZero picked up Spanish and French in April 2024, then German and Portuguese a year later, and by its own published figures German came out weakest of the four. Originality.ai advertises 30 languages, Copyleaks more than 30. Read those totals alongside GPTZero's own admission that most of its training data is English prose.
Detection is weaker in your language, and rephrasing breaks it unevenly
The most useful study here tested all four languages side by side. Researchers built matched sets of human, AI-generated, and AI-rephrased texts, then measured how well classification held up.
| Language | Raw AI text detected | After rephrasing |
|---|---|---|
| Spanish | 99% | 86% |
| English | 98% | 79% |
| German | 97% | 72% |
| French | 95% | 89% |
Raw generated text gets caught at similar rates everywhere. Rephrased text does not, and the spread is wide: German rephrasing slipped past detection far more often than French. Which means your odds depend on your language in a way nobody advertises.
There is a broader effect underneath this. In a benchmark across eleven languages, detectors trained on English dropped 25.7 percent when tested on other languages, and the authors concluded that English behaves as an outlier rather than a representative case. Detection quality is not a property of the tool. It is a property of the tool plus your language.
One finding cuts the other way and is worth knowing before you get clever. When Originality.ai translated text out of English and back again, AI text stayed caught at 100 percent, but human text started failing badly: false positives on human writing rose from under half a percent to about 28 percent after a round trip through Spanish. Translation laundering does not hide AI. It does put honest bilingual writers at risk, which rhymes with how detectors treat non-native English writing generally.
What my own run actually produced
Here is the test. One AI-generated business paragraph per language, scanned, humanized, scanned again.
| Language | Before | After | Word count change |
|---|---|---|---|
| Spanish | 100% AI | 0% | 74 to 81 words |
| French | 100% AI | 0% | 71 to 51 words |
| German | 100% AI | 0% | 63 to 45 words |
The detection result is uniform and the quality result is not. Spanish held its meaning and gained a little length, though it also gained a sentence that was not in the original, which is its own problem if you are writing anything factual.
French dropped a quarter of its substance and slid down a register. "Des outils technologiques avancés" became "la tech." "Rester compétitives" became "rester fortes," which means staying strong rather than staying competitive. "Les dirigeants" became "les chefs." Nothing there is ungrammatical. It is simply no longer business writing.
German came out worst. "Verbesserung der Kundenerfahrung," improvement of the customer experience, became "die Zeit für Kunden wird besser," the time for customers gets better, which is a different claim. And "datenbasierter Lösungen" came back as "Daten Lösungen," split into two words, which German orthography does not allow. That error is common enough to have a nickname, the Deppenleerzeichen, and it is a direct import of English spacing habits into a language that compounds.
One more result surprised me. I first wrote the German with ASCII substitutions, ae and oe and ue instead of ä, ö, ü. That version scored 0 percent AI. The identical paragraph with proper umlauts scored 100 percent. Same sentences, same meaning, opposite verdict, purely from character encoding. It is a good reminder that these scores are more brittle than they look, which the five detectors that disagreed on one human essay already suggested.
The traps a humanizer will not catch for you
Each language has conventions that a rewriting tool has no reliable way to get right, because they depend on who you are writing for.
Spanish: pick a region and stay in it. The clearest example is voseo. The Real Academia Española describes it as the use of vos to address someone, belonging to particular geographic and social varieties of American Spanish and carrying closeness and familiarity. The trap is what goes with it: pronominal voseo takes tuteo clitics and possessives, so it is te, tu, and tuyo, never os or vuestro. "Vos te quedás pensativo" is right. A tool that pairs vos with vuestro produces something no native speaker writes.
Vocabulary carries the same risk with higher stakes. The Academies' dictionary of americanisms marks coger as vulgar and taboo, meaning to have sex, across Mexico, most of Central America, Argentina, Uruguay and more, where Spain uses it for ordinary taking and grabbing. And zumo, the Spanish word for juice in Spain, denotes the oily substance in citrus peel in Mexico, Guatemala, Costa Rica and elsewhere. Neutral output that ignores your market can land somewhere between confusing and mortifying.
Two mechanical tells are worth checking every time. Spanish requires the opening ¿ and ¡, and the RAE is explicit that dropping them in imitation of languages that only use the closing mark is an error. And Spanish reaches for the definite article where English reaches for a possessive: me duele la cabeza, not mi cabeza, and levanta la mano rather than the calque levanta tu mano. Both survive translation intact and both read as imported.
French: the typography is part of the language. Quebec's Office québécois de la langue française wants a non-breaking space before a colon, guillemets instead of straight quotes, and a non-breaking space ahead of a percent sign, giving you 8 % rather than 8%. Numbers work the same way. Decimal comma, currency symbol after the amount. Which makes $24.99 wrong on two counts where 24,99 $ is right. Watch out for blanket rules here. Quebec puts no space before a semicolon, exclamation mark or question mark, while writers in France use a thin one, so anybody selling you the single French rule is half wrong wherever they stand. The undisputed part: plain English spacing and English quote marks are wrong on both sides, and they are the fastest signal that a French text got assembled somewhere else.
Word order gives it away too. A machine that translates "the first fifteen minutes" literally produces les premières quinze minutes, where French wants les quinze premières minutes. Same with verbs: English "ask a question" becomes demander une question instead of poser une question. The vocabulary is French and the joinery is English.
German: the verb goes where the grammar says. German sentences are built on a bracket. In a main clause the finite verb sits early and the participle closes the sentence: "Sie ist gestern Abend nicht gekommen." In a subordinate clause the whole verb complex moves to the end: "weil sie gestern Abend nicht gekommen ist." English rhythm imposed on German breaks that bracket, and so does splitting compounds. Both are things a native reader notices immediately and a detector never will.
Why the output degrades in the first place
None of this is bad luck. Translation studies has named these patterns for decades: explicitation, the tendency to spell out connections a native text leaves implicit, and shining through, where the source language's structures stay visible in the target. Recent work found machine translation shows a stronger and more persistent tendency toward shining through than human translators, who adapt to a register's norms.
The models are also just weaker outside English. An analysis by the localization company LILT attributed the bulk of non-English performance loss to model-level causes, including tokenizer inefficiency and a tendency to reason internally in English before producing another language. That is the same reason a humanizer's structural rewriting is less reliable here: it is doing harder work with thinner training data.
Worth saying plainly: I could not find any credible independent testing of humanizer tools in Spanish, French, or German. Every result that ranks for those queries is vendor or affiliate content with no published methodology. Even Undetectable.ai describes its multilingual support as secondary to English. Treat any specific bypass-rate claim in these languages as marketing.
A workflow that survives contact with a native reader
Write or generate in the target language rather than drafting in English and translating, since the round-trip finding suggests translation adds its own fingerprints. Humanize if you need to, using a tool that actually supports the language rather than one bolting it on, then immediately diff the output against the source and count the words. A 29 percent drop is not a rewrite, it is a deletion.
Then read for the four things the tools miss: register consistency, regional consistency, correct typography, and native word order. Fix the compounds. Restore anything the humanizer quietly cut, and delete anything it quietly invented.
Finally, have a native speaker read it. That is not a hedge to end an article on, it is the actual finding of my test. Every version passed the detector. Two of the three would still embarrass you in front of a reader, and the score gave no hint which two.
