aitextcleaner.top

Remove Accents from Text — Strip Diacritics Online

Convert accented characters like é, ü, ñ, and ç to their plain ASCII equivalents. Paste your text, click Remove Accents, and get a clean ASCII string ready for usernames, filenames, URLs, and legacy systems.

0 characters

What are diacritics and why do they cause problems

A diacritic is a mark added to a base letter to change its pronunciation or meaning. The most familiar examples are the acute accent (é), the umlaut (ü), the tilde (ñ), the cedilla (ç), the circumflex (â), the grave accent (à), and the ring above (å). Together these and dozens of similar marks form a family called diacritical marks or, collectively, accents.

In everyday prose they are essential. French résumé is not the same word as resume. German über loses its meaning when written uber. Spanish año means year; ano means something quite different. For human reading the distinction matters enormously.

For machines, however, accented characters frequently cause failures. A username field that accepts only [a-zA-Z0-9_] rejects José. A URL router that does not percent-encode non-ASCII characters produces a broken link when the path contains résumé. A legacy database using Latin-1 encoding stores ü correctly but throws an error or produces garbled output when it receives a character outside that encoding range. Stripping the diacritics — converting José to Jose and résumé to resume — produces an ASCII string that passes every one of those constraints.

How accent removal works: Unicode normalization and the NFD trick

Unicode represents most accented characters in two forms. The composed form (NFC) stores é as a single code point U+00E9. The decomposed form (NFD) stores it as two code points: the base letter e (U+0065) followed by the combining acute accent (U+0301). The combining character is a zero-width mark that attaches visually to the preceding letter.

The algorithm for stripping accents follows directly from this: normalize the input to NFD so every accented character is split into a base letter plus one or more combining marks, then delete all characters in Unicode category Mn (Mark, Nonspacing). What remains is the sequence of base letters without any diacritics. In JavaScript that is one expression: text.normalize('NFD').replace(/\p{Mn}/gu, '').

This handles the full Latin Extended Additional block — characters like Vietnamese ộ (which decomposes into o + combining hook below + combining dot below), Polish ł (Latin small letter L with stroke — note: this is a separate base letter, not a combining mark, and is not stripped by the NFD method), and the many precomposed characters in ISO 8859-1, ISO 8859-2, and Windows-1252.

Characters that do not decompose — such as ø (Latin small letter O with stroke), ð (eth), þ (thorn), ß (German sharp S), and æ (ash) — are not affected by the NFD approach because they have no combining-mark decomposition. This tool removes only true diacritics; characters with no combining decomposition are left in place. If you need those converted too, a full transliteration table (ø→o, ß→ss, æ→ae) requires a separate lookup — outside the scope of this tool.

Common use cases with concrete examples

Username generation. A new user registers with the display name Björn Müller-Lüdenscheidt. The authentication system accepts only letters, digits, hyphens, and underscores. After accent removal: Bjorn Muller-Ludenscheidt. After removing the space and enforcing the allowed character set: bjorn-muller-ludenscheidt — a valid, readable username candidate.

URL slug creation. A blog post titled Café Culture in São Paulo needs a URL-safe slug. After accent removal: Cafe Culture in Sao Paulo. After lowercasing and replacing spaces with hyphens: /cafe-culture-in-sao-paulo — a clean, crawlable permalink with no percent-encoding required.

Filename sanitization. A batch script processes uploaded files. One file is named Rapport_Année_2024.pdf. Many filesystems and shell scripts fail on non-ASCII filenames, especially across operating systems. After accent removal: Rapport_Annee_2024.pdf — safe on every platform.

Legacy system import. A CSV export from a modern CRM contains French and Spanish customer names. The receiving system uses ASCII-only encoding. Running the names through accent removal before the import eliminates encoding errors without discarding the data.

Full-text search normalization. A search index treats naïve and naive as different tokens. Stripping accents from both the stored content and the query before indexing means that a user who types either spelling finds the same documents.

A step-by-step walkthrough

1. Paste or type your text in the input box above. There is no size limit enforced by the tool; practical limits are set by your browser's available memory, which handles tens of thousands of words without difficulty.

2. Click the Remove Accents button. Processing runs in your browser as JavaScript — no text is sent to any server. The output appears immediately in the result box below.

3. Check the character count displayed beneath the boxes. If the count dropped significantly more than expected, scan the output for characters that were unexpectedly modified. The NFD approach targets only combining diacritical marks; base letters without decomposable combining marks are untouched.

4. Click Copy to copy the output to your clipboard. The button text changes to Copied for two seconds to confirm the operation.

5. For downstream processing — username validation, slug generation, or CSV import — paste the output into your target system or the next step of your pipeline.

Which accents are removed: a reference table

The following table shows common accented characters and their ASCII equivalents after diacritic stripping. The list covers the characters most frequently encountered in Western European languages: French, Spanish, Portuguese, German, Italian, Dutch, Swedish, Norwegian, Danish, and Romanian.

AccentedASCII outputLanguage examples
à á â ã ä åaFrench, Spanish, Portuguese, German, Swedish
è é ê ëeFrench, Spanish, Italian
ì í î ïiFrench, Spanish, Italian, Portuguese
ò ó ô õ öoSpanish, Portuguese, German, French
ù ú û üuFrench, Spanish, German
ý ÿyFrench, Czech
çcFrench, Portuguese, Turkish
ñnSpanish
ž š čz s cCzech, Slovak, Slovenian, Croatian
ă ș ța s tRomanian

Note that uppercase variants (À Á Â Ã Ä Å É Ê Ë etc.) follow the same rule — the diacritic is stripped but the base letter case is preserved. RÉSUMÉ becomes RESUME, not resume.

What this tool does not change

Case. Accent removal does not lowercase the output. Ångström becomes Angstrom, not angstrom. If you need lowercase output, apply a case conversion after pasting the result.

Ligatures and special Latin letters. Characters such as ø (O with stroke), æ (ash), œ (OE ligature), ß (German sharp S), ð (eth), and þ (thorn) do not decompose under NFD normalization and are therefore left unchanged. Converting these to ASCII equivalents requires a hard-coded substitution table, which is outside the scope of this diacritic stripper.

Non-Latin scripts. Arabic, Hebrew, Greek, Cyrillic, and other scripts may have their own combining marks. The Mn category removal used here will strip combining marks from those scripts too if they appear in the input. This is usually the correct behavior for ASCII normalization, but verify your output if the input contains mixed scripts.

Spaces, punctuation, and formatting. Nothing other than combining diacritical marks is removed. If you also need to strip special characters or extra spaces, run the output through the remove special characters tool or the extra spaces tool.

Accent removal in popular programming languages

If you need to automate the conversion in code rather than using this tool interactively, the NFD approach translates directly into every major language.

JavaScript / TypeScript: str.normalize('NFD').replace(/\p{Mn}/gu, '') — the u flag enables full Unicode mode and \p{Mn} matches any nonspacing mark.

Python 3: import unicodedata; ''.join(c for c in unicodedata.normalize('NFD', s) if unicodedata.category(c) != 'Mn') — the unicodedata module ships with the standard library.

PHP: transliterator_transliterate('Any-Latin; Latin-ASCII', $str) using the intl extension, which performs the full NFD + Mn removal in one call and also handles ligatures.

Java: Normalizer.normalize(str, Normalizer.Form.NFD).replaceAll('\\p{InCombiningDiacriticalMarks}+', '') using java.text.Normalizer.

Ruby: str.unicode_normalize(:nfd).gsub(/\p{Mn}/, '') — the unicode_normalize method is part of core since Ruby 2.2.

In all cases the logic is identical: decompose to NFD, then remove code points in the Nonspacing Mark (Mn) category. The browser-based tool on this page uses the JavaScript version.

Performance: how many characters can this handle

The JavaScript implementation uses a single normalize call followed by a single regex replacement pass. Both are linear — O(n) in the length of the input. A modern browser processes approximately 10 million characters per second for this operation.

A typical blog post runs 1,000–2,000 words, roughly 6,000–12,000 characters. Processing time is under 2 milliseconds. A full novel (100,000 words, ~600,000 characters) completes in under 100 milliseconds. A large database export — say, 50,000 customer names concatenated — finishes in under a second. For bulk processing beyond a few million characters, a server-side pipeline in Python or Java is more ergonomic, but the tool handles any realistic copy-paste task without hesitation.

Related tools on this site

If your goal is to make text fully safe for URLs, you likely need more than just accent removal. The remove special characters tool strips punctuation, brackets, and symbols while keeping letters and numbers. Combined with accent removal, the two tools produce a clean alphanumeric string from any European text.

For filenames that need to be safe on every operating system, also check for extra spaces with the extra spaces remover and duplicate lines with the remove duplicate lines tool.

If the text came from a word processor, it may also contain smart quotes and em dashes that need converting before the string is truly ASCII-clean. The smart quotes remover and em dash remover handle those cases.

Frequently Asked Questions

▸Does removing accents change the meaning of the text?

For a human reader, yes — in languages that depend on accents to distinguish words, stripping them changes or loses meaning. <em>résumé</em> becomes <em>resume</em>, <em>año</em> becomes <em>ano</em>. The purpose of this tool is not human readability but machine compatibility: generating usernames, file paths, URL slugs, or CSV fields where only ASCII is accepted. Always keep the original text and use the de-accented version only for the specific technical context that requires it.

▸Why are ø, æ, ß, and ð not converted?

These characters do not contain a combining diacritic — they are independent base letters with their own Unicode code points. The Unicode NFD normalization used here only decomposes characters that are built from a base letter plus a combining mark (like é = e + ́). Converting ø to o, ß to ss, or æ to ae requires a hard-coded substitution table, not a normalization step. If you need those conversions, a library-level transliteration function — such as PHP's <code>transliterator_transliterate('Any-Latin; Latin-ASCII', $str)</code> or Python's <code>unidecode</code> package — covers them.

▸Is this tool safe for passwords or sensitive data?

Processing runs entirely in your browser as client-side JavaScript. Nothing is sent to any server and nothing is logged or stored. You can confirm this by opening your browser's developer tools, going to the Network tab, and running the tool — no outgoing request will appear. That said, copying passwords through any third-party website is not a best practice. For general text — names, titles, descriptions, filenames — there is no privacy concern.

▸Will this handle Vietnamese, Turkish, or Eastern European text?

Yes for true diacritics. Vietnamese characters like <code>ộ</code> decompose under NFD into a base letter plus two combining marks (hook below and dot below), and both combining marks are removed, leaving <code>o</code>. Romanian <code>ș</code> and <code>ț</code> decompose to <code>s</code> and <code>t</code>. Turkish <code>ğ</code> (g with breve) and <code>ş</code> (s with cedilla) decompose to <code>g</code> and <code>s</code>. Czech and Slovak <code>č š ž ř</code> all decompose correctly. The exception is the dotless i (<code>ı</code>) used in Turkish, which has no decomposition and is left unchanged.