Invisible Unicode Characters Explained
A reference to invisible unicode — zero-width spaces, joiners, bidi marks and lookalike spaces — with code points and fixes.
· 7 min read
Unicode contains hundreds of code points that render as nothing at all. They exist for good reasons — typography, script shaping, text direction — and they cause a remarkable amount of trouble once they escape into ordinary text.
The zero-width family
| Code point | Name | Purpose |
|---|---|---|
| U+200B | Zero Width Space | Allows a line break with no visible space |
| U+200C | Zero Width Non-Joiner | Prevents two characters from ligating |
| U+200D | Zero Width Joiner | Forces joining — builds multi-part emoji |
| U+2060 | Word Joiner | Forbids a line break |
| U+FEFF | Zero Width No-Break Space | Legacy byte order mark |
U+200B is the one people mean by "invisible character". U+FEFF is the classic cause of a mystery character at the start of a UTF-8 file. U+200D is load-bearing: strip it and 👨👩👧 falls apart into three separate emoji.
Invisible mathematical operators
U+2061 FUNCTION APPLICATION, U+2062 INVISIBLE TIMES, U+2063 INVISIBLE SEPARATOR and U+2064 INVISIBLE PLUS carry semantic meaning for maths software — they tell a parser that f(x) is a function call rather than multiplication. In plain prose they have no business being there at all, which makes them a favourite carrier for hidden data.
Bidirectional formatting marks
U+200E, U+200F and the range U+202A–U+202E control text direction for mixed left-to-right and right-to-left content. U+202E RIGHT-TO-LEFT OVERRIDE is notorious: it can make a file named exe.txt display as txt.exe, a trick used in real-world malware distribution. Most systems now strip or flag it in filenames for that reason.
Lookalike spaces
These are not invisible — they render as whitespace — but they are not the space character, and code comparing strings does not care how they look.
- U+00A0 Non-breaking space — the single most common intruder, injected by HTML and by Word
- U+2002–U+200A — en space, em space, thin space, hair space and the per-em fractions
- U+202F Narrow no-break space — appears around punctuation in French typography
- U+3000 Ideographic space — full-width space used in CJK text
A good cleaner converts these to a regular space rather than deleting them. Deleting an U+00A0 turns "10 kg" into "10kg" and, worse, can fuse two words permanently.
Variation selectors
U+FE00–U+FE0F select between glyph variants — most visibly U+FE0F, which turns a monochrome symbol into a colour emoji. Because there are sixteen of them and they attach silently to any character, they have become the standard vehicle for unicode steganography: an entire hidden message can be encoded in variation selectors attached to a single visible emoji, and it survives most copy-paste.
What breaks when they get in
- Exact-match search fails. The text on screen is identical; the bytes are not.
- Logins fail inexplicably. A zero-width space pasted into a password field is part of the password.
- Code will not compile. An invisible character inside an identifier produces an error on a line that looks perfect.
- URLs break. A U+200B inside a link makes it a 404 that reads correctly.
- CSV and JSON parsing corrupts. A BOM mid-file, or an U+00A0 where a delimiter was expected.
- Character limits are hit early. Invisible characters still count.
How to find and remove them
Paste the text into the invisible character detector. The Analyze tab lists every non-ASCII code point with its unicode name and count; the Visualize tab labels each one inline so you can see exactly where it sits; the Clean tab strips them, converting invisible spaces to regular spaces so word boundaries survive.
In code, the equivalent is a regex over the relevant ranges:
text.replace(/[\u200B-\u200F\u2028\u2029\u2060-\u206F\uFEFF\uFE00-\uFE0F]/g, "")
.replace(/[\u00A0\u2000-\u200A\u202F\u205F\u3000]/g, " ")
Two passes, deliberately: delete the truly invisible, normalise the space lookalikes. Doing it in one pass is the mistake that fuses words together.
Related
FAQ
+ − What are invisible unicode characters?
Code points that render with no visible width, such as the zero width space (U+200B), word joiner (U+2060) and variation selectors (U+FE00–U+FE0F). They exist for typography and script shaping but frequently end up in ordinary text through copy-paste.
+ − How do I find invisible characters in a string?
Paste the text into an invisible character detector, which lists every non-ASCII code point with its unicode name and count. In code, match the ranges U+200B–U+200F, U+2060–U+206F, U+FEFF and U+FE00–U+FE0F.
+ − Are invisible characters dangerous?
They can be. U+202E right-to-left override has been used to disguise executable filenames, and variation selectors can encode hidden data inside ordinary-looking text. Most occurrences are harmless copy-paste accidents, but they routinely break search, logins, URLs and code.
Clean your text now
Free, instant, and entirely in your browser — nothing is uploaded.
Keep reading
Does ChatGPT Watermark Its Text?
OpenAI built a ChatGPT text watermark and never shipped it. What is confirmed, what is myth, and how to check your text.
Does Claude Watermark Its Text?
Since August 2026 Claude embeds a watermark in all generated text, worldwide. What it is, what it proves, what removes it.