Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes—attackers can exploit differences between an email’s original bytes, the text a security tool extracts, and the HTML a mail client renders. Invisible Unicode characters can split words without changing their visible appearance. HTML and CSS can hide or vary content so a filter and a recipient effectively inspect different messages. Neither technique defeats every product: modern defenses usually combine normalization, link analysis, reputation, authentication, and behavioral signals.

How an email becomes three different things

An incoming message has a MIME source containing headers, encoded parts, HTML, plain text, images, and links. A gateway may inspect raw source, decode transfer encodings, extract visible text, normalize Unicode, or render HTML in a separate analysis engine. The recipient’s mail app then applies its own HTML and CSS rules.

Those stages do not necessarily produce the same representation. A detector may miss text hidden in an image or fail to match a word interrupted by non-rendering characters, while the client displays a convincing sentence. Conversely, a security system may expose suspicious content that the user never sees. The practical risk is this representation gap—not a guarantee that a particular trick bypasses a particular vendor.

How invisible Unicode disrupts text matching

Tag characters and “ASCII smuggling”

Microsoft Security Research described a September 3, 2026 phishing campaign that inserted invisible Unicode Tag characters from U+E0000–U+E007F into financial lures. Characters placed inside a word such as “funding” can interrupt the sequence a rule expects, while the rendered message still appears to say “funding.” Microsoft calls the broader use of invisible or non-rendering Unicode to hide content inside normal-looking text “ASCII smuggling.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This specific mechanism is different from homoglyph substitution (using look-alike letters), bidirectional-control abuse, or every other Unicode spoofing method. A detector that searches only for a contiguous ASCII keyword may therefore miss a transformed string unless it decodes and normalizes the text first.

What Microsoft’s telemetry does—and does not—show

Microsoft reported that hits on a hunting signature for ASCII smuggling rose sharply from February 9, 2026, and remained elevated on weekdays for about three months. That is Microsoft’s telemetry for its hunt and the observed campaign, not an industry-wide prevalence rate. The company also said most messages were caught by layered protections rather than a single Unicode-specific rule.

Can HTML emails hide text from spam filters?

Hidden and conditional content

HTML permits content that is visually absent or presented differently: CSS can set text to display:none, make it transparent or microscopic, position it off-screen, or show different blocks under different conditions. Images can carry words that are unavailable to text-only extraction, and malformed markup can cause parsers to build different document trees.

A 2024 study by Lucas Betts, Robert Biddle, Danielle Lottridge, and Giovanni Russello examined HTML/CSS concealment in unsolicited email and described ways to create multiple message permutations. Its findings support concealment as a filter-evasion risk, not the claim that every method works against every gateway or mail client.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why links are especially deceptive

The text displayed for a link is independent of its destination. Unicode Technical Report #36 gives a historical HTML example in which a familiar-looking URL masks another target, and discusses visually confusable characters. A visible brand name or domain in an email therefore does not prove where a click will go.

What quantitative studies actually measured

A 2025 preprint by Dalmiere, Zhou, Auriol, Nicomette, and Marchand analyzed 386 verified phishing emails. In that dataset, the reported body-obfuscation techniques included:

Technique Share in the study sample Scope and qualification
Text in image 47.0% Percentage of the 386-email dataset, not an industry-wide rate
Base64 encoding 31.2% Percentage of that study’s sample
Invalid HTML 28.8% Percentage of that study’s sample
Model reported in the paper R² = 0.486, p < 0.001 Associations in the paper’s configuration; not proof of universal causation

The paper reported significant antispam-evasion associations for Base64 Encoding and Text in Image in its configuration, and higher scores correlated with Invalid HTML. These are sample-specific findings, not a current benchmark of products or a promise that any technique will evade your mail system.

Defensive controls that address the representation gap

Normalize before matching

  • Decode MIME transfer encodings and quoted-printable or Base64 parts before inspection.
  • Inspect original text and a consistently normalized form, including removal or flagging of suspicious invisible and format characters.
  • Run link extraction and logging on the same transformed representation used for detection, while retaining the original bytes for investigation.
  • Evaluate decoded text inside attachments and images where your risk model requires it; document confidence and false-positive behavior.

No single normalization recipe is established as a universal fix. The required sequence depends on the mail pipeline, supported languages, and the representations each control can safely process.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use layered detections

Combine Unicode anomaly signals with sender authentication, domain and URL reputation, message-level similarity, attachment analysis, and user or organization context. Microsoft’s campaign account specifically illustrates why a character-specific signature should be one signal among several.

Handle internationalized addresses carefully

Unicode UTS #39 version 18.0.0 (August 27, 2026) describes checks for internationalized email identifiers, including NFKC formatting of the local part, restriction-level checks, mixed-number-system checks, filtering certain quoted-string characters, and flagging suspicious incoming addresses. It also warns that bidirectional reordering can alter display and recommends isolates or equivalent handling around address components.

These controls are not a blanket ban on non-ASCII mail. UTS #39 explicitly states, “This profile does not exclude characters from EAI.” The objective is to detect structurally suspicious or unexpected combinations while preserving legitimate multilingual communication.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How recipients can check a suspicious link

  1. Do not rely on the visible label. Hover over the link on a desktop, or use the client’s link-details action, to reveal the actual destination.
  2. Read the registrable domain. Compare the domain immediately before the first single slash after the host with the organization’s known domain; a brand name in a subdomain or path is not the same thing.
  3. Watch for look-alike characters. Mixed scripts, unexpected accents, or an address that changes when copied can indicate confusable Unicode.
  4. Inspect safely. Do not open an unfamiliar link merely to test it. Navigate to the organization through a bookmark or a separately typed address, and verify requests for credentials, payment, or sensitive files out of band.
  5. Report the message. User reporting gives layered filters additional evidence; vigilance does not replace technical controls.

What is not established

  • There is no cited controlled comparison ranking email-security vendors on Unicode or HTML evasion.
  • No universal percentage describes how often these techniques bypass filters.
  • No particular normalization setting is guaranteed to block every attack.
  • The Microsoft figures describe one campaign and Microsoft telemetry; the academic percentages describe one 386-message dataset.

The Bottom Line

Unicode and HTML concealment exploit a mismatch between what software parses and what people see. The durable response is consistent decoding and normalization, internationalization-aware identifier checks, link and content analysis, and layered controls—not a blanket ban on Unicode or confidence in one detection signature.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.