Ian Muscat and Leanne Briffa of Have I Been Squatted, a company that tracks lookalike domains, say they found two gaps in Chromium's protection against internationalized domain name (IDN) spoofing. Using them, they registered 20 proof-of-concept domains that impersonate well-known brands. In their tests, Chromium showed these domains in convincing Unicode rather than the obviously odd Punycode form that begins with xn--.
The two characters
The researchers say Cyrillic ө, a letter used in Kazakh, Mongolian, and Tatar, is not on Chromium's lookalike list, so аррӏө.com displays in Unicode with no warning. The second character is the Latin "k with hook," ƙ, which is used in Hausa and Karai-karai. It keeps its hook through Chromium's internal comparison, so oƙta.com is not reduced to okta.com.
Have I Been Squatted's write-up includes a screenshot of Chrome 154 displaying аррӏө.com, a Cyrillic lookalike of apple.com registered for this research, in Unicode with a valid certificate. The padlock, the valid certificate and the familiar name all look correct, yet none of the five letters is Latin.
Examples from the 20 registered domains, with the Punycode form a careful browser would be expected to show:
| Displayed as | Imitates | Punycode (A-label) |
|---|---|---|
| аррӏө.com | apple.com | xn--80a6aa68c8d.com |
| ѕрасөх.com | spacex.com | xn--80a5aeq0fr0c.com |
| oƙta.com | okta.com | xn--ota-f6a.com |
| niƙe.com | nike.com | xn--nie-g6a.com |
The full list also includes imitations of openai.com, sharepoint.com, skype.com, kaspersky.com, chess.com, vk.com and kickstarter.com. The researchers own the domains, and they point to explanatory demo pages. Nothing in the research suggests the imitated brands were targeted or that these exact names have been used in real phishing campaigns.
Background: the 2017 fix
The flaw being bypassed was patched years ago. In April 2017, researcher Xudong Zheng registered аррӏе.com, a name made entirely of Cyrillic letters, and Chrome, Firefox and Opera all showed it as something nearly identical to apple.com. In response, Chromium introduced a "Whole-Script Confusable" (WSC) check that has held since. The job of the WSC check is to catch DNS labels written entirely in Cyrillic characters that all look like Latin ones. Zheng's domain encodes as xn--80ak6aa92e.com, and Chromium has shown that name as Punycode ever since.
The new research targets the word "all" in that rule.
Gap 1: the whole-script check is all-or-nothing
Chromium's first line of defense is the SafeToDisplayAsUnicode() function. Have I Been Squatted describes it as seven checks that run in order, and the first one that fails sends the name to Punycode. The checks cover:
- ICU spoof checks, including script mixing and an allowed-character list
- Icelandic þ and ð, which are only allowed under .is and .fo
- Azerbaijani schwa ə, which is only allowed under .az
- The middle dot, which is only allowed as Catalan l·l under .cat
- Routing between single-script and mixed-script handling
- The whole-script confusable check
- Dangerous-pattern checks, mainly for mixing Latin with CJK scripts
The header file in Chromium's own source confirms the categories behind these checks. It lists result codes for ICU spoof failures, TLD-specific characters, unsafe middle dots, whole-script confusables and dangerous patterns, and it uses Zheng's Cyrillic "apple" as its example of kWholeScriptConfusable.
Check 6 is where аррӏө.com gets through. According to the researchers, Chromium keeps a list of 29 Cyrillic characters that look like Latin letters. If a single-script Cyrillic label is made entirely of those characters, the browser shows Punycode. If even one character is missing from the list, the check does not reject the label. The researchers call that missing character a "breaker." Swapping the final е in Zheng's 2017 domain for ө is enough to change the result.
The researchers identified three breakers that still work today: ө (resembles o or e), ї (i) and ү (y). Two older breakers, ҏ and ӿ, stopped working in Chrome 148. That release moved to ICU 78.2 with Unicode 17, which reclassified those characters as uncommon, so they now fail check 1. By the researchers' count, 64 lowercase Cyrillic characters still pass the allowlist while being absent from the lookalike list. Not all of them look like Latin letters, but the pool is far from empty.
The researchers also give a version detail that is easy to miss. Chromium added м to the lookalike list on 11 September 2026, raising it from 28 to 29 characters. That change landed after Chrome 155 branched, so Chrome 155 still ships the 28-character list. The Register describes the breaker characters as absent from the list as of Chrome 154, which it says reached stable on September 22. I could not confirm that release date against an official Chrome release post.
Gap 2: hooks and bars survive the skeleton
The second layer, GetSimilarTopDomain(), reduces a hostname to a "skeleton" and compares it with a bundled list of popular domains. The researchers count 8,462 entries. Chromium's header describes two lookups: one uses the Unicode confusables skeleton, and the other uses a diacritic-free version of the hostname.
The ƙ character beats both lookups because its hook is not treated as a removable accent. The Register reports that the skeleton for oƙta.com comes out roughly as o k ' t a . c o r n (the skeleton maps "m" to "rn"), which does not match Okta. The bar on ө also survives, so аррӏө.com ends up as approximately a p p l o - . c o r n and does not match apple.com either.
In short: аррӏө.com passes the first layer because ө is not on the lookalike list, and it passes the second because its skeleton does not match Apple's. With both layers passed, the browser shows the domain in Unicode.
Gap 3: navigation warnings have blind spots
Chromium has another safeguard. In addition to the spoof checks, Chrome also implements a full page security warning to protect against lookalike URLs. You can find an example of this warning at chrome://interstitials/lookalike. This warning blocks main frame navigations that involve lookalike URLs, either as a direct navigation or as part of a redirect. Chrome also shows lighter "Safety Tip" prompts.
According to the researchers, these warnings check for an exact skeleton match, a one-character edit or a swap of two adjacent characters. The comparison runs against both the bundled list and sites the user visits often. That design leaves gaps:
- Too many edits:
аррӏө.comis two edits away from theappleskeleton, so the near-match check does not catch it. - Short names: the near-match check skips targets with fewer than five characters before the TLD. "Okta" has four.
- History-dependent: a user who has never visited the real site gets less protection from the engagement-based part of the check.
The researchers say their navigation tests did not include reputation services, remote allowlists or redirects. Safe Browsing and, in Edge, Microsoft's own reputation layer could still block a domain that is actually malicious. The display gap is real, but it does not mean every lookalike gets through to a credential prompt without a warning.
A note for Windows users: the research focused on Chromium, and the screenshot shows Chrome on macOS. Edge uses the same engine. I found no Edge-specific test results, so I can't say how Edge's additional protections handle these domains.
Webmail: Gmail and Outlook fail in opposite ways
Address-bar protections do nothing for the sender line in an inbox. The researchers sent their 20 test domains, a mix of lookalikes and real IDNs, in the From header to Gmail and Outlook on the web, using default settings.
- Gmail showed every domain in Unicode, including the lookalikes.
- Outlook Web showed every domain as Punycode, including the legitimate ones.
Neither client separated the harmless names from the deceptive ones. Outlook's approach is safer for spotting spoofs, but it is unfriendly to the many people who use non-Latin domains legitimately. The researchers say they did not measure spam filtering, sender reputation or link scanning, so this result only covers how the clients display sender domains, not how well each service blocks phishing overall.
How big is the problem?
The researchers took a snapshot of the .com zone from ICANN's Centralized Zone Data Service. It contained 167,226,216 domains, of which 733,509 were IDNs. They replaced each IDN's non-ASCII characters with the ASCII letters they resemble and found 161,894 pairs where an IDN looks like a registered ASCII domain, covering 133,953 distinct ASCII names. Of the 59,302 pairs they classified as "reverse-confusable," 46,522 would display in Unicode under their analysis.
That is not a count of phishing domains. Many will be defensive registrations, multilingual sites or unrelated businesses. Chromium's documentation acknowledges the trade-off: "Displaying either punycode or a visible security warning on too wide of a set of URLs would hurt web usability for people around the world." Kazakh, Mongolian and Hausa speakers rely on these characters, and banning ө outright would take a working tool away from them.
What admins and users can do
These suggestions build on the researchers' recommendations and common industry practice. They reduce risk but don't eliminate it.
- Monitor lookalike registrations for your brands, especially IDN and Punycode variants. Pay closest attention to short names: a four-letter brand like Okta falls under Chromium's five-character near-match threshold.
- Check your email gateway's impersonation controls. Find out whether it compares IDN sender domains against your own domains and key partners, rather than relying on how the mail client displays them.
- Don't treat a valid certificate as proof of identity.
аррӏө.comhad one. HTTPS only confirms you are connected to whoever owns that domain. - Use password managers and passkeys that match the exact stored domain. This is general practice, not a finding of this research. A password manager won't offer your Okta credentials on
oƙta.com, and that refusal is a useful warning in itself. - Teach staff that a Punycode address (
xn--) is worth a second look, not proof of malice. Plenty of legitimate businesses use IDNs.
Bottom line: Chromium's defenses work as designed, but each layer has a clear gap. The whole-script check is beaten by one unlisted letter, the skeleton comparison by a hook or bar that isn't stripped, and the navigation warnings by two edits or a short brand name. Chromium has closed breakers before, including ҏ and ӿ in Chrome 148 and adding м to the list in September, so a fix for ө seems likely eventually. Until then, a convincing fake can look identical to the real site in the address bar.
References
- Two characters open up a world of typosquatting opportunities in Chromium browsers The Register · 2026-10-10T10:15:00+00:00
- Chromium Docs - Internationalized Domain Names (IDN) in Google Chrome chromium.googlesource.com
- Turning IDN edge cases into typosquats - Have I Been Squatted haveibeensquatted.com