The topic is Linux, not Windows. It still matters to anyone following Windows games on Arm, because the stack involved is Proton, Wine, DXVK and the FEX emulator. It also bears on any x86-on-Arm translation layer.
The problem: Arm64 says no where x86 says yes
The deck says x86 allows atomic instructions on unaligned addresses and Arm64 doesn't. An unaligned atomic on Arm64 raises a SIGBUS. The conference abstract says the process is then terminated. It adds that this is a problem for games because "we can't modify their source code to avoid unaligned atomic operations."
Emulation also makes the situation worse. x86 has a strong memory model and Arm64 has a weaker one. In the deck's words, FEX uses sync instructions (barriers, LL/SC, atomics) to get stronger memory ordering. The slides say this produces many more atomic instructions than the original code contained. Each one is another chance to hit an unaligned address.
What the 5x figure covers
The slide says handling the fault directly in the kernel is more than 5x faster than the current userspace route. That route traps the SIGBUS in FEX and fixes things up there, with a context switch each time.
Please treat this as a microbenchmark-style claim. It compares two ways of handling one fault. It is not a frame rate, a load time or a result for any named game. The deck as presented gives no hardware, workload or sample details that I could check. The deck itself was unavailable to me, so I am relying on the conference abstract and the reporting on it.
The proposal has a public history
The Linux kernel mailing list shows this isn't new. In November 2025 Almeida posted an RFC patch proposing kernel-side emulation of unaligned atomic instructions on Arm64. Applications would opt in through a new prctl flag, PR_ARM64_UNALIGN_ATOMIC_EMULATE.
The cover letter explains why userspace handling hurts. FEX uses code backpatching to avoid repeated faults. That can badly hurt performance if a function like memcpy gets patched, which the patch author says makes games like Mirror's Edge and Assassin's Creed: Origins unplayable without per-application configuration.
The earlier numbers were more modest. The RFC cites microbenchmarks showing a 4x decrease in overhead with kernel-side handling, and says the figure is even larger when FEX runs under Wine. The "more than 5x" from the Plumbers talk is consistent with that direction. I could not verify whether the two figures come from the same benchmark.
I also found no evidence in my search that this has been merged into mainline Linux. Treat it as a proposal and a design under discussion.
Split locks: the unsolved corner
The harder case is the split lock. This is an atomic operation that straddles two cache lines. The deck says x86 permits it and Arm64 has no instruction for it. The talk makes several points about it:
- On x86, the hardware issues a "bus lock" that blocks a whole cache. This adds latency for unrelated tasks, and Linux has a mitigation to stop abuse.
- Armv8.3+ extensions such as FEAT_LRCPC help emulate x86 ordering, and Apple Silicon implements full TSO emulation in hardware. None of them solve split locks.
- Splitting a 64-bit compare-and-swap into two 32-bit ones lets another thread slip in between. The deck labels the outcome "memory tearing."
The closing slide asks open questions, per the reporting. Should every other CPU be halted while one executes the pair? Should there be a giant lock in userspace? Should mainline support the use case at all?
What this means for readers
- Windows-on-Arm users: Nothing here changes Windows' own x86 emulation. The talk is about Linux, Proton and FEX.
- Steam Frame and Arm Linux gamers: A kernel fast path could reduce the cost of a known slow path. It is not a guaranteed uplift for any particular title.
- Developers: Native Arm64 builds sidestep the problem entirely, as the MIXED report notes with Factorio's Arm64 version.
I would caution against reading the headline as "5x faster games." The measured claim is about fault handling. The split-lock question is a policy debate about whether Linux should carry special support for emulating one x86 quirk.
References
- Igalia says the kernel can handle x86 atomics more than 5x faster than FEX does MIXED Reality News · 2026-10-09T05:41:47+00:00
- (RFC PATCH v2 0/1) arch: arm64: Implement unaligned atomic emulation lkml.iu.edu
- Linux Plumbers Conference 2026 (5-7 October 2026): Emulating x86 atomics on arm64 · Indico lpc.events