What PerfOpt actually does
PerfOpt is the AMD IOMMU Performance Optimization from the AMD I/O Virtualization Technology specification. It lets integrated graphics bypass the IOMMU when accessing system memory directly, which avoids that overhead. With the patches going into Linux 7.4, it will be enabled automatically on AMD iGPU setups. It can be forced off with the amdgpu.iommu_perfopt=0 module option, which is how you compare the impact.
Phoronix's earlier report said it is not specific to the high-end Ryzen AI Max parts. It can benefit AMD iGPUs at large, and it is meant only for integrated I/O devices, not discrete GPUs or other external hardware.
The prior-gen test
Last week's tests on current-gen Ryzen AI hardware showed gains of up to 18~23%. Readers asked about older chips, so Michael Larabel tested one. He used an older Lenovo ThinkPad P14s Gen 4 with the Ryzen 7 PRO 7840U and Radeon 780M integrated graphics.
The method was a within-machine A/B comparison. The laptop ran Ubuntu 26.04 LTS with a kernel built from the IOMMU-next Git tree. Nothing was changed beyond toggling PerfOpt.
The workload was local inference through the Lemonade AI server with a Llama.cpp backend. Some runs also used Llama.cpp directly. The results, as reported:
- Latency: The review's second page describes a visible improvement in time to first token in Lemonade. That is how long you wait before the model starts answering.
- Throughput: Token throughput rose only slightly with the Vulkan backend.
- Backends: Both Vulkan and ROCm benefited on Radeon iGPUs.
- Magnitude: Gains of roughly 4–12% were common across the tests on this one machine. That figure comes from my read of the full review's charts, not from the intro text above.
The models tested included Qwen3, MiniCPM4, DeepSeek-Qwen3 and gpt-oss-20b variants, with chat and code-style prompts.
How to read the numbers
- This is one laptop. It is one 7840U ThinkPad with one set of software. It is not a survey of Ryzen APUs.
- The generations aren't directly comparable. The 18~23% and 4–12% figures come from different machines, including a Strix Halo desktop plus Strix Point and Krackan laptops in the first round. Don't conclude that older chips benefit less or more.
- Latency gains are plausible. Direct DMA access mainly cuts memory-access overhead, and prompt processing and first-token time are sensitive to that. Throughput limited by memory bandwidth would be expected to move less. That is my inference from the results, not something Phoronix claims.
- Scope is narrow. Nothing here covers games, desktop responsiveness, CPU-only inference, non-AMD GPUs or Windows. One CachyOS feature request speculates about gaming gains, but it admits there are no benchmarks yet.
- Details are missing. The report doesn't spell out memory configuration, power profile, exact versions or repeat counts.
The security and compatibility trade-off
The kernel commit that wires PerfOpt into amdgpu is explicit about the cost. Arming PerfOpt trades IOMMU DMA containment for lower DMA latency. The commit describes the default as a deliberate, documented policy.
Per that commit:
- It applies when the GPU is in the identity domain. In that case the GPU is already doing direct DMA, and the IOMMU only enforces read/write permission bits with no address translation.
- Enabling it clears ATS, PRI, PASID and SVA for the device.
- It is gated on the device being an APU, since the AMD IOMMU spec says the feature is only supported on integrated GPUs.
- If the IOMMU doesn't implement the feature, the driver warns and continues rather than failing.
- The setting is per IOMMU and reference counted, so one GPU's teardown doesn't clear it while another device still needs it.
The patch history is also still moving. The commit in the IOMMU tree dated September 25 uses a default-on module parameter. A later proposal I can't verify from the commit alone reportedly aims to make the behavior non-optional for iGPUs in identity mode. It would also block VFIO/iommufd claims on that device. Treat the module-parameter details as accurate for the benchmarked code, not necessarily for the final 7.4 release or your distribution's kernel.
What to do about it
- Check your kernel. Linux 7.4 hasn't been released. Phoronix says its merge window opens in late October. Distribution kernels will arrive later. Early adopters such as CachyOS users are already asking for a backport.
- Run an A/B test yourself. On a kernel that has the feature, boot once normally and once with amdgpu.iommu_perfopt=0. Run your actual model and prompts, and compare time to first token and tokens per second.
- Audit your IOMMU use. If you depend on GPU passthrough, VFIO, PASID/SVA or strict DMA isolation involving the integrated GPU, check how your kernel handles PerfOpt first. The opt-out parameter is there for exactly this case.
- Mind the threat model. For a personal laptop running local models, trading some DMA containment for latency may be acceptable. A shared workstation or a hardened environment may not agree.
Bottom line
For Linux users with a Radeon 780M-class laptop, the evidence points to a free, workload-dependent improvement in local AI responsiveness, with the biggest effect on first-token latency. It comes with a real, documented loss of IOMMU containment for the iGPU, and the final upstream form is still being settled. Promising, then, but it needs a benchmark of your own and a quick look at your isolation requirements.
References
- AMD PerfOpt Boosting Prior-Gen Ryzen APU Performance For AI - Phoronix Phoronix · Fri, 09 Oct 2026 14:38:00 GMT
- Diff - refs/heads/amd/amd-vi^! - pub/scm/linux/kernel/git/iommu/linux - Git at Google kernel.googlesource.com
- AMD PerfOpt Boosting Prior-Gen Ryzen APU Performance For AI - Phoronix phoronix.com