AMD IOMMU PerfOpt targets Radeon iGPU AI workloads
AMD queued a Linux 7.4 IOMMU PerfOpt patch that bypasses IOMMU translation for integrated graphics, with a headline AI/LLM gain of 18 to 23 percent.

The supplied Phoronix page does not show the benchmark table behind that range. It names the mechanism: AMD PerfOpt, an IOMMU performance optimization that bypasses IOMMU translation when integrated graphics access system memory directly.
The constrained path is the iGPU memory route, not the silicon. The document proves a kernel option exists, default on for integrated graphics, and can be disabled with amdgpu.iommu_perfopt=0. It does not prove that 18 to 23 percent holds for every Ryzen AI SKU, every model size, or every inference workload.

Still unmeasured are the exact models, prompt lengths, batch sizes, power draw, and latency deltas. The write-up also leaves out whether the gain survives outside the Lemonade and Llama.cpp test path, or under sustained load and thermal behavior.

The constrained path is the iGPU memory route, not the silicon. The headline range is a test result, not a silicon spec.
The document leaves out the benchmark table, model sizes, prompt lengths, batch sizes, power draw, and latency deltas.
PerfOpt removes IOMMU translation overhead for iGPU system-memory accesses. The 18 to 23 percent figure is workload-specific and should be treated as a kernel-path result, not a hardware uplift.
Linux 7.4 · amdgpu.iommu_perfopt=0 · AMD IOMMU Performance Optimization · Strix Halo · Strix Point · Lemonade local AI server
After Phoronix. We did not report this. The pictures, if any, are theirs.


