RDNA 4's FP8 units now run NVIDIA's DLSS 4.5 kernels
Countervolts' d4r 0.1.2 ships native FP8 DLSS 4.5 kernels for gfx1201, but the path has only been verified through emulation, not on a physical RX 9070 XT.

The industry assumption was that NVIDIA's DLSS stack was a closed moat, locked to Tensor Cores and proprietary CUDA. Countervolts' d4r 0.1.2 breaks that assumption by shipping native FP8 inference kernels for RDNA 4's gfx1201 and gfx1200 targets, running NVIDIA's official nvngx_dlss.dll through ZLUDA and ROCm on Linux and Proton.
A translation layer, not a replacement. The project loads NVIDIA's unmodified DLSS DLL, then routes inference through custom Radeon kernels built on ROCm. For gfx1201 hardware—the RX 9070 XT, RX 9070, RX 9070 GRE, and Radeon AI PRO R9700—the kernels exploit RDNA 4's native FP8 matrix operations, with a NativeFp8 flag enabled by default. Disabling it falls back to 16-bit math, the same path the RDNA 3 implementation uses. gfx1200 builds for the RX 9060 XT and RX 9060 are included but have only been compiled, not executed.
The only performance data in the document comes from an RX 7700 XT running SILENT HILL Townfall at 2560×1440, Quality preset. DLSS 4.5 moved from 48.4 to 49.3 FPS between d4r 0.1.1 and 0.1.2; DLSS 4 went from 67.9 to 68.1; DLSS 3 from 71.2 to 71.9. The sub-one-FPS deltas sit inside measurement noise. What the numbers do show is that the ZLUDA-to-ROCm translation stack adds a fixed overhead that does not scale with model complexity, at least on RDNA 3 silicon.
The constraint here is not AMD's hardware. The gfx1201 path has been checked through emulation only; no physical RX 9000 card has produced a frame-time or image-quality result. A buyer evaluating an RX 9070 XT for a DLSS 4.5 workflow is being asked to trust an emulated FP8 pipeline that has never touched the silicon it targets. For AMD, the FP8 matrix units were marketed -inference differentiator; this project repurposes them as a compatibility shim for a rival's software stack, which is a different workload profile than the LLM inference AMD designed the units for.
What remains unmeasured: whether FP8 execution on physical gfx1201 matches the emulated latency, whether the ZLUDA translation overhead holds at higher resolutions or with larger DLSS 4.5 network variants, and whether output image quality matches NVIDIA's reference on RTX hardware. The project is also limited to Super Resolution. Frame Generation, Ray Reconstruction, and DLSS 5 Neural Rendering are absent. Watch for a physical-card benchmark from Countervolts or an independent tester before treating the gfx1201 FP8 path as production-ready.
The FP8 matrix units on RDNA 4 were designed for LLM-scale inference, not for the small, latency-sensitive super-resolution networks DLSS 4.5 runs; this project repurposes them as a compatibility shim, and the emulated-only verification means no buyer can yet confirm the silicon path holds. AMD's competitive position on FP8 is now defined partly by whether a third-party translation stack can make its units run a rival's proprietary models.
gfx1201 FP8 kernels are emulated-only; gfx1200 is compile-only. The RDNA 3 reference (RX 7700 XT) shows a fixed ~1 FPS overhead from the ZLUDA/ROCm stack that does not grow with model size. No physical RX 9000 frame-time or PSNR data exists. The FP8 path is a different tensor-shape and throughput profile than the AI workloads AMD's marketing targeted.
d4r 0.1.2 targets gfx1201 and gfx1200 · NativeFp8 enabled by default on RDNA 4 · RX 7700 XT: DLSS 4.5 at 49.3 FPS, 2560x1440 · Super Resolution only; no FG or Ray Reconstruction
After videocardz.com. We did not report this. The pictures, if any, are theirs.


