DLSS 5 Modders Get Hybrid FP8 and NVFP4 Running on RTX 50 GPUs, but Gains Remain Tiny
NVIDIA DLSS 5 modders have taken another step toward reducing the heavy performance cost of Neural Rendering, successfully combining FP8 and NVFP4 processing on GeForce RTX 50 Series GPUs. The experimental implementation is now available through the community developed OptiScaler DLSS Neural Rendering project, but early testing shows that simply moving portions of the workload to Blackwell's lower precision NVFP4 format does not yet produce the large performance improvement some users expected.
Version 0.7.1 introduced the hybrid FP8 and NVFP4 model specifically for Blackwell GPUs. According to developer wilsjo2, measurements at 4K showed the complete Neural Rendering pass taking approximately 1% to 2% less time than the original FP8 path. Importantly, that number refers only to the time required to process DLSS 5 Neural Rendering and is not a 1% to 2% increase in overall game frame rate. The project itself describes the Blackwell improvement as very minor.
wow! huge update on optiscaler dlss 5 fork from wilsjo2. A hybrid FP8/NVFP4 model. According to a user that tested it with his 5080 this new hybrid model gave him a boost from 130fps to 165fps using dlss 5 🤯. Both opsticaler and renodx dlss 5 mods have seen so many new… pic.twitter.com/JJiUDZ0gGG
— SwurvGaming (@SwurvGaming) September 8, 2026
The experiment is technically interesting because NVFP4 is one of the major AI capabilities introduced with NVIDIA Blackwell. NVIDIA describes NVFP4 as a 4 bit floating point format designed for efficient low precision inference. It uses FP8 scaling across groups of 16 values together with an additional FP32 scaling factor, allowing considerably lower memory requirements while attempting to preserve the accuracy required by neural models. NVIDIA says NVFP4 can reduce model memory requirements by approximately 1.8 times compared with FP8.
That theoretical advantage does not mean an existing FP8 neural network can immediately double its performance when parts of it are converted to NVFP4. Neural models need to be specifically optimized and quantized around the lower precision format, and some operations may still perform better or retain better accuracy in FP8. The current OptiScaler implementation therefore uses a hybrid path rather than attempting to force the entire DLSS 5 workload into NVFP4.
The developer continued optimizing the approach with version 0.7.2, making the combined NVFP4 hybrid the recommended option for Blackwell hardware. A 120 second Baldur's Gate 3 test recorded 55.16 rendered FPS with the recommended hybrid model compared with 54.57 FPS using FP8 and 54.91 FPS with the previous hybrid implementation. That represents a difference of less than 1 FPS, and the developer specifically warns that the small average improvement has not yet been established as repeatable across multiple tests.
The result also puts some of the expectations surrounding NVFP4 into perspective. Blackwell's 5th generation Tensor Cores are designed to accelerate lower precision AI workloads, and NVIDIA has demonstrated substantial NVFP4 gains with models that have been properly optimized for the format. In one official example involving AI model training, NVIDIA measured NVFP4 performance between 1.31 times and 1.73 times faster than FP8 depending on the workload. DLSS 5, however, is a completely different real time graphics model with strict image quality and latency requirements, meaning those improvements cannot simply be transferred directly to Neural Rendering.
The community experiment arrives while NVIDIA is carrying out its own substantial optimization work. As previously covered, DLSS 5 is already approximately 5 times faster than the original GTC 2026 implementation, which initially required 2 GeForce RTX 5090 GPUs. The production implementation can now operate across RTX 50 Series hardware, although Neural Rendering can still impose a considerable performance cost depending on resolution and workload.
Modders have consequently been exploring several different ways to reduce that cost. One recent experiment moved DLSS 5 Neural Rendering onto a second RTX 5060 Ti, effectively treating another GPU as a dedicated neural processor. Other community work has moved the Neural Rendering stage before Super Resolution so that the AI model processes fewer pixels. The new FP8 and NVFP4 hybrid attacks the same problem from another direction by attempting to reduce the computational cost of the model itself.
NVIDIA has also confirmed that DLSS 5 support is coming to GeForce RTX 40 Series GPUs, although Blackwell remains the architecture with native NVFP4 capabilities. NVIDIA is concentrating first on additional RTX 50 optimization before expanding official Neural Rendering support to Ada Lovelace.
The important result here is not the extra 1% or the fraction of a frame gained in Baldur's Gate 3. It is that the modding community has already begun experimenting with precision levels inside the DLSS 5 workload itself.
NVFP4 may eventually become much more valuable for Neural Rendering, but this test demonstrates why hardware capability alone is not enough. A model designed around FP8 cannot necessarily extract the full benefit of 4 bit processing without deeper quantization, calibration and kernel optimization. NVIDIA has already achieved enormous performance gains by optimizing DLSS 5 at the model and pipeline level, so a future model specifically designed around Blackwell's NVFP4 hardware could produce a much more meaningful result than the current hybrid conversion.
For now, however, FP8 remains extremely competitive. The first community attempt at combining it with NVFP4 cuts the Neural Rendering pass by only around 1% to 2%, making this an important technical experiment rather than a major performance breakthrough.
Do you think a fully optimized NVFP4 version of DLSS 5 could significantly reduce the Neural Rendering performance hit on RTX 50 GPUs?
