OpenDLSS: A Vulkan Reimplementation of Nvidia's DLSS 5 Neural Rendering Network
- ID
- 30895
- Status
- summarized
- Published
- 30 Sep 2026, 4:43 PM
- Fetched
- 02 Oct 2026, 2:35 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://github.com/maanHimself/OpenDLSS-NR
- Source URL
- https://hnrss.org/best
Summary
- Score
- 6.5
- Created
- 02 Oct 2026, 2:36 AM
- Tags
- Audience
- developersai_ml_learners
What happened
OpenDLSS-NR is a Vulkan reimplementation of NVIDIA's DLSS 5 Neural Rendering network that claims bit-exact output against DLSS-NR build 310.8.0, matching not just the final image but all 75 block boundaries byte for byte. The network is a 71-block shifted-window Swin/ViT U-net over six pooling levels, 141 MiB of weights, FP8 (E4M3) activations with FP16 accumulation, and it is not an upscaler — it re-renders an already-drawn frame at the same resolution. A second, independent implementation under ports/browser-webgpu/ runs the same bytes in a browser at 2048x1152 with no tensor cores and no FP8 support; users must supply their own weights. The repo has 702 stars, 60 forks and only 4 commits, with a 220-point / 103-comment Hacker News thread.
Why it matters
The browser WebGPU port is the concrete takeaway: the same network runs without tensor cores and without FP8, which means browser-side neural inference at 2048x1152 is demonstrably possible without the hardware features people assume are mandatory — worth testing before you default to server-side GPU inference for a rendering or post-processing feature. Note the practical limits before planning anything: you must supply your own weights (nothing is shipped), the repo is 4 commits deep, and the claims of byte-exactness come from the author, not an independent benchmark. There is no Malaysia or Southeast Asia angle in this item.
Discussion angle
Bit-exactness as a verification bar: matching all 75 block boundaries byte for byte is a much stronger claim than matching the final image — is that reproducible by anyone else, and what does it take to verify a reimplementation you can't easily diff against closed-source hardware behaviour?