NVIDIA CUDA 13.4 Brings Native Windows on Arm Support Ahead of RTX Spark

NVIDIA has released CUDA Toolkit 13.4, expanding its accelerated computing platform to Windows on Arm for the first time while providing developers with preview support for the next generation Rubin GPU architecture. The software release arrives ahead of RTX Spark systems launching in October 2026 and establishes much of the development foundation needed for NVIDIA's new generation of Arm powered Windows PCs and future Vera Rubin AI infrastructure.

The most immediate change is native CUDA development on Windows Arm64. CUDA applications have supported Arm processors under Linux for years, but CUDA Toolkit 13.4 now extends that environment to Windows. Developers can begin porting applications, validating dependencies and testing CUDA software paths for Arm based Windows systems before RTX Spark hardware reaches consumers. NVIDIA's core mathematical libraries also gain Windows on Arm support for the N1X laptop ecosystem, while Nsight Systems 2026.5.1 adds visibility for CUDA 13.4 workloads running on the platform.

That timing directly supports RTX Spark, NVIDIA's new Windows platform combining a Grace Arm CPU with a Blackwell RTX GPU. The highest N1X configuration integrates a 20 core Grace CPU, 6,144 CUDA core Blackwell GPU, up to 128 GB of unified LPDDR5X memory and up to 1 petaflop of FP4 AI performance. NVIDIA is positioning the systems for local AI agents, content creation and gaming, while retaining technologies including DLSS 5, Reflex 2, DirectX 12 Ultimate and the broader CUDA ecosystem. NVIDIA confirmed at IFA that RTX Spark laptops and compact desktops are scheduled to begin arriving in October.

CUDA support is particularly important because RTX Spark represents a considerably different Windows architecture from conventional GeForce PCs. Developers targeting native Arm applications need software libraries, compilers and profiling tools that understand both the CPU architecture and NVIDIA GPU environment. CUDA 13.4 therefore gives the ecosystem time to begin moving applications toward Windows Arm64 before the hardware launch. Earlier RTX Spark driver support already established the initial development path, while CUDA 13.4 substantially expands the software stack around it.

The update is also the first CUDA release to provide functional preview support for NVIDIA Rubin, identified as compute capability 107. Developers can compile for the new SM 107 architecture and begin porting libraries and applications before Rubin reaches full CUDA general availability in a future toolkit release. This early access is particularly relevant for framework and library maintainers, since Rubin introduces significant changes across compute, memory and interconnect architecture.

NVIDIA's detailed Rubin GPU architecture disclosure shows why software preparation is beginning early. Rubin combines 2 reticle limited compute dies through the NVIDIA High Bandwidth Interface and integrates 336 billion transistors, 224 streaming multiprocessors and 896 Tensor Cores. Its third generation Transformer Engine delivers up to 50 petaflops of NVFP4 inference performance, with NVIDIA claiming up to 10 times more agentic throughput per unit of energy than Blackwell for its targeted workloads.

Each Rubin GPU supports up to 288 GB of HBM4 with as much as 22 TB per second of memory bandwidth, representing a 2.8 times bandwidth increase over Blackwell and Blackwell Ultra. NVLink 6 provides 3,600 GB per second of scale up GPU communication, while NVLink C2C delivers 1,800 GB per second between the CPU and GPU. PCI Express Gen 6 contributes another 256 GB per second of host connectivity. Those specifications are designed around increasingly long context reasoning, mixture of experts models and agent workloads where memory movement and communication can become as important as raw Tensor Core throughput.

The connection to CUDA 13.4 goes beyond simply recognizing the new GPU. NVIDIA has added CUDA Compute Fabric Transport, giving communication library developers a lower level method for moving data across NVLink fabrics through named logical endpoints and asynchronous operations. Multi Process Service V3 also provides more granular management of shared GPUs through scriptable controls, named server instances, streaming multiprocessor partitioning and GPU memory limits tied to Linux cgroups. These changes are increasingly relevant as individual accelerators are divided between multiple AI services or agent workloads.

CUDA Python receives another substantial expansion. cuda.core 1.1.0 adds texture and surface programming, improved managed memory controls, stronger CUDA graph integration and more complete type information, while cuda.compute 1.1 introduces ahead of time compilation for CCCL algorithms across multiple GPU architectures. CCCL 3.4 also improves Blackwell performance, with NVIDIA reporting that its new warp specialized DeviceScan implementation can reach up to 92% of available memory bandwidth.

Developer tooling is advancing alongside the runtime. Nsight Python 1.0 introduces Python based GPU kernel profiling, while Nsight Compute 2026.3 adds CUDA Tile IR visibility and Nsight Systems gains support for Rubin and Windows on Arm. NVIDIA is also introducing AI assisted CUDA development through Nsight AI, which connects coding agents to current CUDA documentation and examples through the NVIDIA hosted CUDA MCP Server.

NVIDIA has made the toolkit available through its CUDA download portal, which currently lists CUDA Toolkit 13.4.1 as the latest maintenance release. Developers preparing Windows on Arm software or beginning early Rubin validation can now access the updated compiler, libraries and tools directly from NVIDIA.

The release also strengthens the software groundwork behind NVIDIA's broader local AI strategy. RTX Spark is designed to run CUDA natively under Windows 11 while offering up to 128 GB of unified memory for large local models and persistent agents. NVIDIA recently expanded that strategy through PAIR, which can distribute local AI inference across multiple RTX PCs, while Vera Rubin extends the same broader CUDA ecosystem into massive rack scale AI infrastructure.

CUDA 13.4 is less about one headline feature and more about NVIDIA preparing its software ecosystem for 2 major hardware transitions at opposite ends of computing. Windows on Arm support gives RTX Spark a native development environment before the first systems arrive in October, while Rubin preview support gives AI framework developers time to prepare for an architecture designed around vastly larger memory systems, faster interconnects and sustained agentic workloads.

The strategic advantage is CUDA itself. RTX Spark can introduce a new Arm based PC architecture without asking developers to abandon NVIDIA's established programming environment, while Rubin can radically change the underlying GPU design and still remain accessible through the same software platform. That continuity is increasingly important as NVIDIA expands from discrete GPUs into tightly integrated CPU, GPU and unified memory systems across both personal computing and hyperscale AI.

Could native CUDA support make Windows on Arm a serious platform for AI developers and creators, or will x86 systems remain the preferred choice despite RTX Spark's unified memory and local AI capabilities?

Share
Angel Morales

Founder and lead writer at Duck-IT Tech News, and dedicated to delivering the latest news, reviews, and insights in the world of technology, gaming, and AI. With experience in the tech and business sectors, combining a deep passion for technology with a talent for clear and engaging writing

Previous
Previous

G.SKILL Brings AMD EXPO ULL to Trident Z5 Royal NeoX With DDR5 6000 CL26

Next
Next

Experimental OptiScaler Mod Cuts DLSS 5 Neural Rendering Performance Cost Nearly in Half