NVIDIA Expands Local AI on RTX and DGX With Nemotron 3.5, LTX 2.5 and New Open Models

NVIDIA is expanding its local AI ecosystem across GeForce RTX PCs, RTX PRO workstations, DGX Spark, and DGX Station with support for a new wave of open models covering agentic AI, coding, video generation, robotics, multimodal reasoning, and local model training. The latest rollout includes Nemotron 3.5 Lightning, Meta Muse Glimmer, LTX 2.5, Cosmos 3 Edge, MiniMax H3, Laguna S 2.1, DeepSeek V4 Flash, Inkling Small, Wan Animate 2, and the upcoming Unsloth Desktop application, strengthening NVIDIA's strategy of positioning its hardware as a complete platform for private AI workloads outside traditional cloud infrastructure. NVIDIA says the models are being optimized across its Blackwell hardware portfolio with dedicated precision formats, inference frameworks, and local deployment tools.

One of the biggest additions is NVIDIA Nemotron 3.5 Lightning, a customizable 30B mixture of experts model designed for always active agents and specialized tasks inside larger AI systems. NVIDIA claims the model delivers up to 4x faster token generation and approximately 30% faster agentic task completion than other open models in its class. It supports local deployment across RTX PCs, DGX Spark, DGX Station, Jetson, RTX PRO workstations, data centers, and cloud infrastructure, with support from vLLM, Ollama, llama.cpp, LM Studio, and Unsloth. NVIDIA is also launching NeMo Switchyard, an open source model routing library that can automatically direct individual agent tasks toward different models according to accuracy, speed, latency, and cost requirements. Internal NVIDIA testing claims this approach can maintain frontier level task completion while reducing completion cost to approximately one third of using Opus 4.8 alone.

Meta Muse Glimmer is also receiving extensive NVIDIA optimization. The 30B dense open weight model is designed for local coding and long running agentic workloads and can operate entirely on a single consumer GPU. NVIDIA reports more than 200 tokens per second on a GeForce RTX 5090, while Meta's own optimized 17 GB configuration reaches 233.4 tokens per second using DFlash speculative decoding.

Creative AI is receiving a major upgrade through LTX 2.5. The new open video generation model introduces multishot generation, generative editing, stronger character and scene consistency, improved prompt adherence, and a new diffusion video decoder designed to improve final image quality. NVIDIA says optimizations including NVFP4, FastVideo, and ComfyUI improvements allow an RTX PRO 6000 Blackwell GPU to deliver up to 20% higher performance while reducing memory requirements by 40%. LTX 2.5 is optimized for RTX GPUs, DGX Spark, and DGX Station, allowing creators to run increasingly complex video generation workflows locally rather than depending entirely on cloud services.

NVIDIA is also accelerating a broader group of open models. Cosmos 3 Edge is a 4B world model aimed at robotics, autonomous vehicles, and vision AI that can operate directly on DGX Spark and Jetson. Poolside's 118B Laguna S 2.1 coding model receives an NVFP4 checkpoint capable of running on a single DGX Spark, while DeepSeek V4 Flash brings a 284B mixture of experts architecture with 13B active parameters and a 1 million token context window to DGX Station through community GGUF versions. Thinking Machines Lab's 276B Inkling Small activates 12B parameters per token and can run on either 1 DGX Station or 2 DGX Spark systems.

Wan Animate 2 demonstrates the advantage NVIDIA is targeting in generative media workloads. The 14B model can transfer movement and facial expressions from an existing video onto a static character image, including humans, animated characters, robots, and animals. NVIDIA reports generation performance up to 26x faster on an RTX 5090 and 16x faster on an RTX PRO 5000 Blackwell compared with an Apple M3 Ultra. Unsloth Desktop is also joining the local ecosystem with an open source desktop application combining model inference, image and video diffusion, fine tuning, agent integration, web research, and code execution.

DGX Spark itself is gaining new functionality through NVIDIA Sync Cluster Assistant. Developers can connect 2 or more DGX Spark systems using their ConnectX 7 interfaces, with NVIDIA Sync automatically configuring networking, distributing workloads, and monitoring system health. This effectively allows larger models to access additional memory and compute resources without manually configuring a distributed environment. Google Chrome is also coming to DGX Spark as a native ARM64 Linux build later in August, alongside a new Resource Monitor offering real time and historical CPU and GPU utilization across individual systems or complete Spark clusters. NVIDIA's continued expansion of local agents follows its broader work around DGX Station, Nemotron 3 Ultra, and NemoClaw, extending the same local AI strategy from consumer RTX systems into much larger professional platforms.

NVIDIA's local AI strategy is becoming much larger than simply accelerating language models on GeForce GPUs. The company is building an ecosystem where the same development tools and open models can scale from a gaming PC to DGX Spark, professional workstations, DGX Station, Jetson, and eventually data center infrastructure.

That interoperability could become one of NVIDIA's strongest advantages. A developer can prototype an agent on an RTX PC, move larger workloads onto DGX Spark, and scale into enterprise infrastructure without completely changing the underlying software environment. With models such as Muse Glimmer, Nemotron 3.5 Lightning, and increasingly capable video generators fitting onto local hardware, the AI PC battle is moving beyond dedicated NPUs and toward which platform can actually run the widest range of useful models efficiently.

Would you use your RTX gaming PC for local AI agents and video generation, or would you prefer dedicated hardware such as DGX Spark for these workloads?

Share
Angel Morales

Founder and lead writer at Duck-IT Tech News, and dedicated to delivering the latest news, reviews, and insights in the world of technology, gaming, and AI. With experience in the tech and business sectors, combining a deep passion for technology with a talent for clear and engaging writing

Previous
Previous

Intel Razor Lake AX Appears in HWiNFO as Future Halo Class SoC Takes Shape

Next
Next

Sony Reveals Battle Yellow PS5 for Marvel’s Wolverine With Matching DualSense and Adamantium PS5 Pro Design