NVIDIA's official COMPUTEX article card for RTX Spark and local AI agents
NVIDIA's official COMPUTEX article card for RTX Spark and local AI agents
+ NVIDIA AI News

NVIDIA announces RTX Spark PCs for local AI agents

RTX Spark puts 1 petaflop of AI performance and up to 128GB of unified memory into Windows PCs designed for local agents.

NVIDIA used GTC Taipei at COMPUTEX to announce RTX Spark, a new class of Windows PCs built for local AI agents. The company says RTX Spark systems will offer 1 petaflop of AI performance and up to 128GB of unified memory, with slim laptops and compact desktops expected this fall from ASUS, Dell, HP, Lenovo, Microsoft Surface, MSI, and others.

The point is not just another fast laptop. NVIDIA and Microsoft are trying to make the primary PC a place where agents can run locally with security controls, local model routing, and enough memory for larger workflows. That is the part to watch: if personal agents need to touch local files, apps, identity, and private context, the cloud-only pattern starts to hit limits.

1 petaflopAdvertised AI performanceNVIDIA
128GBMaximum unified memoryNVIDIA
120BLLM parameter class NVIDIA says can run locallyNVIDIA

Local agents need more than a GPU

NVIDIA says RTX Spark uses a Blackwell RTX GPU with 6,144 CUDA cores and fifth-generation Tensor Cores with FP4 precision, connected to a 20-core NVIDIA Grace CPU through NVLink-C2C. Those specs explain the performance claim. They do not explain the product strategy by themselves.

The strategy is security and control. NVIDIA says it is working with Microsoft on Windows security primitives for identity, containment, policy, and end-to-end security. NVIDIA OpenShell is meant to add policy controls, route queries to local models based on privacy settings, and disguise personal information when a request has to go to a cloud model.

That is the right problem to solve. A useful local agent needs permission to act across apps, files, and workflows. Without containment and policy, that becomes a security risk. Without enough local compute, it becomes a cloud proxy. RTX Spark is NVIDIA’s attempt to sell both the silicon and the agent runtime story together.

The workloads are concrete

NVIDIA says RTX Spark systems can render 90GB-plus 3D scenes, edit 12K 4:2:2 video, generate 4K AI videos, run 120-billion-parameter LLMs with up to 1 million tokens of context using agents locally, and play AAA games at 1440p over 100 frames per second. Those are broad claims, and each workload will depend on the model, app, and thermal design of the actual device.

The more useful read is that NVIDIA is pushing unified memory as the practical limiter. Local AI is not only about raw TOPS or benchmark charts. Large context windows, local image/video generation, and multi-app agents all want memory headroom. A 128GB unified-memory PC gives software teams a different target than a standard laptop GPU with a small VRAM ceiling.

NVIDIA’s blog also says OpenShell is coming to Windows, NemoClaw is expanding across GeForce RTX, RTX PRO, RTX, DGX Spark, and DGX Station, and llama.cpp and vLLM are getting multi-token prediction and multi-GPU optimizations for up to 2x inference performance on top agentic models.

Sources

The AI Feed Desk

The AI Feed Desk

Editorial desk

The AI Feed Desk tracks AI provider updates, model releases, agent tooling, and enterprise adoption, turning fast-moving announcements into source-linked context for builders and operators.

Noticed a typo, incorrect information, or translation error?

Tell us so we can fix it.

Help Improve This Article

Related Articles

Editorial illustration of a local text model generating many tokens in parallel on a GPU

Google releases DiffusionGemma for faster local text generation

Google's DiffusionGemma is an experimental open text-diffusion model that generates blocks of text in parallel for lower-latency local workflows.

The AI Feed Desk

By The AI Feed Desk

An open dataset map clusters agent workflow samples beside transparent synthetic persona cards

NVIDIA and Hugging Face publish open data for agents

NVIDIA and Hugging Face published a Nemotron data package for agents, including open pretraining data, post-training samples, an interactive Prompt Atlas, and synthetic persona datasets.

The AI Feed Desk

By The AI Feed Desk

A Japan manufacturing floor connects robotics arms, compact AI PCs, and data-center compute into one NVIDIA stack

NVIDIA uses Japan to package physical AI as a full-stack ecosystem

NVIDIA's July 15 Japan ecosystem update ties RTX Spark, robotics, manufacturing, and local partners into a physical AI stack.

The AI Feed Desk

By The AI Feed Desk

A transparent secure ledger collects agent activity traces from protected workspaces

Open Secure AI Alliance proposes SAFE guidelines for agent security findings

The Open Secure AI Alliance proposed SAFE guidelines for sharing agentic AI cybersecurity findings as Black Hat USA opened.

The AI Feed Desk

By The AI Feed Desk

A governed cloud workspace connects an AI model core to a high-performance compute rack

Claude reaches Microsoft Foundry with Azure governance and GB300 compute

Anthropic made Claude generally available in Microsoft Foundry, while NVIDIA framed the Azure deployment as a GB300 Blackwell Ultra agent platform.

The AI Feed Desk

By The AI Feed Desk