NVIDIA DGX Station for Windows: GB300 Specs and AI Workloads
NVIDIA is expanding its deskside AI computing portfolio beyond Linux with DGX Station for Windows, a system scheduled to launch in the fourth quarter of 2026. Powered by the GB300 Grace Blackwell Ultra superchip, the workstation is designed to bring data-center-class AI computing capabilities directly into Windows-based enterprise environments.
The announcement raises an important question: Why is NVIDIA bringing a Windows edition to a product line traditionally associated with Linux AI development?
The answer goes beyond operating-system compatibility. Enterprise developers often rely on Windows for productivity applications, business systems, and daily workflows, while turning to Linux for demanding AI workloads. Supporting Windows natively could reduce the friction between these environments and make local AI development and deployment more accessible.
With up to 748 GB of combined CPU-GPU memory, advertised AI performance of up to 20 PFLOPS at FP4 precision, Windows Subsystem for Linux (WSL), and support for persistent AI agents, DGX Station for Windows targets organizations that want to develop and run large AI models locally without depending entirely on remote infrastructure.
🪟 Why NVIDIA Is Bringing DGX Station to Windows #
NVIDIA’s decision reflects a broader shift in enterprise AI: AI computing is moving beyond centralized training clusters and cloud-hosted inference toward local development, experimentation, and agent execution.
Bridging the Windows-Linux Divide #
Linux has long been central to AI infrastructure because of its mature GPU software ecosystem, server tooling, and support for distributed computing. Windows, meanwhile, remains deeply embedded in enterprise desktops and developer workflows.
This separation can create operational friction. A developer may use Windows for business applications, documentation, and collaboration, but need a separate Linux workstation or remote server to run large models.
DGX Station for Windows aims to bring these activities closer together. By supporting AI workloads within an existing Windows environment, it can reduce the need to switch machines or maintain entirely separate development environments.
Bringing Local AI Into Enterprise Workflows #
Local AI computing offers several potential advantages:
- Data control: Sensitive datasets can be processed on premises, subject to the organization’s security and governance controls.
- Reduced network dependence: Development and inference can continue without sending every request to a remote AI service.
- Predictable resource access: Teams can reserve local compute resources for experiments and internal applications.
- Workflow integration: AI tools can operate alongside Windows applications, development environments, and enterprise software.
These benefits depend on how the system is deployed. Local execution does not automatically guarantee data privacy, lower costs, or better performance for every workload.
🧠 GB300 Grace Blackwell Ultra: Hardware Architecture #
DGX Station for Windows is built around NVIDIA’s GB300 Grace Blackwell Ultra desktop superchip. Its architecture combines a Grace CPU and a Blackwell Ultra GPU, connected through a high-bandwidth chip-to-chip interconnect.
Grace CPU and Blackwell Ultra GPU #
The system’s Grace CPU features 72 cores based on Arm’s Neoverse V2 architecture. It works alongside the Blackwell Ultra GPU to handle CPU-side processing, model orchestration, data preparation, and accelerator-intensive AI operations.
The CPU and GPU communicate through NVLink-C2C, which provides up to 900 GB/s of interconnect bandwidth according to the announced specifications.
This high-bandwidth connection is important for workloads that need to exchange data between CPU and GPU memory. It also supports the system’s unified-memory design, allowing large models and datasets to be managed across the two memory domains.
Hardware Specifications at a Glance #
| Component | Announced specification |
|---|---|
| AI platform | NVIDIA GB300 Grace Blackwell Ultra |
| CPU | 72-core Grace CPU, Arm Neoverse V2 |
| CPU-GPU interconnect | NVLink-C2C |
| Interconnect bandwidth | Up to 900 GB/s |
| GPU memory | 252 GB HBM3e |
| HBM3e bandwidth | 7.1 TB/s |
| System memory | 496 GB LPDDR5X |
| System memory bandwidth | 396 GB/s |
| Combined memory capacity | Up to 748 GB |
| AI compute performance | Up to 20 PFLOPS at FP4 |
| Networking | ConnectX-8 SuperNIC |
| Maximum network speed | Up to 800 Gb/s |
| GPU partitioning | Multi-Instance GPU (MIG), up to 7 instances |
| Planned availability | Q4 2026 |
These are announced platform specifications. Actual performance, supported configurations, and application-level throughput will depend on the final system and workload.
💾 Understanding the 748 GB Unified Memory Architecture #
One of the system’s most notable features is its combined memory capacity of up to 748 GB. However, this figure needs to be interpreted correctly.
It consists of two distinct memory pools:
- 252 GB of HBM3e GPU memory, with 7.1 TB/s of bandwidth.
- 496 GB of LPDDR5X system memory, with 396 GB/s of bandwidth.
Together, these provide a large CPU-GPU memory environment. The combined figure does not mean the GPU has 748 GB of standalone HBM VRAM.
Why Unified Memory Matters for Large Models #
Large AI models can require substantial memory for model weights, intermediate activations, runtime state, and associated data. When a model does not fit comfortably in dedicated GPU memory, developers may need to distribute it across devices, offload data to system memory, or use quantization and other memory-saving techniques.
A platform that supports coordinated CPU-GPU memory access can make these workloads easier to manage.
Potential benefits include:
- Loading larger models than would fit within the GPU’s HBM capacity alone.
- Managing large datasets and model state within a single workstation.
- Reducing some of the complexity associated with explicit data movement.
- Supporting local inference, fine-tuning, and long-running AI applications.
However, unified addressability does not make all memory physically identical. HBM3e and LPDDR5X have different bandwidth and performance characteristics. Data accessed through system memory may behave differently from data resident in GPU HBM, depending on the access path and workload.
For performance-sensitive inference, memory placement, quantization, batching, and runtime optimization remain important.
🚀 Up to 20 PFLOPS: What It Means for Local AI #
NVIDIA specifies up to 20 PFLOPS of AI compute at FP4 precision and positions the system for models with up to one trillion parameters.
FP4 is a four-bit floating-point format designed to reduce the memory and computation requirements of supported AI workloads. Lower-precision computation can increase throughput and improve the amount of model data that fits within a given memory budget.
However, the advertised peak performance is not equivalent to application throughput. Real-world results depend on model architecture, numerical format support, software kernels, memory bandwidth, and workload characteristics.
Large Language Model Inference #
The system’s memory capacity and accelerator performance make it relevant to local LLM inference, particularly for organizations evaluating models that are too large for conventional workstations.
Model size alone does not determine whether a model will run efficiently. Quantization, runtime overhead, context length, key-value cache requirements, and concurrency all affect memory consumption and throughput.
Fine-Tuning and Model Development #
DGX Station for Windows is also positioned for model experimentation and fine-tuning. A large local memory pool can help researchers test larger models and datasets without immediately moving workloads to a remote cluster.
The practical limits will depend on the training method. Full-parameter training generally has much greater memory and compute requirements than inference or parameter-efficient fine-tuning.
Long-Running AI Agents #
AI agents can maintain context, invoke tools, process files, execute tasks, and repeat inference cycles over extended periods. These workloads may benefit from dedicated local compute resources, especially when agents need to operate alongside desktop applications or internal systems.
Persistent agents also introduce additional requirements around permissions, process isolation, monitoring, and resource management.
🛡️ Windows, WSL, and NVIDIA OpenShell #
Hardware alone does not explain why a Windows edition matters. Its software integration is a central part of the product’s positioning.
Windows Subsystem for Linux #
Windows Subsystem for Linux (WSL) allows developers to run supported Linux distributions and tools within Windows without maintaining a separate physical Linux workstation.
For AI development, this can provide access to familiar Linux-based workflows while retaining Windows applications and services.
A WSL-based development environment can support activities such as:
- Running Linux-oriented development tools.
- Managing Python environments and AI frameworks.
- Working with containerized applications.
- Integrating code editors with Linux project environments.
- Connecting local AI services to Windows-based workflows.
The exact GPU capabilities and supported frameworks depend on NVIDIA’s final drivers, software stack, and compatibility matrix.
NVIDIA OpenShell and Persistent Agents #
NVIDIA also positions DGX Station for Windows to support NVIDIA OpenShell on Windows, using Windows security and isolation mechanisms to host AI agents.
This points toward a use case beyond one-off model inference. Instead of simply submitting a prompt and receiving a response, organizations can run agents that remain active, invoke tools, and perform repeated tasks.
For enterprises, persistent agents require careful control over what they can access and execute. Isolation, permission boundaries, auditability, and resource limits are essential when agents interact with business data or operating-system resources.
The ability to run agents locally may help organizations keep more processing within their own environment, but security still depends on configuration, application permissions, and the safeguards applied to each workload.
🏢 Enterprise Deployment: Why Deskside AI Matters #
DGX Station for Windows is positioned between a conventional workstation and a centralized AI cluster. It offers substantial local computing resources while retaining the desktop-oriented workflow familiar to enterprise developers and researchers.
Local Development and Evaluation #
Teams can use a deskside system to prototype AI applications, evaluate models, test inference pipelines, and perform selected fine-tuning tasks before deploying workloads to production infrastructure.
This can reduce dependence on shared remote resources during experimentation, although local systems still require capacity planning and operational management.
Data-Sensitive Workloads #
Organizations in finance, healthcare, manufacturing, and other regulated sectors may prefer to process certain data within controlled environments.
A local AI workstation can support that objective, provided the deployment also meets applicable requirements for access control, logging, encryption, retention, and data handling.
A Complement to Cloud and Data Center AI #
DGX Station for Windows should not be viewed as a universal replacement for cloud GPU instances or large AI clusters.
Large-scale distributed training, high-concurrency production inference, and workloads requiring extensive multi-node resources may still be better suited to centralized infrastructure.
A more practical deployment model is to use deskside systems for local development and selected inference tasks, then move larger or more demanding workloads to shared clusters when necessary.
🔍 Why This Product Matters to NVIDIA’s AI Strategy #
The Windows edition extends NVIDIA’s AI platform strategy into an environment where many enterprise users already spend most of their working day.
Its significance comes from combining three elements:
- High-capacity local hardware: The GB300 platform provides substantial accelerator performance and a large combined memory pool.
- Integration with Windows workflows: WSL and Windows application compatibility can reduce friction between conventional development and AI computing.
- Support for persistent AI agents: Local agent execution expands the role of the workstation beyond interactive model testing.
Together, these elements reflect a broader shift in AI infrastructure: increasingly capable models are being developed and deployed across a spectrum of environments, from cloud data centers to local workstations and edge systems.
The challenge is no longer just providing enough compute. It is making that compute accessible to developers and organizations through the tools, security models, and workflows they already use.
🚀 Conclusion #
NVIDIA’s decision to introduce DGX Station for Windows reflects the growing demand for local AI computing within mainstream enterprise environments. By combining the GB300 Grace Blackwell Ultra platform with up to 748 GB of combined CPU-GPU memory, up to 20 PFLOPS of FP4 AI performance, WSL support, and an agent-oriented software direction, the system targets workloads that benefit from substantial local compute capacity.
The key distinction is that 748 GB represents combined HBM3e and LPDDR5X memory, not standalone GPU VRAM. Likewise, peak FP4 performance does not guarantee equivalent throughput across all AI models.
If the planned Q4 2026 launch delivers the promised hardware and software integration, DGX Station for Windows could give enterprise developers a more direct path to local model development, inference, and AI agent deployment—without requiring them to abandon their existing Windows workflows.