NVIDIA and OpenAI launched the gpt-oss-20b and gpt-oss-120b models optimized for accelerated performance on NVIDIA Blackwell architecture, achieving up to 1.5 million tokens per second. Trained on NVIDIA H100 Tensor Core GPUs, the models utilize RoPE with 128k context and are released in FP4 precision, fitting on a single 80 GB data center GPU.

The models incorporate features like TensorRT-LLM Gen, CUTLASS MoE kernels, and XQA kernel for enhanced performance. NVIDIA collaborated with OpenAI and the community to maximize efficiency, offering optimized kernels through the FlashInfer library. vLLM verification and analysis were done for Hopper and Blackwell architectures, providing high-performance options for developers.

Optimizations for the gpt-oss models are available on the NVIDIA/TensorRT-LLM GitHub repository, along with a deployment guide for setting up servers with high performance and low latency. The models can be deployed using the Python API in a JupyterLab notebook as an NVIDIA Launchable in the build platform, allowing easy testing across multiple cloud platforms with a single click deployment option.

NVIDIA engineers ensured accelerated performance on the NVIDIA Blackwell and Hopper platforms for the gpt-oss-120b and gpt-oss-20b models, delivering up to 1.5 million tokens per second on a GB200 NVL72 system. The NVIDIA ecosystem enables high throughput and cost-efficient deployment of the latest models on Day 0.

NVIDIA Dynamo, an open-source inference serving platform, aids in deploying models like gpt-oss for large-scale applications, offering LLM-aware routing and elastic autoscaling features. Run locally on NVIDIA GeForce RTX AI PCs for faster iteration and lower latency, with models supported on professional workstations powered by NVIDIA RTX PRO GPUs.

Enterprise developers can access the gpt-oss models for free using the NVIDIA NIM Preview API and web playground environment in the NVIDIA API Catalog. Packaged as NVIDIA NIM microservices, the models can be easily deployed on GPU-accelerated infrastructure with flexibility, data privacy, and enterprise-grade security.

Read more at NVIDIA: Delivering 1.5 M TPS Inference on NVIDIA GB200 NVL72, NVIDIA Accelerates OpenAI gpt-oss Models from Cloud to Edge