Models

Red Hat Halves Nemotron 3.5 Lightning Size with FP8

Red Hat AI has released an FP8-quantized version of NVIDIA's Nemotron 3.5 Lightning 30B model, lowering the hardware requirements for deploying the advanced agentic system.

AlphaSignal3 days agoModels
Image: AlphaSignal

Red Hat AI has launched an FP8-quantized checkpoint of NVIDIA's Nemotron 3.5 Lightning 30B A3B. By converting selected weights and activations from BF16 to FP8, this release cuts the storage and GPU memory footprint of the 30-billion-parameter model by roughly 50 percent. The base model features a hybrid architecture combining Mamba-2, Mixture of Experts, and Attention, utilizing 3 billion active parameters. It supports a massive 1-million-token context window, tool calling, switchable reasoning, and six languages.

To achieve this compression, Red Hat used LLM Compressor to apply static per-tensor FP8 quantization to weights and activations in supported linear operators. To preserve accuracy, several precision-sensitive components remain unquantized, including conv1d layers, embeddings, latent projections, mixture-of-experts gates, multi-token prediction layers, the final normalization layer, and the language-model head. Calibration was performed using 512 UltraChat samples of 2,048 tokens each. On performance benchmarks, the model scores 83.4 on PinchBench agent tasks, though it lags on coding and complex reasoning benchmarks. Early developer interest is high, with Hugging Face recording more than 780,000 downloads.

For practitioners, this optimization makes it possible to deploy the model on a single NVIDIA H100 GPU or a DGX Spark system using vLLM, SGLang, or Docker. While the quantization halves the size of the quantized tensors, overall GPU memory usage will decrease by less than 50 percent because of unquantized parameters, activations, Mamba state, attention key-value caches, and vLLM workspace allocations. Hopper and Blackwell GPUs can execute the FP8 matrix multiplication natively via Tensor Cores. However, actual throughput gains will vary based on batch size, sequence length, and memory bandwidth.

This is our own summary of reporting by AlphaSignal

More in Models