Models

OpenBMB Releases 324M Draft Model for MiniCPM5-2B

OpenBMB has launched MiniCPM5-2B-DSpark, a 324-million-parameter draft model designed to significantly accelerate inference speeds for its 2.52-billion-parameter MiniCPM5-2B model.

AlphaSignal3 days agoModels
Image: AlphaSignal

OpenBMB has introduced MiniCPM5-2B-DSpark, a 323.8-million-parameter speculative decoding draft model designed to accelerate its 2.52-billion-parameter MiniCPM5-2B target model. Under this setup, the smaller draft model proposes seven-token blocks, which the larger target model then verifies in a single parallel pass. This architecture is served via SGLang under the Apache 2.0 license, requiring the speculative algorithm flag set to DSPARK with a block size of seven.

The DSpark framework uses a semi-autoregressive drafting method to solve the acceptance decay typical of parallel draft models. It combines a parallel backbone with a lightweight sequential head to preserve information across the proposed block. Additionally, a confidence scheduler dynamically adjusts the number of proposals based on estimated prefix survival and the serving engine's throughput. The draft model was trained on 7.05 billion tokens over six epochs, utilizing cross-entropy, L1, and confidence loss objectives.

In practice, the draft model averages 5.52 accepted tokens per forward pass of the target model during greedy decoding, and 4.05 tokens at a temperature of 1.0. Performance scales even higher in specialized domains, reaching 6.05 accepted tokens in math and 6.11 in coding per verification step. The base MiniCPM5-2B model itself features 42 layers, grouped-query attention, and a 131,072-token context window. It scores an average of 53.9 across 34 evaluations, outperforming Qwen3.5-4B's score of 51.1, while hitting 69.1 on LiveCodeBench v6 and 97.1 on τ2-Bench Telecom.

For AI practitioners, this release offers a highly efficient way to deploy MiniCPM5-2B without sacrificing accuracy. Because the target model uses a standard LlamaForCausalLM architecture, developers can easily load it into existing inference engines without needing custom attention kernels. The integration of speculative decoding via SGLang allows developers to achieve much higher throughput and lower latency for token-heavy applications like math tutoring or code generation, where the draft model's high acceptance rates are most pronounced.

This is our own summary of reporting by AlphaSignal

More in Models