Posts tagged "Vllm"

1 post

July 31, 2026
Serving DeepSeek-V4-Flash-0731 in Tensor Parallel Across Two DGX Sparks
DeepSeek shipped the official V4-Flash-0731 release today, folding its speculative-decoding module into the base checkpoint and beating V4-Pro-Preview on agentic benchmarks despite a far smaller activated parameter count. Here's what's actually inside the checkpoint (mixed FP8/FP4, not the plain FP8 config.json advertises), whether 155 GiB fits across two 121 GiB unified-memory GB10 boxes, the vLLM tensor-parallel launch recipe over the RoCE link, and the JIT-warm-up trap that will silently wedge a multi-node deployment on its first real request.