ai_explainer_worthySeptember 1, 2026Issue #101

Nvidia ships DeepSeek V4 in FP4 — 37B params, half the size, 70B quality

Nvidia is releasing a quantized version of DeepSeek V4 Pro under the Hugging Face name DeepSeek-V4-Pro-0813-NVFP4. The model has 37 billion parameters and runs in 4-bit floating point (FP4). The claim is that it delivers performance roughly on par with a 70B model while using far less memory and compute.

Quantization is one of those moves that turns big open models from lab curiosities into something you can actually run. An 800B model needs a full server rack; 37B in FP4 can sit on a single GPU. That opens up a world of inference for shops and tinkerers who don't have datacenter budgets — the kind of setups that run on a 4090 or two.

Nvidia's involvement is notable for its own reasons. The company is pushing NVFP4 as its next-generation quant format and is making it available on Hugging Face so anyone can use it. The model itself was developed by DeepSeek, not Nvidia, so this is a partnership-style release rather than a first-party model.

Why this matters for us: the open models that actually fit on one GPU are the ones that small shops, local orgs, and side-hustle devs can run without begging for cloud credits.

800B needs a rack. 37B in FP4 fits a 4090 — that's the difference between a lab toy and something you actually run.

arxiv.org

Read the originalOpen in new tab
#quantization#fp4#deepseek#nvidia#open_models

Daily issue · no spam

Get the daily on your stoop

One short email a day — AI, tech, and what it means for our communities. Plain language, cultural lens, no Silicon Valley jargon.