Nvidia ships a 700B model that fits in a single A100 — for real this time
Nvidia dropped DeepSeek-V4-Pro-0813-NVFP4 on Hugging Face, and it does what the big models have been promising since last year: a 700-billion-parameter model that actually runs on one A100 GPU. The trick is NVFP4, a compressed 4-bit format Nvidia built for this. Without it, the model needs eight 80GB GPUs. With it, you need one.
DeepSeek V4 was already the cheapest open-weights model of its class — better than Llama 405B on most benchmarks, priced at a fraction of what GPT-4o or Claude 3.5 charge per token. Now Nvidia's quantization makes it something a single small shop can actually host. A one-A100 setup is still not a toy; that card runs you $7k–$10k new. But it's a fraction of a full 8-GPU cluster and a fraction of cloud API fees at scale.
Why this matters for us:
For the primos running small shops and the aunts who post on Facebook, this means the models that used to live in Silicon Valley data centers are finally moving onto hardware that fits in a server closet — and that changes who gets to build with them.
“A 700B model that fits in one A100 — that's the shift, not the benchmark.”