Comparing prices across every provider we track...
Serverless Inference · $0.01/mo
Vultr's Llama-3.1-Nemotron pricing is on par with OVHcloud's AI Endpoints at the headline rate, but both carry material hidden costs that vary by region and use case. Neither is clearly cheaper; your actual bill depends heavily on bandwidth consumption, data center choice, and whether you need Windows licensing.
Identical headline pricing ($0.01/mo, $0.01/M input tokens) with fp16 quantization support, but serves a different model class (embedding vs. LLM inference).
Caveat: APAC regions lose free egress after mid-2026; IPv4 and Windows surcharges apply; not a direct model replacement for Nemotron.
Before you switch