Comparing prices across every provider we track...
AI Endpoints · $0.26/mo
Vultr's Llama-3.1-Nemotron offering is dramatically cheaper at $0.01/mo with identical token pricing, though it runs a different (smaller, safety-focused) model. OVHcloud's Qwen3-Coder-30B is a larger, more capable coding model, but you'll pay 26x more for the base fee and face hidden costs on IPv4, Windows licenses, and APAC egress after mid-2026.
Identical $0.01/M token pricing on both input and output, with a 26x lower base fee and no region-based egress cliffs.
Caveat: Model is 8B (not 30B), optimized for safety guardrailing rather than coding; bandwidth overages vary by data center and Windows licensing costs extra.
Before you switch