Comparing prices across every provider we track...
AI Endpoints · $0.18/mo
Vultr's serverless inference offering is substantially cheaper on both base price and per-token costs, with no hidden IPv4 or Windows license surcharges. However, the model (Llama-3.1-Nemotron-Safety-Guard-8B) is smaller and different from OVHcloud's gpt-oss-20b, so verify it meets your accuracy and capability requirements before switching.
Input and output tokens both cost 18x less ($0.01/M vs $0.18/M), with no separate IPv4 or license fees on top.
Caveat: Model is 8B parameters, not 20B; verify inference quality and latency meet your use case, and confirm bandwidth overage rates for your data center.
Before you switch