Comparing prices across every provider we track...
AI Endpoints · $0.1/mo
Vultr's Llama-3.1 serverless inference option is substantially cheaper on both base price and per-token costs, with no region-based egress cliffs or separate IPv4 billing. However, the models are different (Llama vs Mistral), so verify the model choice meets your use case before switching.
10x cheaper monthly base ($0.01 vs $0.10), identical input token rate ($0.10/1M), and includes output tokens at no extra charge versus OVHcloud's input-only pricing.
Caveat: Different model family (Llama vs Mistral); verify Nemotron-Safety-Guard meets your accuracy and latency requirements, and confirm bandwidth overage rates for your data center.
Before you switch