Comparing prices across every provider we track...
AI Endpoints · $0.1/mo
Vultr's Llama-3.1 serverless inference option is substantially cheaper on both base price and per-token costs, with no region-based egress penalties or separate IPv4 billing. However, the models are different (Llama vs Mistral), so verify the model choice meets your use case before switching.
10x cheaper monthly fee and 100x cheaper per-token input cost ($0.01/M vs $0.10/M), with no hidden IPv4 or region-based egress clawbacks.
Caveat: Different model family (Llama vs Mistral); verify Nemotron-Safety-Guard meets your accuracy and latency requirements, and confirm output token pricing applies to your workload.
Before you switch