Comparing prices across every provider we track...
AI Endpoints · $0.05/mo
Vultr's Llama-3.1 serverless inference option is significantly cheaper on both base price and per-token costs, with no region-based egress penalties or quantization limitations. However, the model is smaller (8B vs 20B) and you should verify output token pricing and SLA coverage before committing.
Input and output tokens both cost $0.01/M (vs OVHcloud input-only at $0.05/M), with no hidden egress cliffs or region surcharges, plus 100% network uptime SLA.
Caveat: Model is 8B parameters, not 20B; verify whether output token pricing applies to your use case and whether smaller model meets accuracy requirements.
Before you switch