Comparing prices across every provider we track...
AI Endpoints · $0.47/mo
Vultr's Llama-3.1 serverless inference is 47x cheaper on base price and 47x cheaper per input token, with no hidden IPv4 or license surcharges. However, the models are different (Qwen vs Llama), so verify the model choice meets your use case before switching.
Input token cost is $0.01/M (vs OVHcloud $0.47/M), base fee is $0.01/mo, and no per-instance IPv4 or Windows license charges apply to this serverless offering.
Caveat: Different model family (Llama vs Qwen); verify output token pricing and bandwidth overage rates for your traffic pattern, and confirm 100% SLA meets your availability needs.
Before you switch