Comparing prices across every provider we track...
AI Endpoints · $0.18/mo
Vultr's serverless inference option is substantially cheaper on both base price and per-token costs, with no region-based egress cliffs or separate IPv4 billing. However, verify model capability parity (Llama vs Qwen) and actual bandwidth costs for your traffic patterns, as Vultr's overage rates vary by location.
Input and output tokens both cost 100x less ($0.01/M vs $0.18/M), with no hidden IPv4 or region-based egress penalties after a cutoff date.
Caveat: Model is Llama-based (not Qwen); verify it meets your inference accuracy and latency requirements, and confirm bandwidth overage rates for your data center.
Before you switch