Comparing prices across every provider we track...
AI Endpoints · $0.47/mo
Vultr's serverless inference offering is dramatically cheaper on both base price and per-token costs, though it uses a smaller model (8B vs 120B parameters). OVHcloud's 120B model is more capable but carries significant hidden costs including IPv4 charges, potential Windows license fees, and upcoming APAC egress billing changes that could erode the value proposition.
Per-token pricing is 47x cheaper ($0.01/M vs $0.47/M for output), with no monthly base fee and a 100% network uptime SLA.
Caveat: Model is significantly smaller (8B vs 120B parameters), so inference quality and capability will be substantially lower for complex tasks; bandwidth overages vary by data center and are not included.
Before you switch