Comparing prices across every provider we track...
AI Endpoints · $4.25/mo
Vultr's serverless inference offering is dramatically cheaper on both base and per-token costs, though it uses a smaller model (8B vs 397B parameters). OVHcloud's Qwen3.5-397B is a much larger, more capable model; the price difference reflects that fundamental gap rather than OVHcloud being uncompetitive for its tier. Choose based on model capability needs, not price alone.
Negligible base cost and per-token pricing 425x cheaper than OVHcloud, with 100% network uptime SLA.
Caveat: Model is 50x smaller (8B vs 397B parameters), so inference quality and capability are not equivalent; bandwidth overages vary by data center and may add cost at scale.
Before you switch