Comparing prices across every provider we track...
AI Endpoints · $0.12/mo
Vultr's Llama-3.1-Nemotron offering is significantly cheaper on both base price and token costs, with a strong 100% uptime SLA. However, OVHcloud's Qwen3.5-9B may suit you better if you need that specific model, have predictable traffic patterns, or prioritize APAC region stability before mid-2026.
Input and output tokens both cost $0.01/M (vs OVHcloud input-only at $0.12/M), plus identical $0.01/mo base, backed by 100% network uptime SLA.
Caveat: Different model (Llama vs Qwen); bandwidth overages vary by data center and Windows licensing not included in base price, similar to OVHcloud.
Before you switch