Comparing prices across every provider we track...
AI Endpoints · $0.71/mo
Vultr's Llama-3.1 serverless inference is dramatically cheaper at $0.01/mo with identical token pricing, though it runs a different (smaller, safety-focused) model. OVHcloud's Qwen3.5-397B is a much larger model, so direct comparison is difficult; however, Vultr's pricing structure has no hidden regional surcharges or upcoming egress changes, making it genuinely better value if model capability meets your needs.
Transparent per-token pricing with no regional surcharges, IPv4 fees, or upcoming egress penalties; 100% network uptime SLA provides predictability.
Caveat: Model is 8B parameters (safety-guard variant), not 397B; suitable only if smaller model meets your inference needs.
Before you switch