Comparing prices across every provider we track...
AI Endpoints · $3.19/mo
Vultr's serverless inference offering is dramatically cheaper on both base price and per-token costs, though the model (Llama-3.1-Nemotron-Safety-Guard-8B) is smaller and different in capability than Qwen3.6-27B. OVHcloud's hidden costs around IPv4 billing, regional egress changes post-2026, and storage retrieval fees add material expense that Vultr's simpler pricing avoids.
Input and output tokens both cost $0.01/M versus OVHcloud's $3.19/M output, with no separate IPv4 or regional egress surprises.
Caveat: Model is 8B parameters (smaller, different architecture) versus Qwen3.6-27B; verify inference latency and accuracy match your workload before migrating.
Before you switch