Comparing prices across every provider we track...
AI Endpoints · $3.19/mo
Vultr's serverless inference offering is dramatically cheaper on both base price and per-token costs, though the model (Llama-3.1-Nemotron-Safety-Guard-8B) is smaller and different in capability than Qwen3.6-27B. OVHcloud's hidden costs around IPv4 billing, regional egress changes post-2026, and storage retrieval fees add material expense that aren't reflected in the headline $3.19/mo figure.
Input and output tokens both cost $0.01/M versus OVHcloud's $3.19/M output, with no separate IPv4 or regional egress surprises documented.
Caveat: Model is 8B parameters (much smaller than 27B), so capability and accuracy will differ significantly; verify model fit for your use case before migrating.
Before you switch