Comparing prices across every provider we track...
AI Endpoints · $37.23/mo
Vultr's serverless inference offering is dramatically cheaper for typical usage patterns, charging only per token consumed rather than per-hour billing. However, the models are different (Llama vs Whisper), so this is only a true alternative if your use case can switch from speech-to-text to a language model.
Token-based pricing with no monthly minimum means you pay only for actual inference, not reserved hourly capacity.
Caveat: This is a language model, not a speech-to-text model like Whisper; bandwidth overages are location-dependent and could add cost; different model capabilities entirely.
Before you switch