Comparing prices across every provider we track...
Serverless Inference · $0.01/mo
Vultr's Llama-3.1-Nemotron pricing is competitive at $0.01/M tokens, but OVHcloud's bge-m3 embedding model offers identical per-token rates on a different model class. Neither has a clear cost advantage; the choice depends on model fit for your use case and tolerance for hidden costs like bandwidth overages and regional traffic policy changes.
Identical $0.01/M token pricing with fp16 quantization, but targets embedding workloads rather than instruction-following LLM inference.
Caveat: Different model class (embedding vs. generative LLM); APAC outbound traffic loses free-egress after mid-2026; Object Storage has retrieval fees and minimum storage durations.
Before you switch