Comparing prices across every provider we track...
Serverless Inference · $0.15/mo
Vultr's Nemotron-Cascade pricing is reasonable for a large language model endpoint, but OVHcloud's embedding model is not a direct substitute—it serves a different use case (semantic search vs. text generation). On token pricing alone, OVHcloud is cheaper, but you're comparing different model classes and capabilities.
Dramatically lower per-token cost ($0.01/1M input) and base fee, but this is an embedding model, not a generative LLM.
Caveat: Not a functional equivalent: bge-m3 generates vector embeddings for retrieval; Nemotron-Cascade generates text. APAC egress loses free status after mid-2026, and IPv4 addresses are billed separately.
Before you switch