Comparing prices across every provider we track...
Serverless Inference · $0.85/mo
OVHcloud's bge-m3 embedding model is dramatically cheaper on base pricing ($0.01/mo vs $0.85/mo) and per-token costs ($0.01/1M input tokens vs $0.85/M), but it is a different model class (embedding vs generative LLM). If your use case is semantic search or embeddings rather than text generation, OVHcloud offers exceptional value; if you need GLM-5.1-FP8 specifically for generative tasks, Vultr remains your only listed option.
Sticker price is 98.8% lower and per-token input cost is 99% lower than Vultr's GLM-5.1-FP8.
Caveat: bge-m3 is an embedding model, not a generative LLM; only suitable if your workload is semantic search, clustering, or similarity tasks, not text generation.
Before you switch