4 configs 1x-8x
Cerebrium
Serverless GPUs for real-time AI inference
Founded in Cape Town in 2021 and now headquartered in New York, Cerebrium runs serverless GPU infrastructure for real-time AI workloads: voice agents, image and video pipelines, LLM endpoints, and custom Python apps. Hardware is declared in a cerebrium.toml file and the code deploys as a container that scales from zero, with compute billed by the second.
Example customers include Tavus, Deepgram, Resemble AI. More here.
Cerebrium Homepage
What's good about Cerebrium
- Per-second billing, with no charge for idle GPUs
- Cold starts of 2-4 seconds via memory and GPU snapshots
- SOC 2 Type II, HIPAA, GDPR and ISO 27001 certified
- Ten GPU types, from T4 to B200
- Region pinning for data residency, on US and EU capacity
Cerebrium pricing examples
Cerebrium bills compute by the second while a container is running, with the GPU, vCPU and memory each priced separately, so a deployed app costs more than its GPU rate alone.
Platform plans sit on top of compute: a free tier, a flat monthly subscription, and custom enterprise pricing with volume discounts. Persistent storage is billed per GB-month.
Below are some example configurations and their estimated costs:
| Example configuration | Estimated cost |
|---|---|
| VM Small | $22.73 / mo 1 vCPU, 1 GB RAM (CPU only) |
| VM Medium | $113.94 / mo 4 vCPU, 8 GB RAM (CPU only) |
| VM Large | $227.89 / mo 8 vCPU, 16 GB RAM (CPU only) |
| Block Storage | $5.00 / mo 100 GB beyond free allowance · First 100 GB included |
Cerebrium GPUs
Here are some of the GPU configurations offered by Cerebrium:
4 configs 1x-8x
4 configs 1x-8x
4 configs 1x-8x
8 configs 1x-8x
4 configs 1x-8x
4 configs 1x-8x
4 configs 1x-8x
No GPU models match .
Heads up: A provider's own page may quote a different figure, for example on tax, region or a promotion. A monthly-only plan shows a derived hourly rate. Verify before provisioning. More on how we price. See Cerebrium's pricing.
Which services does Cerebrium offer
Here are some of the services that Cerebrium offers:
Data center locations
Based on our records, Cerebrium operates in the following locations:
| Country | Location | Slug |
|---|---|---|
| Sweden | Stockholm | eu-north-1 |
| USA | Kansas City, MO | us-central1 |
| USA | Northern Virginia, VA | us-east-1 |
Alternatives to Cerebrium
Compare Cerebrium against other cloud providers:
-
Runpod
Runpod is a GPU cloud offering serverless inference, on-demand pods and community-hosted capacity.
-
Replicate
Replicate offers a serverless platform to run and fine-tune open source AI/ML models with an extensive model library.
-
Lambda Labs
Lambda Labs specializes in GPU cloud computing with a focus on deep learning training and research.
-
Beam
Beam runs AI workloads as serverless endpoints, task queues and sandboxes, defined from a Python, TypeScript or Go SDK.
-
Fal.ai
Fal.ai is a serverless platform focused on inference for generative image, video and audio models.
Our data for Cerebrium was last updated on Sept. 4, 2026.