Ollama Pricing: Managed vs Self-Hosted
Managed Ollama is billed by usage. Self-hosted, this stack measures 4 GiB across 1 service and fits a $25/month flat plan.
#the numbers
Managed prices as of June 2026, each linked to its source. The self-hosted row is not an estimate: the RAM figure is what this stack allocates in its compose file, and the plan is the smallest one that holds it.
| Option | Plan | Monthly | What you get |
|---|---|---|---|
| Self-hosted on Miget ★ | 4 GiB, 2 vCPU | $25 | the whole stack, flat - no per-user or per-request metering, and the plan holds your other apps too |
| OpenAI API | usage-based | usage-based | per-token, per call - scales with usage |
Ollama runs open models with no per-token bill; the trade is you provide the compute (a GPU is strongly recommended).
#on usage-based pricing
OpenAI API bills by usage rather than a fixed monthly price, so a table cannot compare it honestly. The structural difference is the one that matters: a flat plan has a known worst case, and a meter does not.
#what it costs on other platforms
Same stack, published rates elsewhere - estimates from each platform's own pricing page, not quotes.
| Platform | Estimated monthly | Why |
|---|---|---|
| Miget ★ | $25 | flat plan, 4 GiB - databases included, no meters |
| Heroku | ~$200 | no volumes; nothing between 1 GB ($50) and 2.5 GB ($250) dynos - 2 GB containers cost far more than shown |
| Render | ~$63 | per-service instances (0.5 GB $7, 2 GB $25) - every container is its own paid service |
| DO App Platform | ~$53 | no persistent volumes - stateful containers need managed DBs/Spaces (base $5 Spaces included here) |
| Railway | ~$48 | usage-based ($10/GB RAM-mo); vCPU billed separately at $20/vCPU-mo on top |
| Fly.io | ~$31 | cheapest sticker price - but burstable shared CPUs (1/16 core; dedicated vCPUs cost ~2-3×), no compose deploys (one app per container, manual wiring), managed DBs billed extra |
#sizing, if you outgrow the small plan
The $25 plan is the smallest that fits. For production traffic you generally want the next tier up, and dedicated vCPU if the workload is CPU-bound rather than idle-waiting.
| Use case | Plan | Monthly |
|---|---|---|
| Fits the stack | 4 GiB, 2 vCPU shared | $25 |
| Comfortable, room for your own apps | 8 GiB, 4 vCPU shared | $49 |
| Dedicated vCPU (production) | 8 GiB, 4 vCPU dedicated | $85 |
RAM tracks the model: a 3B needs a few GiB, 7-8B more. The model volume is large (~2 GB for 3B, ~5 GB for 7-8B, ~40 GB for 70B) - size it or models re-download on restart. Without a GPU, stick to small models.
Run Ollama yourself
The compose file is open source and runs locally with docker compose up. One step to deploy it to
Miget on the $25/month plan, databases included.