On August 5, 2025, OpenAI released gpt-oss-120b and gpt-oss-20b as open-weight models under the Apache 2.0 license — the first time you can legally download an OpenAI reasoning model, run it on your own GPU, and never send a token to their API. The obvious question every Indian SMB CTO asked us that week: "should we self-host this and stop paying OpenAI?" The honest answer is usually no — but the cases where it's yes are specific and worth knowing. Here's the break-even math and the three questions that actually decide it.
The 60-Word Answer
Self-host gpt-oss-120b only if one of three things is true: you have hard data-residency rules that forbid sending data to a US API, your token volume is high and steady enough to beat GPU rental on cost, or you need full control over latency and uptime. For everyone else — spiky traffic, small volume, no residency constraint — the API is cheaper and far less work. Most SMBs should stay on the API.
Why This Matters Now
OpenAI's gpt-oss release changed the choice set. gpt-oss-120b uses a mixture-of-experts design: 117B total parameters but only 5.1B active per token, which is why it runs on a single 80GB GPU. The smaller gpt-oss-20b (21B total, 3.6B active) runs in 16GB — laptop-and-edge territory. Both ship natively quantized in MXFP4, support a 128K context, and carry the Apache 2.0 license, meaning you can use them commercially without asking anyone (per the model card on Hugging Face). For the first time, "run an OpenAI-quality model entirely inside our own network" is a real, legal option for Indian businesses with data-residency worries.
| Spec | gpt-oss-120b | gpt-oss-20b |
|---|---|---|
| Total parameters | 117B | 21B |
| Active per token | 5.1B | 3.6B |
| Minimum GPU memory | ~80 GB (one GPU) | ~16 GB |
| Context length | 128K | 128K |
| Quantization | MXFP4 (native) | MXFP4 (native) |
| License | Apache 2.0 | Apache 2.0 |
Self-Host vs. API: The Honest Comparison
The Break-Even Math
Here's the calculation we run for clients. An 80GB GPU instance (an H100 or A100-80GB) rents by the month on cloud GPU providers, or you can buy/colocate. The API cost depends on your token volume. The chart shows where the lines cross for a workflow averaging a moderate token mix.
The 3 Questions That Decide It
The Self-Host Readiness Checklist
Before you commit to running gpt-oss yourself, you should be able to tick every box below. If you can't, the API is still your answer.
- A written data-residency rule, regulation, or client contract that forbids the API — not a vague preference
- A realistic monthly token volume estimate, and the break-even math run against current GPU pricing
- Confirmation your traffic is steady, not spiky — a GPU you can keep busy most of the day
- An owner for GPU ops: drivers, the inference server, scaling, monitoring, and on-call
- The right-sized model chosen — gpt-oss-20b in 16GB may clear your bar far cheaper than 120b
- A fallback plan to the API (one base-URL change) if self-hosting underperforms
If You Do Self-Host: The Minimal Setup
For the cases where self-hosting wins, the fastest path to a working endpoint is vLLM, which serves gpt-oss with an OpenAI-compatible API — so your existing code barely changes.
# On an 80GB-GPU box (H100 / A100-80GB)
pip install vllmServe gpt-oss-120b with an OpenAI-compatible endpoint
vllm serve openai/gpt-oss-120b \
--tensor-parallel-size 1 \
--max-model-len 128000Your app points at the local server instead of api.openai.com:
base_url = "http://localhost:8000/v1"
The Chat Completions calls are otherwise identical.
The OpenAI-compatible endpoint is the quiet superpower here: code written against the OpenAI SDK works against your self-hosted model with a one-line base_url change. That makes "try self-hosting, fall back to API" a low-risk experiment rather than a rewrite.
Common Mistakes
When NOT to Self-Host (the Default for Most SMBs)
Real Example: When We Said Yes, and When We Said No
A fintech lender in Mumbai had a contractual rule that borrower data could not leave their VPC. For them, self-hosting gpt-oss-120b on a colocated GPU was the only compliant path, and the steady document-processing volume kept the GPU busy. We said yes and built the vLLM endpoint. A 30-person D2C brand asked the same question the same week — but their traffic was spiky support chat at modest volume, no residency rule, no infra owner. For them self-hosting would have meant a costly idle GPU to replace a small monthly API bill. We said no, firmly. The discipline is the same as our GPT-5 launch-day analysis: measure your actual situation, don't follow the headline.
Our AI automation team runs this exact three-question screen before any self-hosting build, and we've reused the model-as-config pattern from our 2025 n8n workflows so clients can switch between self-hosted and API endpoints with a config change.
Frequently Asked Questions
What hardware do I need to run gpt-oss-120b?
A single GPU with about 80GB of memory — an H100 or A100-80GB — thanks to its mixture-of-experts design and native MXFP4 quantization. The smaller gpt-oss-20b runs in roughly 16GB, which fits a high-end consumer or edge GPU. Both natively support a 128K context window.Is gpt-oss free to use commercially?
Yes. Both gpt-oss-120b and gpt-oss-20b are released under the Apache 2.0 license, which permits commercial use, modification, and redistribution without a separate agreement. You still pay for the hardware to run them, but there are no licensing fees to OpenAI.Is self-hosting gpt-oss cheaper than the OpenAI API?
Only at high, steady volume. A self-hosted GPU is a flat monthly cost whether idle or busy, while the API charges per token. Below your break-even (often several million tasks a month), the API is both cheaper and zero-ops. Run your own numbers with your real token mix.When does self-hosting actually make sense?
Three cases: a hard data-residency rule that forbids sending data to a US API; high and steady volume that keeps a GPU busy; or a strict need to control latency and uptime yourself. If none of these apply, the API is the better default for an SMB.Can I switch between self-hosted gpt-oss and the OpenAI API easily?
Yes, if you serve the model with an OpenAI-compatible server like vLLM. Your code only changes the base URL — the Chat Completions calls stay identical. That makes self-hosting a low-risk experiment you can roll back to the API in one config change.Should a small business self-host gpt-oss in 2025?
Usually no. Most SMBs have spiky traffic, modest volume, no residency requirement, and no dedicated infra owner — exactly the profile where the API wins on both cost and effort. Self-hosting is a powerful option for regulated or high-throughput workloads, not a general cost-saver.Want an honest self-host vs. API decision for your AI workload?
We run a fixed-scope assessment: your data-residency rules, real token volume, and ops capacity — then a clear recommendation with the break-even math, and the build either way. The assessment fee is credited toward the build. Suitable if you're weighing gpt-oss self-hosting against the API. First call is technical.
Book a 20-min CallRelated reading
As Hrishikesh, our CTO, puts it: open weights are a gift, but a GPU you forgot to keep busy is the most expensive idle hardware you'll ever own. Answer the three questions first.
