Monday, August 31, 2026
LLM API Providers Are Paying You to Switch: 7 Deals, Mapped
Posted by

Last month I ran my usual agent workload — coding sessions, research runs, long retrieval-heavy nights — across seven small AI inference providers. Not the big labs. The resellers: companies that rent GPU time wholesale and serve open-weight and frontier models over OpenAI-compatible APIs, competing mostly on price.
The bill for the month: $0. Cash paid, across all seven: nothing.
That is not a hack, and this is not a guide to gaming anyone. It is just what this corner of the market looks like in 2026. These providers are fighting hard for developer attention, and their favorite weapon is handing out tokens: real free tiers, signup credits, deposit bonuses, entire models permanently marked 100% off. Used sensibly, that generosity covers serious personal workloads.
Here's the map, as of August 2026 — what each provider actually offers, what it was like to use, and the etiquette that keeps the whole thing alive.
Full disclosure: B.AI, GMI Cloud, and ToAPIs pay me a small referral kickback if you sign up through my links below. The biggest single line in my receipts — TokenRouter — isn't one of them.
Everyone's Competing for Your API Key
A quick bit of market structure, because it explains everything else. Between the big labs and you sits a layer of resellers and aggregators. They buy capacity wholesale, serve a dozen-plus model families behind one API, and compete on price and perks. Margins are thin, and switching providers costs almost nothing — a new base URL and an API key.
When switching costs are zero, customer acquisition gets expensive. That's where the freebies come from. Every offer below is a marketing budget, and knowing that makes the rules obvious.
The seven I used:
| Provider | The deal | Protocols | The catch |
|---|---|---|---|
| B.AI | 6 free models, $0.30 signup credit, 50% top-up match | OpenAI Chat + Responses, Anthropic | Bonus credits expire in 30 days |
| TokenRouter | A small free shelf (GLM-5.3, Nemotron, more) | OpenAI, Anthropic, Gemini | Free list is short |
| OpenRouter | 18 :free models, 50 req/day (1,000/day with $10 in credits) | OpenAI | Rate limits; fees on deposits |
| GMI Cloud | MiniMax-M3/M2.7 at 100% off | OpenAI, Anthropic Messages | Free list is a handful |
| AIHubMix | 54 -free models advertised | OpenAI | In practice, about 10 calls |
| ToAPIs | 10 starter credits (~$0.05) | OpenAI | Tiny |
| CheaperInference | No free tier — 15–60% discounts, +$10 on first $5 deposit | OpenAI, Anthropic Messages | Needs a deposit |
What Each One Actually Gives You
Referral links:
Links marked with ~ are referral links: bonus credits for you, a kickback for me, same price either way.
B.AI: the most generous free tier
B.AI (the inference reseller, not a search company) currently has six models free with no card: DeepSeek-V4-Flash plus its vision variant, Hy3, MiMo-V2.5, GLM-5.3-Flash, and Qwen3.8-Flash. All billed at zero.
Sign up with referral code BTTAEJ~ and you start with 300,000 credits (about $0.30 at their million-credits-per-dollar rate). The more interesting part is the top-up bonus: every deposit gets matched 50% — pay $1, get $1.50 in credits — capped at $100 of bonus per user. Bonus credits expire after 30 days, so it's a discount on your first month, not forever money. The code doubles as the ticket to their Pro plan if you ever want that.
One key speaks three protocols — OpenAI Chat, OpenAI Responses, and Anthropic Messages — which is unusual and quietly useful.
This is where most of my B.AI volume went: GLM-5.3-Flash and the DeepSeek-V4-Flash trio. 141M tokens, $0.
TokenRouter: a small free shelf
TokenRouter (a PaleBlueDot.AI product) lists 126 models and keeps a free shelf: z-ai/glm-5.3-free and NVIDIA's Nemotron-3-nano-omni. When I did most of my volume there, a free Qwen3.8-Max line was also live, and it carried 131.8M tokens with a 91% cache-hit rate. That single line of my usage log is most of the value I'll show you later.
They also run a standing 50% promo on glm-5.3-flash ($0.075/$0.25 per 1M) — cheap even before you consider the free shelf.
OpenRouter: the free-model menu
OpenRouter is the aggregator everyone knows: 395 models, pass-through pricing with zero markup (they charge on deposits instead — about 5.5% via Stripe). The free side: 18 models with a :free suffix, limited to 50 requests a day, or 1,000/day once you've held $10 in credits. There's an openrouter/free auto-router that just picks a free model for you.
If you want one account that touches every model family, this is it. My month there was modest — mostly MiniMax-M3:free, 20M tokens.
GMI Cloud: 100% off, and I still haven't paid
GMI Cloud is a GPU cloud that also runs a model-serving endpoint: 123 models, fp8 quantization, both OpenAI and Anthropic protocols. A few models sit at what looks like a permanent 100% discount: MiniMax-M3, MiniMax-M2.7, and an OCR model. I've been using MiniMax-M3 there for weeks. I have not deposited anything. It works.
The paid catalog is reasonably priced too — GLM-5.3-Flash lands at an effective $0.075/$0.25 per 1M with their standing discount — and if you ever need raw GPUs, they rent H100s by the hour like everyone else.
Referral link if you want it: console.gmicloud.ai/ref/B8SAFGTQ~
AIHubMix: a sample, not a tier
AIHubMix's catalog advertises 54 -free models at $0. In practice, I got about ten calls before the tap closed. That's the honest summary: treat it as a trial for evaluating their 409-model catalog, not a tier you can lean on. The daily promos (GLM-5.3-Flash at half price, rotating discount lines) are the real draw if you end up paying.
ToAPIs: ten credits and a handshake
ToAPIs gives you 10 credits to start — about a nickel at their rate of 1 credit ≈ $0.005. Not much, but it's zero commitment, and their credit-priced catalog (140 entries across text, image, and video) is fun to poke at.
Referral if you want it: toapis.com/login?aff=hSsl~
CheaperInference: not free, just cheap
The honest odd one out: no free tier. You deposit at least $5, and first funding gets a $10 bonus — a silly ratio on the minimum. After that, everything is 15–60% off list: DeepSeek-V4-Flash at $0.048/$0.097 per 1M, GPT-5.6-Terra at $0.80/$4.80, Claude Sonnet 5 at $1.40/$7.00.
It earns its place because any rotation needs a floor: the place you land when the free shelves don't carry a model you need. Sixty-two models, both OpenAI and Anthropic protocols, prices low enough that the deposit doesn't feel like a bet.
A Month In Practice
How did this look day to day? Undramatic. I keep a handful of these configured, and I use whichever free model is strong for the job at hand. When one runs dry for the evening, I pick up another tomorrow — the caps reset, and the catalogs shift often enough that a model missing today shows up free next week.
The heavy lifting went to the free Qwen3.8-Max line on TokenRouter and GLM-5.3-Flash on B.AI. MiniMax-M3 on GMI covered a long background-research stretch, and OpenRouter filled the gaps. Nothing about this required scripts, automation, or anything I'd be embarrassed to show a provider's support team.
That last part matters. The whole setup fits in a config file I'd happily email to any of these companies — which is roughly the test I'd suggest.
The Cache-Read Reason This All Works
Here's the detail that made the month work, and it isn't a trick either: 74% of my token volume was cache reads.
Agents are cache monsters. The system prompt, tool definitions, and conversation context repeat on every call, and prompt caching means repeat prefixes get billed at a fraction of input price — or not at all. Most of the free tiers I used didn't bill cache reads.
That's why a free model plus a big static prompt goes so far: the provider eats the cheap part (a cache read is nearly free to serve — the compute already happened), and only the genuinely new tokens get metered. This is also the quiet reason the economics hold up. My month was worth about $79 at list prices, but the marginal cost of serving it was a fraction of that. The generosity is sustainable precisely because cache-heavy workloads are cheap to host.
If you want your prompts to be cache-friendly in general — paid providers included — I wrote a whole piece on the mechanics: Prompt Caching: Cut LLM Costs by 90%.
The Receipts
A note on method: "value" below is what the same volume would have cost at each provider's own paid list rates for the same models — the honest what-would-this-have-cost-here number, not a made-up retail price. B.AI's DeepSeek models are valued at their idle (weekend) rates.
| Provider | Tokens | Requests | Cash paid | List-price value |
|---|---|---|---|---|
| TokenRouter | 132.9M | 1,322 | $0 | ≈ $58.22 |
| B.AI | 141.0M | 1,328 | $0 | ≈ $14.19 |
| GMI Cloud | 27.3M | 273 | $0 | ≈ $4.64 |
| OpenRouter | 20.0M | 185 | $0 | ≈ $1.65 |
| Cloudflare Workers AI | 0.5M | 47 | $0 | ≈ $0.20 |
| AIHubMix, ToAPIs | ~0 | a few | $0 | ≈ $0 |
| Total | ~322M | $0 | ≈ $79 |
(Cloudflare's free Workers AI allocation isn't one of the seven — it's just where a few thousand tokens of odd jobs landed.)
One experiment note: I rotated models as much as providers, and a lot of the value hides in cache-heavy lines — the TokenRouter figure is 91% cache reads on a single model. Your mileage will differ. The shape is what matters: free tiers plus cache-friendly workloads equals a very small bill.
Don't Ruin This for Everyone
Now the part that actually protects this deal.
Every offer in this post is a customer-acquisition budget. The providers are paying for your attention with tokens, betting that you'll like the service and start paying. That bet only pays off — and the offers only survive — if people behave like customers instead of locusts.
So, plainly:
- One account per person. No farming signups with fresh emails, no VPN identity games, no referral rings. It's the fastest way to get free tiers killed for everyone, and it's fraud in most jurisdictions anyway.
- Respect rate limits and the ToS. The caps exist so thousands of people can try the product. Hammering them isn't clever — it's exactly why some offers end up gated behind deposits.
- Personal-scale only. Free tiers are for evaluating, learning, and personal projects. The moment a workload looks like production traffic, it belongs on a paid plan. Your startup shouldn't run on someone else's marketing budget.
- Pay the ones you lean on. If a provider becomes your default, deposit something. The bonus structures (B.AI's 50% match, CheaperInference's first-funding bonus) make it cheap to do the right thing.
I've deliberately written this as a map, not a manual. There's nothing here about squeezing limits or juggling identities, because that behavior is what ends the golden window. The current generosity is unusual, and it will not last forever. Use it like it might not.
Where to Start
If you want to try this yourself, the short version:
- Start with the real free tiers — B.AI, TokenRouter, and OpenRouter have the deepest free catalogs right now, and GMI Cloud if the MiniMax models cover your work.
- Claim signup credits where they exist. They're small, but they're free.
- Keep prompts cache-friendly: static system prompt, stable tool order. It multiplies what free models give you.
- Keep one cheap paid fallback for models the free shelves don't carry.
- Pay whoever becomes your default.
For the wider picture on spending less on tokens, I keep notes on API cost optimization and model routing. And if you're wiring several providers into one tool, the OpenCode configuration reference shows how they coexist in a single config.
Related Articles & Deep Dives
#agent-skillsCross-Platform Agent Skills: Writing Universal SKILL.md for Antigravity, OpenCode & Claude Code
A complete guide to authoring portable Agent Skills that run seamlessly across Google Antigravity, OpenCode, Claude Code, and Cursor without rewrite.
#copilotCoding Agents at Production Scale: What 761M LLM Calls Reveal
A production-scale study of Copilot's agent: 761M calls, 95T tokens. KV cache collapses across turns and after model switches; failures trigger 4x compute.
#claude-codeClaude Code Burns 33,000 Tokens Before Your Prompt Arrives — We Counted Every One
A systima.ai API-level analysis reveals Claude Code sends 33k tokens of system prompt and tool schemas per request vs 7k for OpenCode. Cache instability makes it worse — up to 54x more cache writes. Here's what it means for your daily workflow.