01. What FreeLLMAPI Is

The short answer. FreeLLMAPI is a free, MIT licensed router that runs on your own machine. It puts the free tiers of many providers behind one endpoint, and Claude Code can talk to that endpoint instead of burning your paid usage on small jobs.

You keep one key instead of juggling a login for every provider. The router picks a free model, and when that model is rate limited it moves to the next one.

How the router sits between Claude Code and the free providers Claude Code your terminal FreeLLMAPI localhost:3001 one key, your machine free provider 1 free provider 2 and the rest

Your request never leaves your machine until the router sends it to the provider you gave a key for.

02. What Goes Free, What Stays on Claude

Decide this before you install anything. Route by risk, not by price.

Send to free modelsKeep on Claude
Simple coding tasksArchitecture decisions
Student projectsReal bugs you cannot explain
Research and summariesRefactors across many files
Documentation and README workAnything a client will see
Routine operations, commit messages, test scaffoldingAnything with keys, customer data or an NDA

That split is the whole point. Your paid limits stop being spent on work a free model handles fine.

03. The Repo You Are Installing

  • Project: tashfeenahmed/freellmapi, MIT licence, 28,188 stars, last pushed 23 September 2026.
  • Scale, from the README: 34 free providers, 474 model families and 635 free model endpoints, around 7.4 billion tokens a month.
  • Live right now: the model catalog is showing 333 models across 23 providers today. Free tiers change weekly, so the two counts will rarely match.
  • Where your keys live: in your own database on your own machine, encrypted at rest. The router is local first and single user.

These are the providers behind the free pool. Every one of them has its own free tier and its own limits, which you can read on the catalog page.

Google Gemini
Mistral AI
OpenRouter
Cloudflare
NVIDIA NIM
Hugging Face
GitHub Models
Ollama
Groq
Cerebras
Cohere
Z.ai

A sample of the providers in the pool. The catalog page lists every model, its context window and its rate limit.

If you would rather wire up one provider by hand instead of running a router, Awesome Free LLM APIs is a good reference list. It is a list, not a router, so no switching and no handoff.

04. Set It Up in Claude Code

Five steps. Nothing here touches your paid Claude plan.

1. Install the router

// terminal

curl -fsSL https://freellmapi.co/install.sh | bash

It starts a local server on port 3001. Read the script first if you prefer, it is the same URL in the repo's own quick start.

2. Add one free key

Open http://localhost:3001 and paste a free provider key. A Google AI Studio key is the quickest first one. One key is enough to test, more keys give the router somewhere to fail over to.

3. Preview the change to Claude Code

// prints a diff, writes nothing

npx freellmapi setup-claude --url http://localhost:3001 --dry-run

4. Apply it

// merges into your existing config, backs it up first

npx freellmapi setup-claude --url http://localhost:3001 --api-key <your-unified-key>

The real write merges with your existing configuration and creates a timestamped backup before it changes anything. You can also keep separate setups with --profile <name>.

5. Run something small

Open Claude Code and give it a throwaway job: write a README section, or explain one file. If it answers, you are routed.

05. Prove It Is Routing

Every response carries a header naming the provider and model that served it.

// what you are looking for

X-Routed-Via: <platform>/<model>

If that header names a free provider, your everyday work is no longer spending paid usage. If it is missing, Claude Code is still pointed at its normal endpoint and step 4 did not apply.

06. When a Model Hits Its Limit

This is the part that makes a router worth more than a second API key.

What happens on a rate limit 429 or 5xx skip provider cool the key next model keep going

The router tracks requests per minute and per day against each provider's own cap, so it tries to stay under the limit before it ever has to recover from one. The catalog it checks refreshes twice a day.

07. The Context Handoff

A model switch normally means explaining your project again. Here it does not.

  • A session stays on one model for 30 minutes, so a switch is not happening mid thought for no reason.
  • When the model does change for that session, the router injects a short handoff note so the new model knows what the session was doing.
  • The note is only injected on a real model change. Not on your first request, not when the same model continues, and not twice.
  • It lives in memory only. Nothing is written to disk or logged, and it expires after three hours.
The handoff between two models model A limit reached handoff note what the session was doing in memory only, 3 hour life never written to disk model B carries on

Like handing the job to another developer with the briefing already attached.

08. Stretch the Free Tiers

Free tiers are counted in requests and tokens, so spending fewer tokens per call means more work per day. The router ships an optional compression pass that dedupes repeated prompt content, filters noisy tool output, compacts repeated JSON and trims stale context.

Turn it on once you trust the setup, not on day one. You want a clean baseline first.

09. Your Routing Policy

Copy these three tiers into your own notes and stop deciding case by case.

TierWorkRule
Tier 1Throwaway: docs, commit messages, boilerplate, learningAlways free. Never escalate.
Tier 2Real but low risk: tests, small features, researchStart free. Move to Claude when the free answer wastes your time.
Tier 3Client code, production, secrets, anything under an NDANever leaves your paid plan.

10. What It Saves, What It Costs

Example math, not a measured result. Say 60 percent of your prompts are Tier 1 and Tier 2. Moving those to free models leaves your paid limits for the 40 percent that actually needs them, which is usually the difference between finishing on a Thursday and waiting for a reset.

The honest side: free models are slower, quality moves around between providers, and nobody owes you an uptime guarantee. You will hit a bad answer and have to re-run it on Claude. That is the trade.

11. Using It at Work

The repository is direct about this, so the guide will be too. It describes itself as being for personal experimentation and learning, not production, and it tells you to move to paid APIs before you ship real applications.

Some providers' free tiers may also log or train on what you send. So: internal tooling, prototypes, learning and documentation are fine. Client code, customer data, API keys and anything under an NDA stay on your paid plan.

12. FAQ

Is FreeLLMAPI actually free?

Yes. The tool is MIT licensed and it routes to providers' own free tiers. You bring your own free keys and stay inside their limits.

Does it work with Claude Code?

Yes. It exposes an Anthropic compatible endpoint and ships a setup-claude command that merges into your existing configuration after writing a backup.

What happens when a free model hits its rate limit?

The router skips that provider, cools the key down, rotates to the next model in your chain and carries on. Your run does not stop.

Does it keep my context when the model changes?

A session holds one model for 30 minutes. On a real switch it injects a short handoff note, kept in memory only, expiring after three hours.

Can I use it for client work?

No. The repo says personal experimentation and learning, not production, and free tiers may log prompts. Keep client work on your paid plan.

Does it replace my Claude subscription?

No, it protects it. Free models take the everyday work so your Claude limits are still there for the hard parts.

Related Guides

13. Sources

// Free newsletter

I send out guides like this every week

Real setups, real sources, no hype. Drop your email and I'll send you the next one.