The short version
- Most AI work is routine: classifying, extracting, drafting, summarizing. It does not need a frontier model.
- Open models such as Llama, Qwen, Mistral and Gemma run locally through Ollama or LM Studio, with no per-token cost.
- When you do call an API, a router like OpenRouter, prompt caching and the smallest model that does the job all cut the bill.
- Start with your single highest-volume routine task: move it off the frontier API and measure the savings.
The bill that never stops
Frontier AI is priced per token, and the bill grows with every request. The trap is sending all of it to the most expensive model, including the easy work that a far cheaper model would handle just as well. You also run into rate limits and session caps at the worst moments, because you are leaning on a metered service you do not control. The fix is not to use less AI. It is to stop overpaying for the parts that do not need the best model.
Most of your AI work does not need a frontier model
Look at what your AI actually does all day. Classifying tickets. Pulling fields out of a document. Drafting a first version. Answering a simple question from a known source. Tagging, routing, summarizing. This is the routine majority, and it is not hard reasoning. A small, cheap model handles it reliably. The hard minority, the genuinely tricky reasoning and the high-stakes call, is where a frontier model earns its price. Treating every request the same is what makes AI expensive.
In a service business, the routine majority looks like tagging incoming service requests, pulling job details out of emails, and drafting follow-up texts. None of that needs the most expensive model.
A useful rule of thumb mirrors how power and water bills work: most of the load is steady and predictable, and only a little is peak. You do not buy peak-rate power for your whole house. You should not buy frontier tokens for your whole workload.
Run the routine on a model you control
You can run capable open models on your own hardware with no per-token cost. Ollama and LM Studio make it a few clicks to download and run models like Llama, Qwen, Mistral and Gemma locally. A modern laptop runs the small ones well. A machine with a capable GPU runs the larger ones and serves a whole team. Once it is running, every routine request that hits it costs nothing per token and never counts against a vendor rate limit.
Local models are strong at exactly the routine work above. They are not the best at the hardest reasoning, and that is the point. You keep a frontier model on call for that small share rather than paying for it on everything.
Cheaper routes when you do call an API
Some work still belongs in the cloud. When it does, you do not have to pay top-of-menu prices. OpenRouter gives you one API that routes across many providers, so you can pick a cheaper or free model per task and fall back automatically if one is down. Beyond that, the basics add up: turn on prompt caching so you stop paying again for the same context, pick the smallest model that does the job, and batch where you can. None of this is exotic. It is just refusing to pay frontier rates by default.
A simple rule for what to keep local
Keep it local or cheap when the task is
routine and high volume (classify, extract, draft, summarize, route), latency tolerant, or involves sensitive data you would rather not send out. Send it to a frontier model when the task is genuinely hard reasoning, low volume, or high stakes where the best answer is worth the price.
You do not have to get this perfect on day one. Start by moving your single highest-volume routine task off the frontier API and onto something cheaper, measure the savings, and expand from there.
Where this goes
The more of your routine you can run on models you control, the less of your budget is exposed to someone else's pricing and limits. And once an agent is running that routine work, the procedures it follows are AOPs, the agent-runnable version of your SOPs.
Common questions
What is OpenRouter?
OpenRouter is a single API that routes your requests across many model providers, so you can pick a cheaper or free model per task and fall back automatically if one is unavailable. It is a simple way to stop paying frontier prices for work a smaller model handles fine.
Is a local AI model good enough?
For the routine majority of work, yes. Open models in the 7B to 70B range handle classification, extraction, summaries, drafting, and structured output reliably. They are not the best at the hardest reasoning, which is exactly why you keep a frontier model on call for that small share instead of paying for it on everything.
How do I use AI for cheaper or free?
Run open models locally for zero per-token cost, route API calls through a cheaper provider like OpenRouter, turn on prompt caching, and pick the smallest model that does the job.
If you want help
Tell us where your AI spend goes today. We will show you which routine work can run cheaper or local, and what to keep on a frontier model. Book a free 20-minute call.
Sources
- Ollama, official site
- LM Studio, official site
- OpenRouter, official site
Update log
- Moved to the new Signals layout and trimmed, including a section on coding agents and two questions that repeated the post. Added an example from a service business. The guidance is unchanged.
- First posted.
Current as of June 27, 2026. Signals is general information, not legal, security or financial advice. How we source Signals.