Five minutes of scrolling. One idea: score every request by difficulty, then pay only what the work deserves.
Frontier models get smarter every quarter — and your dollar buys fewer of their tokens. Agent harnesses multiply the effect: one instruction can fan out into dozens of model calls.
The opinion wars never stop — whose model is best this week, whether you should switch. So you tried them all.
Renaming a variable costs the same as designing a migration. The model doesn't mind. Your bill does.
Local models are now genuinely good at small and medium coding work. Yours is rendering a desktop wallpaper.
Frontier subscriptions, usage plans, free tiers, your own GPU. ThinRouter catalogs every model behind them.
Your agent harness sends the request. ThinRouter measures its complexity and picks the lane.
Cheapest, cost-effective, top-down, round robin, or random — per lane. Caps and limits never stop the work. Your agents run 24/7.
A real request being typed into the playground and routed — scrub it with your scroll.
Lanes, thresholds, and strategies are all configurable — or hit Smart fill and ThinRouter pre-populates them for you.
Paste any prompt into the playground and watch it get scored and routed, live.
Repeated history and tool output get compacted before they're forwarded — code, JSON, and errors stay protected, and it fails open if the compressor isn't available.
The local-versus-cloud split, spend after quota and cache, how requests were classified into lanes, and what provider caching saved you.
And every single decision is legible: the score, the signals that fired, the candidates considered, and why the winner won.
One Node process. One SQLite file. Point your harness at localhost:4100/v1, set the model to "auto", and stop overpaying for easy work.