Optimise & save

Cheaper models. Same quality bar. Your real token mix.

A cheaper model that is just as good for your workload ships every month — and nobody has time to keep re-evaluating. StackSpend does it daily: for every LLM you use, it finds alternatives that match or beat it on quality benchmarks, prices them at your actual usage, and shows the projected saving.

On the Business plan — free to try during your 14-day trial.

Model recommendations

~$1,240/mo potential

3 open

gpt-5gpt-5-mini

Matches or beats on coding · within tolerance elsewhere

~$680/mo

−72% at your mix

claude-opus-4-8claude-sonnet-5

Matches or beats on coding · within tolerance elsewhere

~$410/mo

−58% at your mix

gemini-2.5-progemini-2.5-flash

Matches or beats on reasoning · within tolerance elsewhere

~$150/mo

−31% at your mix

Projected savings at your token mix — illustrative example.

Direct answer

What are Model Recommendations in StackSpend?

Model Recommendations continuously check every LLM your team uses against the market and suggest cheaper models that are at least as good on published quality benchmarks — priced at your actual token mix, with a projected monthly saving for each switch. Prices and benchmark scores refresh daily, so a price cut or a newly released model shows up in your recommendations without anyone watching announcement feeds.

The method

How a recommendation earns its place.

Step 1

Measure your real mix

Recommendations start from what you actually use: tokens per model over the last 30 days, split into input, output, and cache traffic. Headline per-token prices mislead — your mix decides your real cost.

Step 2

Find the quality bar

Every model is scored on three published benchmark axes — coding, reasoning, and math. The strongest axis of the model you use is treated as the reason you chose it, and becomes the bar a candidate must clear.

Step 3

Screen with guardrails

A candidate must match or beat your model on its strongest axis and stay within tolerance on every other axis it is benchmarked on. That rule is what stops flagship-to-nano downgrades that look cheap and fail in production.

Step 4

Price at your mix, rank by saving

Qualifying candidates are priced at your real input/output ratio, ranked by projected saving, and surfaced with up to three alternatives — each with per-axis scores and a full comparison page.

Built for trust

Honest numbers, not optimisation theatre.

A recommendation is only useful if you can defend it — to your team, your CTO, or your board. Every StackSpend recommendation is built to survive that scrutiny.

Projected, never booked

Savings are the price difference at your recent mix, normalised to a month. StackSpend labels them as projections and expects you to validate with a guarded trial before moving production traffic.

Ranges when data is thin

If a provider doesn’t report token direction, the input/output split is estimated and the saving publishes as a low–high range — not false precision.

Conservative by construction

Candidate costs use list prices without assuming batch or cached-input discounts, and fine-tuned models are excluded because a base benchmark says nothing about your fine-tune.

Approved providers only, if you choose

Scope suggestions to the whole market or a selected vendor panel — so recommendations never propose a provider your organisation hasn’t approved.

A lifecycle, not a list

Dismiss it and it stays dismissed. Implement it and the saving is credited. If prices or models change and a recommendation stops being valid, it closes itself. Each one can file a Linear or Jira issue.

Questions

Frequently asked

What are Model Recommendations in StackSpend?
Model Recommendations continuously check every LLM your team uses against the market and suggest cheaper models that are at least as good on published quality benchmarks — priced at your actual token mix, with a projected monthly saving for each switch.
How do I optimise my LLM selection?
Optimising LLM selection means matching each workload to the cheapest model that meets its quality bar: measure your real token mix, compare candidate models on the benchmark axis that matches the workload, price alternatives at that mix rather than headline rates, and re-evaluate as prices and models change. StackSpend automates this loop and re-runs it daily.
How do I know a cheaper model is good enough?
StackSpend only recommends a model that matches or beats your current model on its strongest benchmark axis and stays within tolerance on every other axis it is benchmarked on. Savings are projected, not booked — validate with a guarded trial before switching production traffic.
Are the savings figures real?
They are projections: the price difference between your current model and the alternative, computed at your recent token mix and normalised to a month. When token direction data is missing the figure is published as a low–high range rather than a single number, and estimates are deliberately conservative.
Can I restrict recommendations to approved providers?
Yes. An organisation-level setting scopes suggestions to the whole market or to a selected panel of approved model providers, so platform teams can standardise which vendors engineers may use.
Do recommendations fit into our engineering workflow?
Each recommendation can create a Linear or Jira issue automatically, and every one carries a dismiss / mark-implemented lifecycle. Implemented switches credit the projected saving; stale recommendations close themselves when prices or models change.
Which plan includes Model Recommendations?
Model Recommendations are part of the AI Explorer on the Business plan, and free to try during your 14-day trial — connect your providers and run a model audit on your real usage before you commit.

Find out what your model choices are costing you.

Connect a usage feed, and StackSpend audits every model you use against the market — benchmark-guarded, priced at your real mix.

Want the usage picture first? Explore the AI Explorer.
Model Recommendations — Cheaper LLMs, Benchmark-Guarded — StackSpend