Cut your AI API bill. Prove quality holds.
Send us one month of OpenAI or Anthropic usage. In five business days you get a line-by-line report of where your LLM spend can come down, and a side-by-side test on your own examples before anything changes.
Most LLM spend goes to habits nobody revisits.
A feature ships on the biggest model because that's what worked in the prototype. Six months later it's still there, paying frontier prices for work a smaller model handles fine.
Oversized models
Tagging, extraction and routing running on a model built for hard reasoning.
Repeated prompts
The same system prompt, tool list or reference document paid for at full price on every call.
Real-time pricing for batch work
Nightly summaries and bulk jobs that nobody waits on, billed at interactive rates.
Stale model versions
Older models kept around after a newer one in the same class came out at a lower price.
Three steps. About 15 minutes of your time.
Export your usage
Download one month of usage from your OpenAI or Anthropic console, or fill in a one-row-per-feature sheet. Token counts by model are all we need.
We analyse it
We price every workload at current list rates, then check for cheaper same-class models, caching, batching and right-sizing. Savings are applied in sequence, so nothing is double-counted.
You get a report, then decide
A one-page report with a conservative and an optimistic estimate for each change. Nothing touches your production code; your team decides what ships.
Cheaper only counts if the answers stay right.
Before we recommend a smaller model, we run the same examples through both models and show you the answers side by side. You pick the examples, with anything sensitive removed.
- 50–200 of your own examples, not a generic benchmark
- Exact-match rate where there's a right answer; every difference highlighted for review
- You can run it yourself with your own key; the script is open source
- Model downgrades only count in the optimistic estimate until they pass
| Example | Current model | Candidate | |
|---|---|---|---|
| "Charged twice this month" | billing | billing | Match |
| "Export button does nothing" | bug | bug | Match |
| "How do I add a teammate?" | how-to | how-to | Match |
| "Change my login email" | account | how-to | Review |
| "Love the product!" | other | other | Match |
A report your engineers can act on the same day.
- Current monthly spend per workload, recomputed from token counts
- Each recommended change with a low and high monthly saving
- Concrete instructions: which model, which prompt prefix to cache, which jobs to batch
- A CSV of every number, so you can check our math
Prices come from OpenAI's and Anthropic's official pricing pages, dated in every report.
From a sample report built on illustrative data. View the full sample report →
We see token counts, not your business.
The audit is built so you never have to trust us with anything sensitive.
No keys, no access
We never ask for API keys, passwords or access to your systems. You send an export file; that's it.
No customer data
Usage exports contain model names and token counts. For the quality test you choose the examples and remove anything personal first.
Deleted in 30 days
Your files are used only to prepare your report and are deleted within 30 days of delivery, or sooner on request.
NDA on request
Happy to sign your mutual NDA before you share anything. We never name you or share your numbers without written permission.
Open-source methodology
The audit engine and quality-check script are public on GitHub. Read exactly how every number is calculated.
Nothing changes without you
We recommend; your engineers implement. No changes are made to your code, prompts or providers by us.
Start with a free pilot.
We're working with a small number of pilot companies this quarter. The pilot is free with no obligation; all we ask is honest feedback on whether the report was useful.
One full audit, no commitment.
- Audit of one month of usage
- Savings report with low/high estimates
- Side-by-side quality test on up to 200 examples
- 15-minute walkthrough call
For teams that want ongoing help.
- Monthly re-audit as prices and models change
- Help implementing caching and batching
- Quality tests before each model switch
One focus: cutting wasted tech spend.
Pantheon helps small companies spend less on the technology they already run, and be found where their customers now search.
AI Cost Audit New
Lower OpenAI and Anthropic bills through right-sizing, caching and batching, with quality proven on your own examples.
Cloud Cost Optimization
Find idle and oversized cloud resources and bring your monthly infrastructure bill down.
AI Search Visibility
Understand and improve how your company shows up in answers from ChatGPT, Claude and other AI assistants.
Arnav Thakur, Founder
Pantheon is a small, founder-led company based in Georgia. You work directly with me from first email to final report, and I answer every message myself.
I build the tools we use in the open. Along with the audit engine, I've contributed forecasting and inventory-planning software to Ruby for Good's Human Essentials, an open-source platform used by diaper banks across the U.S.
Questions founders ask
What exactly do you need from us?
One month of usage exported from your OpenAI or Anthropic console (Usage page → export CSV). Optionally, a short note on what each feature does and whether it's latency-sensitive, which lets us find more savings. For the quality test, 50–200 example inputs you've chosen and scrubbed.
Why is it free?
We're running pilots to prove the audit on real workloads. In return we ask for honest feedback, and, only if you're happy, permission to mention that we worked together.
How accurate are the savings estimates?
Every estimate is a range. The low end counts only the safer changes (same-class model updates, caching of prompt parts you've confirmed are fixed, batching work nobody waits on). The high end adds smaller-model moves, which count only after passing the quality test. We price at public list rates, so private discounts or credits aren't included.
We already track our model costs. Is this still useful?
Possibly not, and we'll tell you if so. The audit helps most where nobody owns the AI bill full-time, for example when AI features were added to an existing product and the defaults from the prototype stuck.
Which providers do you support?
OpenAI and Anthropic today. Other providers can be added to the price table on request.
Do you ever touch our production systems?
No. We produce a report and, if you want, a quality-test comparison. Your team decides what to change and makes the changes.
Find out what your AI bill should be.
Email us with your company name and roughly what you spend on LLM APIs per month. We'll reply within one business day with next steps.