Calling OpenAI or Anthropic straight from your application means one dependency, one SDK, no extra hop, no second vendor to trust with your traffic, and nothing between you and the provider when you are debugging. For a single product calling a single provider, that is not a compromise. That is the correct architecture, and adding a gateway to it buys you complexity.
✓ You call more than one provider and are writing your own routing code
✓ A provider outage takes your feature down because you have no failover
✓ You need to know spend per team, per product or per client, and the provider bill is one number
✓ You need hard spending caps rather than a bill you read afterwards
✓ Multiple teams share one API key and nobody can say who spent what
One of those is a reason to look. None of them is a reason to pay before you have tried a free option.
Prices below come from each vendor's own pricing page. Note how much of this list is free.
| Option | Price | Shape |
|---|---|---|
| LiteLLM | Free forever, self-hosted | MIT licensed, 140+ providers, virtual keys and budgets. You run it. |
| Portkey Developer | Free forever | 10k recorded logs/month, 3 day log retention. They say it is not for production. |
| Portkey open source | Free, self-hosted | Routing, retries, fallbacks, guardrails, basic dashboard. |
| Helicone Hobby | Free | 10,000 requests, 7 day retention. Observability rather than routing. |
| Portkey Production | $49/month | 100k recorded logs, then $9 per additional 100k. 30 day retention. |
| Helicone Pro | $79/month | Unlimited seats, alerts and reports, 1 month retention. |
In every one of these you still pay OpenAI and Anthropic directly for tokens. A gateway is a layer, not a reseller. Be suspicious of any comparison that presents a gateway fee as replacing your provider bill.
Every tool on this page is built for engineers. They answer which model handled a request, how long it took, and what it cost in aggregate. That is the right shape for a product team.
It is the wrong shape for an agency. If you are billing eleven clients and running AI inside deliverables for all of them, the month-end question is not what your OpenAI bill was. It is which client is eating your margin. That means tying every call to a client, a deliverable and a rate card, then putting it somewhere your billing process can read. A gateway gives you raw usage. It does not give you attribution, and it does not put it on an invoice.
That is the part I build. Not a gateway, and not a product with a monthly fee. A cost attribution and reporting layer that sits on top of whichever gateway you already picked, wired into how you actually bill.
If your AI bill is fine but your client margins are not, that is a different problem.
Worth 30 minutes.
Talk through your AI costs →