Site icon Kahawatungu

How to Compare LLM API Providers: Stop Comparing Models

How to Compare LLM API Providers

How to Compare LLM API Providers

The usual way to compare LLM API providers is to line up model prices and pick the lowest row. That method is exactly backwards: it compares models, not providers, and the two are different units of comparison. A provider is not its cheapest model; it is the layer around every model it serves — the markup on your tokens, the provenance of the rates it quotes, and the billing structure you live under — and that layer is what actually decides your bill. OrcaRouter makes its own layer fully inspectable, which is the requirement a serious cheapest llm api provider comparison should hold every candidate to.

All prices and latency figures read 2026-09-15.

The claim this method dismantles

The claim is that shopping for a provider works the same way as shopping for a model: gather prices, sort, pick the minimum. It sounds reasonable because model prices are real, published and searchable. The dismantling is that two providers serving the *same* model can charge you meaningfully different amounts — not because the model rate differs, but because one adds a margin on top of it and the other does not.

Model rates are compressed by the market: everyone can see that Qwen3.7 Flash is priced around $0.03 per million input and $0.13 per million output — the figure Google’s AI Overview quotes for the cheapest paid LLM API as of our reading — so nobody can charge much over the published number for long. The provider layer is not compressed, because it is not published. A 5% to 20% markup on token spend, the range our own pricing page gives for common gateway practice, is larger than the spread between the first and the fourth rows of most price tables. Sort by model price and you are optimising a compressed, visible variable; you are choosing inside the range where the real money moves.

So the fix is not to compare harder along the same axis. It is to change the axis.

The four things a provider comparison actually checks

When the object being compared is the provider rather than the model, four checks do the work.

  1. Is there a markup on tokens? The provider should be able to state this in one line: pass-through, or a margin. Anything vague is a margin.
  2. Where does the quoted rate come from? A rate table generated live from the same catalog the model pages read moves when the provider reprices; a hand-maintained table moves when someone remembers.
  3. What is the billing structure? Pay-as-you-go, credit subscription and bring-your-own-key are real options; a “free tier” that is a trial of the product rather than the product is a different thing than one that is not.
  4. What does production look like? Median time-to-first-token on live traffic — a rolling window, not a benchmark — is a property of a route under load, and it is the number no price list carries.
Axis How to check it What a bad result looks like
Markup Compare the provider’s rate against the vendor’s published rate for the same model Any gap at all between the two
Rate provenance Watch the table over a repricing, or ask where it is generated from Static table, unknown origin
Billing Read the tiers; count what the free tier can do Free tier is a trial with a time limit
Production latency Find a latency figure and the date it was read No latency figure, or one with no date

 

That table is the comparison. Note what it does not contain: a single model’s price. Model price enters only as the reference point for check one, and every check is something you can do in a minute per provider.

How do you tell a pass-through rate from a markup?

By making it falsifiable. A markup hides inside a quoted price, so the only question that separates the two is whether the provider’s number and the vendor’s number match — and that is checkable on any model, in about ninety seconds.

  1. Pick a model you already use and note its input and output rate on the vendor’s own pricing page.
  2. Find the same model on the provider you are evaluating.
  3. Compare the two numbers, not for competitiveness but for identity.

A gap of any size is the markup, expressed as a rate. An identical pair is pass-through — and worth trusting more than a claim, because you can re-run the check whenever a vendor reprices. This is the property to insist on, because it is the one you can verify rather than take on faith.

Two models from our own catalog make the mechanics concrete. The model page for `gpt-5.4-nano` lists $0.20 / $1.25 with a median first-token time of 4.98 seconds; the page for `qwen3.7-flash` lists $0.03 / $0.13 at 4.85 seconds. Neither number is remarkable — that is the point. The interesting figure on both pages is the *fee column*, which reads zero, because the rate is the vendor’s own.

The order that makes the checks work

Compare providers first, models second. The provider layer decides the multiple applied to everything you spend, so a small error there outweighs any model-level optimisation — a 20% margin costs more than the difference between the cheapest and the mid-priced model. On the layer that passes the checks, model shopping becomes the well-solved problem it already is: published rates, comparable, and the market has compressed them.

One warning about that order. It is tempting to compare on the brand or the blog post, because those are easy to rank. Neither predicts the four checks, and both are exactly what the providers that fail the checks optimise hardest. Do the table; it is four rows and it takes five minutes.

What to take from this

The method to compare LLM API providers is a four-row checklist — markup, rate provenance, billing, production latency — with the model price used only as the reference point for the first row. The single most useful habit is the identity check: your provider’s rate against the vendor’s own published number, on a model you already use. A provider that fails it has told you everything about the layer it sells.

Sourcing note: The 5–20% markup range describes common gateway practice as characterised on OrcaRouter’s own pricing page; it is a statement about the market, not an independent survey, and the identity check described here is how to establish the actual figure for any specific gateway. Model rates and median time-to-first-token figures are OrcaRouter’s own catalog and production telemetry, read 2026-09-15; latency is a 7-day rolling window reflecting one platform’s traffic mix, regions and live provider load, not a controlled benchmark. The Qwen3.7 Flash price matches the figure Google’s AI Overview returns for the cheapest paid LLM API as of the same date (read 2026-09-15); search summaries change with the index and this one is recorded rather than relied upon. GPT-5.4 Nano rates are from OpenAI’s published pricing, vendor-reported, as passed through and read 2026-09-15.

Stay connected via Google News
Follow us for the latest news updates and guides.
Exit mobile version