AI inference price arbitrage — every provider serving a given model with live input/output $/Mtoken, blended 3:1 cost, cheapest-vs-dearest spread multiple, quantization caveats. Sa
AI inference price arbitrage — every provider serving a given model with live input/output $/Mtoken, blended 3:1 cost, cheapest-vs-dearest spread multiple, quantization caveats. Same open-weights model spans 2-5x across providers; for agents that buy inference this is self-referential spend recovery. Live listings, no LLM.
price-ascending: candidates are ordered by the listed price of the operation, with no managed-first preference applied. Unlike /v1/apis, this surface does not rank Apiosk-settled listings above federated ones — a comparison that reorders on commercial grounds is not a comparison. Pass sort=managed_first for the catalog's ordering.
| Provider | Operation | Price | Settlement | Notes |
|---|
Compared: listed price, settlement rail and input compatibility. Not measured yet: latency, reliability, result quality, and provider terms such as rate limits, jurisdiction and licensing.