
Uber, AT&T, and Pinterest cut AI costs up to 92% by moving to open models
Uber cut cost per AI request 34% and per session 52%, AT&T saved 56% routing through LiteLLM for a 2% quality drop, and Pinterest now runs at under 8% of closed-model cost by post-training open models on its own data. Open models on inference services run 2-20x cheaper than frontier ones, so a developer can route simpler tasks like code summaries and subagent work to open weights and reserve frontier models for complex generation.
Source: newsletter.pragmaticengineer.com ↗
With open models, we are achieving cost per transaction at less than 8% of the cost of comparable closed proprietary models.
William Ready, CEO of Pinterest
Why this matters
- → Open models deliver 2–20x cost savings vs. frontier models with minimal quality loss
- → Major platforms (Uber, Pinterest, AT&T) prove production-scale feasibility at enterprise scale
- → Strategic model routing lets teams preserve frontier capacity for complex tasks only
Open models go mainstream