Uber has entered a phase of keeping costs under control even as AI usage surges.
Axios reported on Aug. 27 that Uber's weekly agent requests have risen 9.4 times since February, but total AI spending has stayed steady since April.
A blog post by Uday Kiran Medisetty (우다이 키란 메디세티), a senior engineer at Uber, shows that the cost per 1,000 requests for the same AI model fell nearly 34 percent from an April peak. Cost per session fell 52 percent from a June high.
AI agents now handle more than 70 percent of code change submissions. Uber engineers run more than 30,000 AI agent tasks a day, and the number of AI tool users has risen more than fourfold. Even so, token costs have fallen.
Uber introduced a system that assigns tasks to suitable models while considering both cost and performance. It routes smaller tasks to relatively cheaper models. It capped conversational session tokens at 400,000. It applies that limit even if a model in use can handle up to 1,000,000 tokens.
Engineers can check real-time costs by session directly in the terminal. Uber also extended prompt cache retention time to 1 hour from 5 minutes, because engineers often left sessions running while away for more than 5 minutes.
A trend is also spreading in which companies cut costs by increasing use of open-weight models. Uber Chief Technology Officer Praveen Neppalli (프라빈 네팔리) also cited experiments with open-weight models as one factor in reducing costs in a post on social media platform X (Twitter). "We are continuously evaluating and deploying models for each use case," he said.
Earlier, Neppalli said in an April interview with The Information that Uber had already used up its 2026 AI budget.