Tech • AI • Robotics • Game

VIDEO
ENFR

Stop Overpaying for Intelligence | DevDay 2026

7/10
AIOpenAIOctober 7, 2026 at 09:10 PM23:01
Audio player
0:00 / 0:00

TL;DR

Optimizing AI spend depends less on token list prices than on cost per task, with major savings coming from model right-sizing, prompt caching, tool orchestration in code, tuned reasoning effort, and delayed batch processing.

KEY POINTS

Cost per task matters more than cost per token

A lower token price does not guarantee lower real-world cost. A weaker model may consume more tokens, fail to complete the job, or require a human to step in, raising the true cost of finishing the task. In customer support, for example, a failed chatbot interaction can add both model usage and human labor to the same request.

Model quality can lower overall spend

OpenAI pointed to cases where more capable models reduced cost on a task basis by completing work more efficiently. Perplexity reported that 6 Astra produced 9% more accurate research results at half the cost versus an earlier model. Notion similarly reported that moving from 5.5 to 5.6 Sol improved accuracy while cutting cost per task by half.

Developers are urged to define accuracy targets first

The recommended starting point is to specify the task clearly and set a business threshold for acceptable performance, such as resolving 80% of support requests automatically. Teams can then test models, prompts, and settings against that benchmark and choose the lowest-cost configuration that still meets the quality bar.

A Pareto approach is recommended for model selection

The trade-off to optimize is accuracy versus cost. Teams are advised to evaluate several models and configurations, then plot a Pareto frontier to find the point where performance is good enough at the lowest practical cost. Simpler jobs such as classification or field extraction may fit smaller models like Luna, while more complex planning or reasoning may justify 6 Astra.

Smaller models can handle most workflows

Luna was highlighted as a model that can outperform larger previous-generation systems when paired with high reasoning settings. OpenAI said some customers use Luna for 80% to 90% of developer workflows, reserving larger models only for tasks that truly need higher intelligence. A newly announced Decisions API was also positioned for simpler tasks.

Prompt caching is one of the biggest cost levers

Repeated instructions and shared context can be reused instead of reprocessed on every call. Cached inputs can cut cost by up to 90%, while also reducing latency. The main advice is to keep the beginning of prompts as consistent as possible so stable instructions and policy text can be reused across requests.

Programmatic tool calling cuts unnecessary model work

Instead of sending every intermediate tool result back into the model, applications can use code to coordinate tool calls, compare outputs, and return only the final packaged result for judgment. Legal software company Clio reported 38% fewer prompt tokens in a multi-step document analysis workflow after adopting this approach, with no quality loss.

Reasoning effort should be tuned, not maximized

Higher reasoning effort can be worth paying for on difficult or differentiated tasks, but it should be treated as a testable setting rather than a default. For routing and other straightforward jobs, the recommendation is to start low and increase only if evaluations show quality problems. Short outputs can still require substantial reasoning, so answer length is not a reliable proxy.

Batch processing can sharply reduce price when latency is flexible

For non-urgent work such as processing large document sets, Batch API jobs can run in a 24-hour window at 50% lower token pricing than synchronous requests. This makes delayed processing a practical way to lower spend without changing the underlying application logic.

New cache controls aim to improve reuse

With newer model families, developers can set explicit breakpoints to cache stable prompt sections while allowing changing task details to remain dynamic. An explicit-only mode offers tighter control over exactly what gets cached. OpenAI also said newer systems can change reasoning effort during a conversation without breaking cache reuse, and tools can be changed without invalidating cached prompt sections.

Customer results show the impact of cache optimization

Software development company Blitzy reported that switching from a single structured-output call to Luna’s tool-calling loop lifted cache reuse from 24% to 90%. Across production traffic, the company reported 8.5x fewer output prompt tokens, 87% lower cost than GPT 5.4 Mini, and support for 2.2x more context.

CONCLUSION

The central message is that AI cost control depends on measuring successful outcomes, not just token prices. Teams that define quality targets, evaluate models systematically, and reduce unnecessary model work can cut costs sharply without sacrificing performance.

Ask a question

More from AI