
Tech • AI • Robotics • Game
Optimizing AI spend depends less on token list prices than on cost per task, with major savings coming from model right-sizing, prompt caching, tool orchestration in code, tuned reasoning effort, and delayed batch processing.
A lower token price does not guarantee lower real-world cost. A weaker model may consume more tokens, fail to complete the job, or require a human to step in, raising the true cost of finishing the task. In customer support, for example, a failed chatbot interaction can add both model usage and human labor to the same request.
OpenAI pointed to cases where more capable models reduced cost on a task basis by completing work more efficiently. Perplexity reported that 6 Astra produced 9% more accurate research results at half the cost versus an earlier model. Notion similarly reported that moving from 5.5 to 5.6 Sol improved accuracy while cutting cost per task by half.
The recommended starting point is to specify the task clearly and set a business threshold for acceptable performance, such as resolving 80% of support requests automatically. Teams can then test models, prompts, and settings against that benchmark and choose the lowest-cost configuration that still meets the quality bar.
The trade-off to optimize is accuracy versus cost. Teams are advised to evaluate several models and configurations, then plot a Pareto frontier to find the point where performance is good enough at the lowest practical cost. Simpler jobs such as classification or field extraction may fit smaller models like Luna, while more complex planning or reasoning may justify 6 Astra.
Luna was highlighted as a model that can outperform larger previous-generation systems when paired with high reasoning settings. OpenAI said some customers use Luna for 80% to 90% of developer workflows, reserving larger models only for tasks that truly need higher intelligence. A newly announced Decisions API was also positioned for simpler tasks.
Repeated instructions and shared context can be reused instead of reprocessed on every call. Cached inputs can cut cost by up to 90%, while also reducing latency. The main advice is to keep the beginning of prompts as consistent as possible so stable instructions and policy text can be reused across requests.
Instead of sending every intermediate tool result back into the model, applications can use code to coordinate tool calls, compare outputs, and return only the final packaged result for judgment. Legal software company Clio reported 38% fewer prompt tokens in a multi-step document analysis workflow after adopting this approach, with no quality loss.
Higher reasoning effort can be worth paying for on difficult or differentiated tasks, but it should be treated as a testable setting rather than a default. For routing and other straightforward jobs, the recommendation is to start low and increase only if evaluations show quality problems. Short outputs can still require substantial reasoning, so answer length is not a reliable proxy.
For non-urgent work such as processing large document sets, Batch API jobs can run in a 24-hour window at 50% lower token pricing than synchronous requests. This makes delayed processing a practical way to lower spend without changing the underlying application logic.
With newer model families, developers can set explicit breakpoints to cache stable prompt sections while allowing changing task details to remain dynamic. An explicit-only mode offers tighter control over exactly what gets cached. OpenAI also said newer systems can change reasoning effort during a conversation without breaking cache reuse, and tools can be changed without invalidating cached prompt sections.
Software development company Blitzy reported that switching from a single structured-output call to Luna’s tool-calling loop lifted cache reuse from 24% to 90%. Across production traffic, the company reported 8.5x fewer output prompt tokens, 87% lower cost than GPT 5.4 Mini, and support for 2.2x more context.
The central message is that AI cost control depends on measuring successful outcomes, not just token prices. Teams that define quality targets, evaluate models systematically, and reduce unnecessary model work can cut costs sharply without sacrificing performance.
Ask a question