Daily Podcast full article
Introducing the Decisions API
OpenAI’s Decisions API has moved into public beta, turning text and image inputs into typed, low-latency answers on GPT-6 Luna. The launch matters less because it adds another chatbot surface than because it formalizes a narrower pattern: give the model evidence, ask a bounded question, and let software act immediately on the selected option.

A new endpoint for bounded choices
OpenAI has released the Decisions API in public beta, making it available to developers as a dedicated way to choose a model, tool, category, score, or action in near real time . The API is powered by GPT-6 Luna and accepts both text and image inputs, but its purpose is deliberately narrower than a general conversation or reasoning endpoint: it turns evidence into typed answers that an application can use immediately .
The official API changelog records the beta release on October 6, 2026, under gpt-6-luna and v1/decisions, describing it as a way to turn text and images into typed answers “10x faster than the Responses API” . That speed claim is central to the product’s positioning. Instead of asking a model to write a full response, maintain a dialogue, or call arbitrary tools, a developer gives the endpoint a question and a controlled output space. The model’s job is to select, estimate, or score within that space.
The immediate implication is architectural. Many production AI systems do not need a paragraph of prose at every step. They need a fast answer to a constrained question: Is this refund request urgent? Which support queue should receive this ticket? Does this product photo show visible damage? Which action should a voice agent take next? The Decisions API packages that recurring pattern into a first-class API surface.
What the API returns
OpenAI’s public beta announcement describes three supported output types: predicates, choices, and scores . A predicate estimates the probability that a statement is true; a choice selects from predefined options and can include confidence scores; a score evaluates input against a numeric range . Unite.AI’s report on the release adds that the new endpoint is served through POST /v1/decisions and returns typed answers to developer-defined questions over text and image inputs .
That structure is important because it separates “AI judgment” from “AI freedom.” In a conventional chat or agent flow, developers often have to constrain a model with schemas, system prompts, or post-processing logic. With Decisions, the finite answer set is part of the product model. The application defines what may be chosen; the model interprets the input and selects within the permitted frame.
This design also makes the API easier to reason about operationally. A customer-service platform can route a complaint to billing, technical support, shipping, fraud, or “other.” An image workflow can classify a photo as damaged, not damaged, or uncertain. A live assistant can decide whether to reload a page, go back, do nothing, or escalate to a reasoning model. These are not open-ended conversations; they are decision points inside larger software systems.
Why speed comes from limiting the answer space
The launch demos emphasized that the performance gain is not magic so much as scope control. Orply’s summary of the OpenAI demonstration reports that the API is designed for applications needing quick choices from limited action sets rather than open-ended responses . In one sales-routing demo, the status display showed 81 milliseconds of API processing for six decisions . That figure fits the product’s pitch: if the application only needs a bounded decision, it should not pay the latency cost of a broader reasoning call.
The same demonstration also showed the text path through sales leads and form-filling, where unstructured order or lead information could be classified and routed into a structured workflow . The visual path used image frames from a driving-style game, with the API choosing among lanes as obstacles appeared . These examples show the same underlying contract: the app supplies the context and the options; GPT-6 Luna supplies the fast judgment.
The practical difference is especially relevant for interactive products. A voice interface, robot, browser assistant, or game loop cannot always wait for a long chain of reasoning. Even a few hundred milliseconds can change how responsive a product feels. By narrowing the output to a small set of alternatives, Decisions aims to make model judgment feel like part of the interface rather than a visible pause.
Text, images, voice, and agents
The API’s multimodal support gives it broader use than text classification alone. OpenAI’s announcement says the endpoint accepts text and image inputs . Unite.AI reports that independent questions can share a single request and that different question types can be combined, while dependent decisions should be sent as separate requests so the first result can gate the follow-up . That distinction matters for application design: batch independent judgments together, but keep conditional workflows explicit.
Voice-driven applications are another natural fit. The current documentation covered by recent reporting describes using Decisions with voice flows, where client delegation lets an agent choose actions from voice requests and report results back to the user . In practice, that means a live model can handle conversation while the Decisions API handles small control choices: which UI action to take, which expression to show, whether to continue, or when to escalate.
The demos also put the API near robotics and physical interaction. Orply describes a Microduck robot demonstration in which GPT-Live-1 handled voice intelligence while the Decisions API used camera frames to decide where the robot should look . The point is not that a bounded-choice endpoint becomes a full robotics stack. It is that many embodied or live systems contain frequent micro-decisions, and those decisions benefit from being fast, explicit, and limited.
Pricing and developer reaction
The early developer ecosystem moved quickly. Simon Willison released llm-openai-decisions version 0.1a0 on October 6, describing it as a plugin for the OpenAI Decisions API . He noted that the new gpt-6-luna decision model supports image input as well as text and compared its conceptual shape to Jev-style decision models . His post also highlighted pricing: OpenAI’s Decisions API charges 10 cents per million input tokens, while output is not the billed unit in his comparison .
Unite.AI separately reported the same $0.10 per million input tokens figure for gpt-6-luna on the Decisions endpoint and said only input tokens are billed, with no cache-read, cache-write, or output-token charges . That pricing model reinforces the intended usage pattern. If a decision endpoint is meant to run frequently inside workflows, developers need to understand cost per input-heavy classification or scoring call rather than cost per generated answer.
The comparison with Jev is useful but should not obscure the OpenAI-specific angle. Decisions sits inside the OpenAI platform, runs on GPT-6 Luna, and connects to the broader stack of Responses, Live, tools, and agent-oriented APIs. For teams already building on OpenAI, the new endpoint is less a standalone novelty than an optimization layer for repeated, bounded judgments.
What developers should watch next
The public beta is not the end state. Unite.AI reports that OpenAI expects the Decisions API to reach general availability in the coming weeks . Until then, teams should treat it as a new production candidate rather than a solved abstraction. The hard questions are the same ones that apply to any decision system: Are probabilities calibrated on the developer’s own data? What happens at low confidence? Which mistakes are costly? When should a human or a larger reasoning model take over?
The strongest early use cases are therefore not vague “AI decides everything” scenarios. They are narrow, measurable workflows: routing inbound requests, classifying images, selecting an interface action, scoring severity, or choosing the next step in a voice or agent loop. The API’s real promise is that it gives developers a cleaner primitive for those moments when an app does not need a speech, a plan, or a chain of thought. It needs one fast, bounded answer.
That is why the Decisions API feels like a small but meaningful shift. It acknowledges that intelligent applications are built from many kinds of model calls. Some are deep reasoning steps. Some are generative responses. Some are tool calls. And now, OpenAI is giving developers a dedicated fast lane for the humble but crucial act of choosing among predefined options.
Sources from the last 72 hours
- [1]Decisions API is now available in Public BetaOct 6, 2026, 10:53 PM
- [2]Changelog | OpenAI APIOct 6, 2026, 2:00 AM
- [3]OpenAI Releases Decisions API in Public Beta, Powered by GPT-6 LunaOct 6, 2026, 2:00 AM
- [4]A Small Set of Choices Makes Model Decisions Nearly 10 Times FasterOct 6, 2026, 2:00 AM
- [5]llm-openai-decisions 0.1a0Oct 7, 2026, 1:04 AM
AI-generated article based on recent web research, then preserved as a dated editorial snapshot.

Comments
Be the first to comment.