Press Esc to close

OpenAI's Decisions API opens to everyone, making snap calls 10x faster

OpenAI's Decisions API opens to everyone, making snap calls 10x faster
OpenAI Developers art for the Decisions API docs: white Decisions title and OpenAI logo on a pink and orange gradient

Image: OpenAI

OpenAI has opened its Decisions API to every developer in public beta. It's a new endpoint built for one job: helping apps make quick calls, like which team should get a support ticket, whether a photo shows damage, or which button an agent should press next.

OpenAI says it answers about 10 times faster than sending the same work to GPT-6 Luna through its regular Responses API. Luna is the only model available for now, and OpenAI expects the API to leave beta "in the coming weeks."

That's a quick turnaround from DevDay, where the Decisions API appeared as a limited preview for selected customers. OpenAI's DevDay recap promised a broad release "in the coming days," and it arrived about a week later.

The idea is simple. Instead of asking a chatbot to write a reply and hoping it sticks to your format, you send text, images or both, plus a list of questions. Each question comes in one of three types.

A "predicate" returns the probability that something is true, such as "does this product have a crack or dent?" A "choice" picks one option from a list you supply, like billing, technical or shipping. A "score" rates something against levels you define, such as how severe a bug is.

The smartest part is that you don't just get an answer, you get odds. Choice and score answers come with a probability for every option plus a separate confidence number. So an app can handle the clear-cut cases automatically and send shaky ones to a person. OpenAI suggests setting those cut-offs using labeled examples from your own app.

Pricing is the other twist. Decisions charges only for input, at $0.10 per million tokens, with no output or caching fees. The same GPT-6 Luna model used through the Responses API costs $0.10 per million input tokens plus $0.50 per million output tokens.

In plain numbers, sorting a million 500-token support tickets would cost about $50 in input charges, before any regional processing premium or long-context surcharge.

Why does speed matter? Apps and AI agents make lots of small choices behind the scenes: which tool to use, whether a request looks risky, which model should take a task. Every slow step adds lag. OpenAI's own list of early uses reads exactly like that, from routing requests to flagging risky tool calls and picking buttons from screenshots.

Voice is another fit. OpenAI's docs show Decisions working with GPT-Live through the Live API. The voice assistant keeps talking and listening while the app asks Decisions which action to run, such as reloading a web page.

There are limits. Decisions doesn't write anything. OpenAI says to stick with Structured Outputs or function calling when you need a model to fill in fields, explain itself or call a tool. Images must be sent as inline base64 data, not links or uploaded files, and questions that depend on an earlier answer need separate requests.

For businesses, it supports Zero Data Retention and HIPAA use for eligible customers, with data residency and regional processing in the US and Europe.

One caution: the 10x claim is OpenAI's own, measured against its own API, and the docs don't list response times in milliseconds. Independent tests under real traffic will tell the full story. Developers can try it now in the Playground.

Comments