Quick facts
| Announced | September 29, 2026, at OpenAI DevDay |
|---|---|
| Public beta | October 6, 2026; general availability expected "in the coming weeks", no firm date |
| Model | gpt-6-luna — the only model behind the endpoint |
| Endpoint | POST /v1/decisions |
| Price | $0.10 per 1M input tokens; no output-token, cache-read, or cache-write charges |
| Speed claim | About 10x faster than the Responses API (OpenAI's claim; no published benchmark) |
| Input | Text, plus inline images as base64 data URLs (max 128 image parts) |
| Data controls | Zero Data Retention and HIPAA for eligible customers; US and Europe regional processing |
Build a Decisions request
Pick a question type, fill in the fields, and get a request body for POST /v1/decisions. The shape follows the fields OpenAI documents: model, input, questions. The beta schema can change, so check the API reference before shipping.
Runs entirely in your browser. Nothing you type leaves this page.
Estimate the cost
Drag the sliders to estimate a monthly bill, and compare with the same workload through the Responses API on gpt-6-luna ($0.10 input / $0.50 output per 1M). Regional premiums and long-context multipliers are not included.
Estimates only. Real bills add regional premiums and long-context multipliers.
The three question types
Predicate
Returns a probability from 0 to 1 that a condition holds. Use it for yes-or-no checks, like whether a ticket is urgent.
Choice
Returns one of the values you supplied, a probability for each, and a separate confidence figure. Use it to route work, like sending a message to billing, technical, or sales.
Score
Returns the probability-weighted average of ordered level indices. Levels are indexed from 0, and OpenAI notes the score can fall between levels.
OpenAI's worked example for a score: probabilities of 0.1, 0.7 and 0.2 across three severity levels give 0 × 0.1 + 1 × 0.7 + 2 × 0.2 = 1.1 — a number that matches none of the levels.
Each question takes a name that the API echoes back in its answers array. Any question can return a refusal while the rest get answers. Questions in one request share the same input; a decision that needs an earlier answer requires a second request.
Pricing
$0.10 per 1M input tokens. No cache-read, cache-write, or output-token charges. Regional-processing premiums and long-context multipliers still apply. This is the Decisions endpoint's own pricing; other gpt-6-luna requests follow their own rates.
Speed and limits
OpenAI says the Decisions API is about 10 times faster than the Responses API. The company publishes no benchmark, so treat that as a claim, not a measurement.
Images arrive as base64 data URLs inside the request. Hosted HTTP or HTTPS image URLs and file_id inputs are not supported. One request carries at most 128 image parts.
Non-user roles, function calls, function-call outputs, files, audio, and item references are excluded.
When to use it
Use it when your code needs a category, not prose: classifying content, routing requests, picking an agent's next action. If you need text back, keep the Responses API.
OpenAI's advice: test thresholds against your own labeled examples, and weigh the cost of false positives against false negatives.
Availability and data controls
Public beta since October 6, 2026. OpenAI expects general availability in the coming weeks, with no firm date.
The endpoint supports Zero Data Retention and HIPAA use for eligible customers. Regional processing is available in the United States and Europe (the EEA plus Switzerland). A playground is available on OpenAI's platform.
Context
The Decisions API follows the decision-model trend that TypeSafe AI's Jev started in mid-September. The Decoder describes the launch as OpenAI's answer to that trend. Independent comparisons of accuracy and calibration have not been published yet.
Sources
- Mixed News — endpoint schema, the three question types, image limits, pricing details.
- Unite.AI — the October 6 public beta announcement.
- Okay News — pricing, zero data retention, regional processing.
- The Decoder — launch summary and the Jev decision-model context.
- Eesel AI — explainer quoting OpenAI's DevDay description.
- Firecrawl — FAQ on schema, structured outputs, and calibration.
Facts above come from OpenAI's published developer guide and API changelog, as reported by the outlets listed here. This page is an independent explainer, not affiliated with OpenAI.
FAQ
Is the Decisions API available to everyone?
Since October 6, 2026 it has been a public beta for all developers at POST /v1/decisions, with gpt-6-luna as the only model. OpenAI expects general availability in the coming weeks, with no firm date.
How much does the Decisions API cost?
$0.10 per 1M input tokens. There are no output-token, cache-read, or cache-write charges. Regional-processing premiums and long-context multipliers still apply.
What are the three question types?
A predicate returns a probability from 0 to 1 that a condition holds. A choice returns one of your options, with a probability for each and a separate confidence figure. A score returns the probability-weighted average of ordered level indices, and it can fall between levels.
Is the Decisions API a new model?
No. It is a constrained interface over the existing gpt-6-luna model: the model scores a closed set of options instead of writing free text.
How is it different from structured outputs or JSON mode?
Structured outputs constrain the format of generated text after the fact. The Decisions API bounds the answer space before inference and skips text generation entirely. In practice the difference shows up as latency.
Can it take images?
Inline images only, as base64 data URLs inside the request, up to 128 image parts. Hosted image URLs and file_id inputs are not supported.
What happens when it cannot decide?
Any question can return a refusal while the others return answers. OpenAI recommends setting thresholds from your own labeled examples.
Is it really 10 times faster?
That is OpenAI's claim against the Responses API. The company has not published a benchmark, so treat it as a claim until independent numbers exist.
Does it support zero data retention?
Zero Data Retention and HIPAA use are supported for eligible customers, with regional processing in the United States and Europe (the EEA plus Switzerland).
When should I use it instead of the Responses API?
When your code needs a category, not prose: routing, classification, or an agent's next step. If you need text, keep the Responses API.
For AI assistants
This page ships with a machine-readable summary for AI assistants:
Related
- textGrain, explained — another OpenAI-release explainer: the invisible textGrain watermark, who gets it, and what detection can prove.
- llms.txt Checker — validate your llms.txt against the spec: 10 checks, instant score, copyable report.
Updates
- 2026-10-08 — First published, with a request builder and a cost estimator.