Explainer · Public beta

What is the OpenAI Decisions API?

OpenAI has a new API that does not write sentences. The Decisions API takes a question with a fixed set of answers and hands back a number or a pick. The public beta opened on October 6, 2026, and so far only GPT-6 Luna sits behind it. OpenAI claims it is roughly ten times faster than the regular Responses API. The price is $0.10 per million input tokens, with no fee for output. This page covers the three question types, the limits, and when the endpoint makes sense, based on OpenAI's own docs.

Last verified: October 8, 2026 · Status: public beta, details can change · Independent explainer, not affiliated with OpenAI

Quick facts

AnnouncedSeptember 29, 2026, at OpenAI DevDay
Public betaOctober 6, 2026; general availability expected "in the coming weeks", no firm date
Modelgpt-6-luna — the only model behind the endpoint
EndpointPOST /v1/decisions
Price$0.10 per 1M input tokens; no output-token, cache-read, or cache-write charges
Speed claimAbout 10x faster than the Responses API (OpenAI's claim; no published benchmark)
InputText, plus inline images as base64 data URLs (max 128 image parts)
Data controlsZero Data Retention and HIPAA for eligible customers; US and Europe regional processing

Build a Decisions request

Pick a question type, fill in the fields, and get a request body for POST /v1/decisions. The shape follows the fields OpenAI documents: model, input, questions. The beta schema can change, so check the API reference before shipping.

Runs entirely in your browser. Nothing you type leaves this page.

Estimate the cost

Drag the sliders to estimate a monthly bill, and compare with the same workload through the Responses API on gpt-6-luna ($0.10 input / $0.50 output per 1M). Regional premiums and long-context multipliers are not included.

$15.00 / mo
$21.00 / mo
$6.00 / mo (29%)

Estimates only. Real bills add regional premiums and long-context multipliers.

The three question types

Predicate

Returns a probability from 0 to 1 that a condition holds. Use it for yes-or-no checks, like whether a ticket is urgent.

Choice

Returns one of the values you supplied, a probability for each, and a separate confidence figure. Use it to route work, like sending a message to billing, technical, or sales.

Score

Returns the probability-weighted average of ordered level indices. Levels are indexed from 0, and OpenAI notes the score can fall between levels.

OpenAI's worked example for a score: probabilities of 0.1, 0.7 and 0.2 across three severity levels give 0 × 0.1 + 1 × 0.7 + 2 × 0.2 = 1.1 — a number that matches none of the levels.

Each question takes a name that the API echoes back in its answers array. Any question can return a refusal while the rest get answers. Questions in one request share the same input; a decision that needs an earlier answer requires a second request.

Pricing

$0.10 per 1M input tokens. No cache-read, cache-write, or output-token charges. Regional-processing premiums and long-context multipliers still apply. This is the Decisions endpoint's own pricing; other gpt-6-luna requests follow their own rates.

Speed and limits

OpenAI says the Decisions API is about 10 times faster than the Responses API. The company publishes no benchmark, so treat that as a claim, not a measurement.

Images arrive as base64 data URLs inside the request. Hosted HTTP or HTTPS image URLs and file_id inputs are not supported. One request carries at most 128 image parts.

Non-user roles, function calls, function-call outputs, files, audio, and item references are excluded.

When to use it

Use it when your code needs a category, not prose: classifying content, routing requests, picking an agent's next action. If you need text back, keep the Responses API.

OpenAI's advice: test thresholds against your own labeled examples, and weigh the cost of false positives against false negatives.

Availability and data controls

Public beta since October 6, 2026. OpenAI expects general availability in the coming weeks, with no firm date.

The endpoint supports Zero Data Retention and HIPAA use for eligible customers. Regional processing is available in the United States and Europe (the EEA plus Switzerland). A playground is available on OpenAI's platform.

Context

The Decisions API follows the decision-model trend that TypeSafe AI's Jev started in mid-September. The Decoder describes the launch as OpenAI's answer to that trend. Independent comparisons of accuracy and calibration have not been published yet.

Sources

  • Mixed News — endpoint schema, the three question types, image limits, pricing details.
  • Unite.AI — the October 6 public beta announcement.
  • Okay News — pricing, zero data retention, regional processing.
  • The Decoder — launch summary and the Jev decision-model context.
  • Eesel AI — explainer quoting OpenAI's DevDay description.
  • Firecrawl — FAQ on schema, structured outputs, and calibration.

Facts above come from OpenAI's published developer guide and API changelog, as reported by the outlets listed here. This page is an independent explainer, not affiliated with OpenAI.

FAQ

Is the Decisions API available to everyone?

Since October 6, 2026 it has been a public beta for all developers at POST /v1/decisions, with gpt-6-luna as the only model. OpenAI expects general availability in the coming weeks, with no firm date.

How much does the Decisions API cost?

$0.10 per 1M input tokens. There are no output-token, cache-read, or cache-write charges. Regional-processing premiums and long-context multipliers still apply.

What are the three question types?

A predicate returns a probability from 0 to 1 that a condition holds. A choice returns one of your options, with a probability for each and a separate confidence figure. A score returns the probability-weighted average of ordered level indices, and it can fall between levels.

Is the Decisions API a new model?

No. It is a constrained interface over the existing gpt-6-luna model: the model scores a closed set of options instead of writing free text.

How is it different from structured outputs or JSON mode?

Structured outputs constrain the format of generated text after the fact. The Decisions API bounds the answer space before inference and skips text generation entirely. In practice the difference shows up as latency.

Can it take images?

Inline images only, as base64 data URLs inside the request, up to 128 image parts. Hosted image URLs and file_id inputs are not supported.

What happens when it cannot decide?

Any question can return a refusal while the others return answers. OpenAI recommends setting thresholds from your own labeled examples.

Is it really 10 times faster?

That is OpenAI's claim against the Responses API. The company has not published a benchmark, so treat it as a claim until independent numbers exist.

Does it support zero data retention?

Zero Data Retention and HIPAA use are supported for eligible customers, with regional processing in the United States and Europe (the EEA plus Switzerland).

When should I use it instead of the Responses API?

When your code needs a category, not prose: routing, classification, or an agent's next step. If you need text, keep the Responses API.

For AI assistants

This page ships with a machine-readable summary for AI assistants:

Related

  • textGrain, explained — another OpenAI-release explainer: the invisible textGrain watermark, who gets it, and what detection can prove.
  • llms.txt Checker — validate your llms.txt against the spec: 10 checks, instant score, copyable report.

Updates

  • 2026-10-08 — First published, with a request builder and a cost estimator.