Skip to content
All tool news

Tool desk · Updated daily

Tool news·Vercel·

Vercel Adds Ultrafast Mode to AI Gateway.

Vercel adds Ultrafast mode to AI Gateway, offering faster OpenAI output for supported models at six times the standard token rate.

CW

Create With tool desk

3 sources checked · 2 min read · 2 sections

ShareLinkedIn
Vercel Adds Ultrafast Mode to AI Gateway

Ultrafast mode is now available through Vercel’s AI Gateway, giving builders faster OpenAI output for interactive apps and rapid coding iterations.

The changelog entry lists two supported models: GPT 6 Astra and GPT 6.1 Sol. Builders can request the service tier through Vercel’s AI SDK, Chat Completions API or Responses API.

This extends Vercel’s recent run of AI Gateway updates, with this release focused on response speed rather than routing or fallback logic.

How Vercel Ultrafast Mode Works

To use it through the AI SDK, select openai/gpt-6-astra or openai/gpt-6.1-sol, then set OpenAI’s serviceTier option to ultrafast. The request can also be restricted to OpenAI through the gateway provider settings.

The same service tier is available through the Chat Completions and Responses APIs. If no service tier is specified, the request uses the standard tier.

RequestTier UsedToken Rate
Ultrafast with a supported model and regionUltrafast6× standard
Ultrafast request that falls backTier that serves the requestRate for that tier
Ultrafast pinned to an unsupported regionStandardStandard tier rate
No service tier specifiedStandardStandard tier rate

Ultrafast supports processing in the US and globally. Requests pinned to an unsupported region, including the EU, run at the standard tier instead.

Billing follows the tier that actually handles the request. A request served through Ultrafast costs six times the standard per-token rate. If it falls back to another tier, Vercel charges the rate for that tier rather than the Ultrafast rate.

The update does not require every AI Gateway request to use the faster service. Builders choose it per request, while existing requests without the setting continue on the standard tier.

Frequently asked questions

What is Vercel Ultrafast mode?

Ultrafast mode is an OpenAI service tier available through AI Gateway. It provides faster model output for interactive applications and rapid coding iterations.

Which models support Ultrafast mode?

GPT 6 Astra and GPT 6.1 Sol are currently supported. Their AI Gateway model identifiers are openai/gpt-6-astra and openai/gpt-6.1-sol.

How much does Ultrafast mode cost?

Requests served through Ultrafast are billed at six times the standard per-token rate. Requests that fall back are billed at the rate for the tier that actually serves them.

Sources

3 checked

How we cover tool news: Create With's tool desk drafts these reports with AI from the sources listed above and checks them against those sources before publishing.

Worth passing on?

ShareLinkedIn

Go deeper on Vercel

Related reading, watching and going.

Everything on Vercel →

The briefing

19 May 2026

How to Migrate Your App from Lovable to Claude Code in 2025

Step-by-step guide to migrating your Lovable app to Claude Code. Learn how to move from Lovable's managed platform to a stack you own and control.

Read →

The briefing

10 Feb 2026

OpenClaw Review: 2 Weeks With This AI Agent

Honest 2-week review of OpenClaw — an autonomous AI agent powered by Claude. We test setup, safety, autonomy, and real value to see if it's worth using.

Read →
34 min

Podcast

My Daughter Built a Website Business, Claude Design First Look & Going All-In on Cloudflare

James and Kieran discuss practical AI-powered development through real-world examples, including a seven-year-old building a website business with Claude Code voice mode, Anthropic's new Claude Design tool, API security challenges in the AI agent era, and Cloudflare's expanding AI infrastructure offerings. The episode covers hands-on experiences with vibe coding, emerging design tools, and the evolving landscape of AI-assisted development.

Watch now ↗

Podcast

53. Why Your Vibe Code Needs Tests, Ditching SaaS & AI Agents That Work While You Sleep

James and Kieran discuss the practical realities of building AI-powered software, emphasizing the importance of testing in vibe-coded applications, exploring alternatives to traditional SaaS tools, and implementing autonomous AI agents for business operations. The episode covers technical topics like Convex databases, app store bottlenecks from AI-generated submissions, and using OpenClaw as a marketing automation agent.

Watch now ↗
38 min

Podcast

52. Your AI Co-Founder: OpenClaw, Manus & the Zero-Person Business

In this Create With podcast episode, James and Kieran are joined by Matt Roberts from Happy Operators for an in-depth discussion on using autonomous AI agents in production. The conversation covers practical experiences with OpenClaw, security considerations, comparison with emerging tools like Manus, and the broader shift toward AI-assisted development. They also discuss the VibeCoding Olympics results, Bubble's new AI capabilities, and research on Claude.MD file effectiveness.

Watch now ↗

Latest tool news

What else changed this week.

All tool news

The Create With Briefing

Don't watch forty changelogs. Read one email.

Every Tuesday: the tool changes worth knowing, real business use cases, and what's on near you. Free, unsubscribe any time.