OpenAI DevDay 2026

OpenAI DevDay 2026: What shipped and how to use it

DevDay 2026 was less about one new model than about turning AI into a persistent work layer. Here is what shipped, where it fits, and why usage transparency matters as API and Codex products converge.

Published
Reading time
9 min read
Author
Umer Farooq

Short answer

What matters most

OpenAI DevDay 2026 was a shift from chat as a destination to AI as a persistent work layer. The headline releases are Dots, GPT-6.1 Sol, faster Astra access, cloud Codex, an Agents API with computer use, Decisions API, plugin extensions, and shared ChatGPT workspaces. For most teams, the sensible path is to start with one bounded workflow, use Sol or Codex where code and tools matter, and measure usage and performance instead of assuming that a plan label tells the whole story.

What changed at OpenAI DevDay 2026?

OpenAI says DevDay 2026 included more than 20 major announcements across ChatGPT, Codex, its models, and new ways of working with AI. The common thread was not simply a larger model. It was a move toward agents that can keep working, developer tools that can act inside real systems, and ChatGPT surfaces where people, agents, and software can share context.

That makes this DevDay easier to understand as a platform update than as a single launch. The useful question is not which feature looked most impressive on stage. It is which new layer belongs in the workflow you already need to improve.

What are Dots and ChatGPT Space for?

Dots are OpenAI's always-on agents for ongoing responsibilities. They are designed to learn what matters to a person and take work off their plate over time. ChatGPT Space is the shared layer around that idea: teammates, ChatGPT, and a Dot can work from common knowledge and keep a project organized.

Pages extend the shared surface into collaborative documents where people and agents can write, research, create charts, generate images, and exchange feedback. Collaborative slides, team tasks, and scheduled work push the product further toward a workspace rather than a sequence of isolated chats.

  • Use a Dot for recurring coordination, monitoring, or preparation where a person can review the result.
  • Use ChatGPT Space when a team needs shared context, project memory, and a place to pick work back up.
  • Use Pages for an artifact that needs contribution and review, not just a one-off answer.
  • Keep high-impact decisions behind an explicit human checkpoint until the workflow has evidence that it is safe.

What did developers get in GPT-6.1 Sol and the new APIs?

GPT-6.1 Sol is the main model release. OpenAI describes it as a major upgrade to GPT-6 Sol with strong performance on agentic coding, computer use, and professional work. The headline pricing claim is that it delivers near-Astra intelligence at one-fifth of Astra's standard input and output token prices. OpenAI lists it for API, Plus, Pro, Business, Enterprise, and Edu users.

Ultrafast is a separate premium speed tier. OpenAI says it can reach up to 300 tokens per second, or up to 8x faster generation in Codex and up to 6x in the API. GPT-6 Astra Ultrafast is available today in the API and in ChatGPT Work and Codex on Pro 500 and Enterprise plans. GPT-6.1 Sol Ultrafast is coming later.

The Decisions API is aimed at smaller, sharper decisions. A developer supplies text or images and a finite set of questions with predefined answers, then uses the result to classify content, route a request, or choose an agent's next action. It is in limited preview, which makes it a good candidate for a narrow evaluation rather than a foundation for a critical workflow on day one.

The Agents API adds computer use, tool search, tool calling, context compaction, and Codex-style multi-agent capabilities. Bedrock Managed Agents brings the core agent capabilities into AWS so teams can connect them to AWS resources and run the managed experience in that environment.

  • Choose GPT-6.1 Sol when a complex workflow needs stronger reasoning but still has a meaningful cost budget.
  • Choose Ultrafast when latency is the bottleneck and the workload can justify a premium speed tier.
  • Choose Decisions API for bounded classification and routing, not open-ended reasoning.
  • Choose Agents API when the system must use tools or interact with software, then define the permissions and human handoffs before expanding its scope.

How can people use the Codex updates?

Codex can now run on a computer, remotely from a phone, or in the cloud from any device. Reusable development environments are meant to make tasks start quickly and give a team a shared setup with approved settings and permissions. That is useful for longer-running work, but only when the environment is reproducible and the task has a clear stopping point.

The refreshed Codex CLI adds voice control, a `/agents` view for delegating and tracking work, better prompt editing, resumable sessions, worktrees, and a cleaner terminal interface. The new code review experience in the ChatGPT desktop app can summarize changes, explore diffs, and help investigate a pull request or merge request. Automatic reviews can also run in the cloud while the developer is away.

Codex Security Cloud is the more specialized release. It can scan GitHub repositories on demand or on a schedule, keep checking new commits, investigate findings, remove duplicates, and prepare fixes in the cloud. This is valuable for teams that want security checks to continue after a laptop is closed, but it also raises the bar for permissions, audit trails, and review before a fix is accepted.

  • Use Codex Cloud for tasks that benefit from a persistent environment or background execution.
  • Use `/agents` when work can be split into independent tasks with clear ownership and review points.
  • Use Code Review as a second pair of eyes, not as a replacement for understanding the risk of a change.
  • Use Security Cloud for continuous checks only after repository access, scheduling, and patch approval rules are explicit.

What do plugins, MCP events, and ChatGPT collaboration add?

Plugin extensions let developers give an experience a home in the ChatGPT sidebar, add interactive panels, and create viewers for supported file types. OpenAI also introduced Plugin Creator, a clearer submission flow, and improved discovery. Users still choose which plugins to use and approve the access each plugin receives.

MCP events let plugin automations start when something happens in a connected system. A new project-board task, for example, can trigger ChatGPT to read linked material and draft a plan. Sites can host supported plugins for teams, while @ChatGPT in Slack and Microsoft Teams brings the work into the conversations where it is already happening.

Meetings can turn a conversation into notes and action items, and Sign in with ChatGPT can carry an eligible allowance across participating tools. That convenience has a cost that deserves attention: OpenAI says eligible usage through partners counts toward the plan's limits, so connected tools are part of the same usage story rather than free extra capacity.

  • Build a plugin when the value comes from an interactive tool or document view inside the conversation.
  • Use MCP events for event-driven work that has a clear trigger, scope, and owner.
  • Use Slack or Teams access when the team needs answers in place, but keep permissions aligned with the connected workspace.
  • Treat cross-product sign-in as a shared budget decision and review which tools can consume the allowance.

Which DevDay release should you try first?

The best first release depends on the work, not the novelty of the announcement. Start with the smallest workflow that can show whether the new capability is useful, safe, and affordable. This is the same principle I use when choosing an AI workflow to automate: begin with repeated work, accessible information, a clear owner, and a safe handoff.

  • For software teams: start with GPT-6.1 Sol or Codex Cloud on a bounded issue, with tests and review kept in the loop.
  • For operations teams: evaluate Decisions API on a finite routing or classification task before adding open-ended agent behavior.
  • For product teams: prototype an Agents API workflow with one or two tools and measure successful completion, not just generated text.
  • For security teams: pilot Security Cloud on a repository where findings and proposed fixes already have an owner.
  • For knowledge work: use Space, Pages, or a Dot for a recurring project with clear source material and human review.

Are OpenAI's usage limits fair?

OpenAI's current documentation makes the accounting difference clear. ChatGPT plans that include Codex use a plan allowance or shared credit pool, and Codex, ChatGPT Work, ChatGPT for Excel, and other agentic features can draw from the same pool depending on the plan. The amount consumed depends on the model, where the task runs, task complexity, context, reasoning, speed, and tools. API usage follows token pricing and rate limits instead.

A usage limit is not automatically unfair. Agentic work can consume much more compute than a short request, and a service needs a way to protect capacity. The problem is when the meter is difficult to understand, the reset is unpredictable, or a paid user cannot tell whether a task used more because it was genuinely harder or because the product silently changed its routing and reasoning behavior.

My view is that fair usage needs to be predictable before it needs to be generous. OpenAI should show the remaining allowance, the exact reset time, the model and reasoning settings that affected consumption, and a useful credit or token-equivalent explanation for long tasks. When a model or default changes, the product should say so and preserve a reasonable way to finish work already in progress.

  • Show consumption in a way users can compare across models and tools.
  • Explain how context length, reasoning, speed tiers, and tool calls affect the meter.
  • Make reset times and available top-up options visible before a task starts.
  • Keep a task's status and remaining path clear when it approaches a limit.
  • Publish enough usage detail for users to distinguish capacity protection from unexplained quality or routing changes.

Should API and Codex subscriptions perform the same?

The same model name does not guarantee the same end-to-end product. The API gives developers control over the model, prompt, tools, reasoning settings, and billing. Codex adds an agent harness, repository context, file permissions, worktrees, execution tools, and product-specific defaults. OpenAI's model guidance also warns that availability, tools, reasoning settings, and usage limits differ by product and model version.

That explains why two users can run what looks like the same task and get different results without either person imagining the difference. A different system prompt, context assembly, tool policy, retry strategy, model snapshot, or reasoning budget can change the outcome. The public documentation does not establish that Codex subscriptions are inherently lower quality than the API, so the claim should be tested rather than repeated as fact.

The complaint is still reasonable. When OpenAI advertises the same model across the API and Codex, users should expect a comparable capability baseline at an equivalent model and reasoning setting. The products do not need identical pricing, limits, or interfaces, but the differences should be visible and explainable.

The fairest solution would be a parity matrix with the model snapshot, reasoning setting, tool access, context policy, rate or credit cost, and representative evaluation results for both surfaces. Let users compare successful task completion, time to result, tokens or credits consumed, and failure or retry rates. That would turn a recurring argument about performance into an engineering question with evidence.

The practical takeaway

DevDay 2026 makes OpenAI's direction clear: persistent agents, shared work surfaces, more capable coding tools, and APIs that can act inside software. The opportunity is real, but the implementation still starts with ordinary engineering discipline: a narrow job, trusted context, explicit permissions, measurable success, and a human path for uncertain cases.

For builders, I would start with GPT-6.1 Sol or Codex Cloud on one workflow and keep the API available when you need precise control. For subscribers, I would judge the new experience by the work it completes per unit of allowance, not by the plan multiplier alone. Better tools deserve a clear meter. Better models deserve a transparent comparison.

Which business workflow should you automate with AI first?A practical technical SEO checklist for an AI engineering siteHow to write content that earns a useful answer

Sources

Direct to Umer

Let's talk about your project.

Tell me what you're trying to solve. I'll reply directly by email.