OpenAI DevDay 2026: Everything New in the Agent Stack

DevDay 2026 connected persistent agents, cloud coding, plugins, events, and shared work. Here is what shipped, what is still coming, and where AppSec boundaries lie.

José Palanco José Palanco
Last Updated:
9 min read
Compartir
OpenAI DevDay 2026: Everything New in the Agent Stack

Más allá de ASPM

AppSec basada en pruebas para equipos que desarrollan con IA

Plexicus usa AI Swarm Pentest para explorar rutas de aplicaciones autorizadas, validar qué es explotable y dar a los equipos evidencia para priorizar la remediación.

Explorar AI Swarm Pentest

OpenAI DevDay 2026: Everything New in the Agent Stack

OpenAI’s September 29 DevDay announcements point to a shift in how AI software operates. A model can now sit inside a longer-running workflow: receive a goal, gather context from connected apps, use a browser or coding environment, and continue when new work arrives. The DevDay recap spans consumer, developer, and enterprise updates, but the common thread is an agent that can act across systems. Some announcements are new launches, others expand existing products, and several are previews.

That creates useful automation. It also makes permissions, untrusted input, and evidence of what an agent actually did central application security questions.

Diagram of the DevDay agent stack, from models and tools to connected applications

Original Plexicus illustration of the announced agent stack; it is not an OpenAI product screenshot.

Persistent agents meet connected work

Dots are OpenAI’s always-on agents. A dot has a cloud computer, browser, persistent context, and access to connected applications, subject to configured permissions and approvals. OpenAI describes proactive research as read-only background work; other actions depend on the dot’s broader permissions. This distinction matters: observing a new issue, preparing a change, and executing that change should not be treated as one permission.

OpenAI has also described specialist dots with organizational identities and dedicated access. That is a pilot direction, not a generally available enterprise fleet. For teams considering persistent agents, each identity should have a defined owner, narrow access, and a record of its actions.

GPT-6.1 Sol supplies another part of the stack. OpenAI positions it for coding, computer use, and longer workflows (model documentation). Lower cost per task can make repeated agent runs practical, but a capable model does not establish that its output is secure or that its actions are authorized.

Faster models and private infrastructure

OpenAI also expanded Ultrafast, a premium speed tier. At DevDay it reported up to 300 tokens per second for GPT-6 Astra Ultrafast, with GPT-6.1 Sol Ultrafast planned later (DevDay recap). These are OpenAI’s performance claims, not a prediction for every workload. Speed matters because an agent may perform dozens of sequential reasoning and tool steps. Cutting latency in each step can change whether a workflow feels interactive. It does not reduce the need to check the result.

The privacy announcements target a different constraint. Zero Data Retention with Private Safety Processing is described for eligible API customers who need automated safety processing without OpenAI retaining prompts and responses; the architecture can use customer-managed cloud storage. Private Inference was previewed for later availability, with confidential computing and verifiable controls in the proposed design. Teams should assess the actual service and contract available to them before treating a preview as an existing data-handling guarantee. This becomes material when agents can read source code, internal messages, and customer records.

From coding sessions to ongoing software work

Codex Cloud gives coding agents reusable cloud environments. Combined with code review and Codex Security Cloud, this suggests a workflow in which agents inspect a repository, propose changes, run checks, and review security findings without depending on a developer’s open laptop.

The security boundary is the repository and its surrounding systems: source code, dependencies, secrets, CI, issues, and pull requests. An agent-generated patch still needs review of its affected behavior. A passing test says something about the tested cases; it does not prove that every authorization path or exposed endpoint is safe.

The refreshed Codex CLI adds voice steering, an /agents view, and improvements for tracking and resuming multiple agent tasks (DevDay recap). The important operational change is that one developer may supervise several concurrent agents. That makes task ownership and the provenance of each code change more important than the interface used to start it.

The desktop Code Review experience brings summaries, diffs, questions, and cloud review into the developer’s workflow. Automatic cloud review depends on repository connection and configuration. Codex Security Cloud can scan connected repositories on demand or on a schedule, investigate and deduplicate findings, and prepare fixes while the local machine is offline. These are useful ways to shorten the feedback loop, but findings still need evidence of reachability and impact. Proposed fixes need a retest of the behavior they change.

The Agents API computer-use guide adds another interface. Once an agent can click through applications, a visible page becomes both task context and a possible carrier for malicious instructions. Keep browser sessions scoped, constrain what they can reach, and require review before consequential changes.

The wider Agents API brings orchestration, tool use, context management, MCP connections, and sandboxes into one development surface. OpenAI hosts the browser for computer use, while the developer’s application still controls site access and surrounding workflow. The Decisions API, announced in limited preview, takes a narrower approach: developers give it a finite set of allowed answers and it selects among them (DevDay recap). Bounded outputs can help with routing or classification, but a model-selected answer should not by itself authorize a sensitive operation.

The DevDay recap also highlights Amazon Bedrock Managed Agents powered by OpenAI. This is an expansion of an earlier AWS partnership rather than a technology that first appeared at DevDay. AWS describes agent identity, memory, compute, security controls, and audit logs within its infrastructure. Enterprises adopting it still need to map those controls to their own permissions and review process.

MCP Events adds an inbound path

OpenAI is adding support for the proposed MCP Events specification. Instead of polling a service, ChatGPT can subscribe to events from an MCP server. OpenAI documents signed webhook delivery, event IDs, filters, authorization checks, and protections against feedback loops.

An event can be useful: a new bug report could trigger investigation and a draft fix. But event content originates outside the agent’s instruction boundary. A forged or maliciously written report must remain task data, even when it arrives through a valid webhook. Authentication establishes which service sent an event; it does not make every sentence in its payload safe to obey.

Diagram of an external event entering an agent workflow

Original Plexicus illustration of an MCP Events workflow; it is not an OpenAI product screenshot.

Plugins and their UI extensions widen this path. A plugin may expose tools, data, and interactive surfaces to an agent. Review its requested scopes, provider, data handling, and write actions as you would any other software integration (OpenAI’s plugin security guide).

OpenAI’s plugin documentation now describes packages combining skills, an MCP server, UI, and external connections. Plugin Extensions can provide sidebars, panels, editors, forms, and context shared with a conversation. Plugin Creator and a revised submission flow aim to make building and publishing easier; OpenAI describes a directory shared between ChatGPT and Codex. A larger distribution channel also increases the importance of vetting the software and scopes behind a convenient install button.

Sites can host supported plugins, so a shared AI-powered page can expose connected capabilities to several people. The page may be common, but each user’s data and actions must still follow their own permissions. For developers, this is an application-hosting route. For security teams, it is another place to check the boundary between shared context and individual authority (DevDay recap).

Space turns the stack into shared work

ChatGPT Space is OpenAI’s home for pages, files, spreadsheets, presentations, and agents. Pages let people and agents co-edit a document that can use connected context and update over time. Collaborative slides were previewed as forthcoming, including generation, editing, comments, and export. A shared page that refreshes from tools is useful, but its source data, update instruction, and audience all need to be visible to its owners.

Teams and Team Tasks put scheduled or event-driven work into a shared organizational setting. OpenAI also plans @ChatGPT in Slack and Microsoft Teams, where an agent can participate in a channel or thread using approved tools and relevant user permissions (DevDay recap). This reduces friction between discussion and action. It also raises a straightforward question: whose authority does a request in a group conversation carry?

The Meetings plugin can turn recorded discussion into notes and action items in the macOS app, then feed follow-up work. OpenAI says meeting audio is deleted after the notes are ready. Teams should still consider consent, retention of the resulting notes, and whether a spoken suggestion should become an executable task. Shareable profiles make selected work easier to discover; workspace sharing controls still determine its audience.

Identity, plans, and distribution

Sign in with ChatGPT lets eligible users bring plan usage into participating third-party applications, with controls over each application’s allowance (OpenAI help). It makes ChatGPT both a sign-in option and a portable usage source. The relevant security review is familiar: identify the third party, inspect the access granted, and know how to revoke it.

Pro 500 is the new high-usage individual tier, with access to Astra Ultrafast where supported. The OpenAI Marketplace lets eligible enterprises apply part of their commercial commitment to approved partner software. Both were announced in the DevDay recap. They change how agent capacity and integrations may be purchased; neither changes the need to assess the tools that receive organizational data.

What AppSec teams should verify

The main risk is a chain of individually reasonable capabilities: an agent reads a message, consults a repository, uses a browser, then changes a ticket or opens a pull request. The harm depends on where untrusted content is promoted into authority and what the agent is allowed to do next.

Diagram of trust boundaries across an agent, tools, and connected applications

Original Plexicus illustration of an agent’s attack surface; it is not an OpenAI product screenshot.

For each workflow, security teams should check:

  • Identity and scope: which agent owns each credential, what it can read or change, and when access expires.
  • Input boundaries: whether retrieved pages, messages, event payloads, and tool results remain untrusted data.
  • Action gates: which writes require human approval and whether approval covers the exact action and destination.
  • Traceability: whether logs connect an event to the agent’s decision, tool calls, resulting change, and reviewer.
  • Validation: whether a finding is reachable and exploitable within authorized scope, and whether the fix was retested.

OpenAI’s GPT-6.1 Sol safety addendum reports cybersecurity evaluation results and describes deployment safeguards. Those are vendor-reported benchmark results under specific evaluation conditions, not a measure of how secure every agent deployment will be. The application, permissions, integrations, and human controls still determine the real exposure.

DevDay’s lasting AppSec question is practical: as agents operate across more systems for longer periods, can teams show which actions were permitted, what evidence supported them, and whether the resulting software is secure? Proof-driven verification needs to follow the whole workflow, from the incoming event to the final change.

Escrito por
José Palanco
José Palanco
José Ramón Palanco es el CEO/CTO de Plexicus, una empresa pionera en ASPM (Gestión de Postura de Seguridad de Aplicaciones) lanzada en 2024, que ofrece capacidades de remediación impulsadas por IA. Anteriormente, fundó Dinoflux en 2014, una startup de Inteligencia de Amenazas que fue adquirida por Telefónica, y ha estado trabajando con 11paths desde 2018. Su experiencia incluye roles en el departamento de I+D de Ericsson y Optenet (Allot). Tiene un título en Ingeniería de Telecomunicaciones de la Universidad de Alcalá de Henares y un Máster en Gobernanza de TI de la Universidad de Deusto. Como experto reconocido en ciberseguridad, ha sido ponente en varias conferencias prestigiosas, incluyendo OWASP, ROOTEDCON, ROOTCON, MALCON y FAQin. Sus contribuciones al campo de la ciberseguridad incluyen múltiples publicaciones de CVE y el desarrollo de varias herramientas de código abierto como nmap-scada, ProtocolDetector, escan, pma, EKanalyzer, SCADA IDS, y más.
Leer más de José
More to read

Related posts

¿Listo para validar lo que importa?

Listo para validar lo que importa.

Plexicus es Proof-Driven AppSec: hallazgos validados, comprensión contextual y remediación revisada — anclada en evidencia, acotada contigo.

Calificación

Comprueba si el AI Swarm Pentest encaja en tu entorno.

Déjanos el contexto mínimo. Revisaremos el alcance y te indicaremos el siguiente paso comercial.

Antes de enviar — verifica que encajas

0 / 280

Sin compromiso. Si no encajas, te lo decimos.

SAMPLE HANDOVER · ILLUSTRATIVE

Sample evidence handover

A trimmed view of what your team receives at the end of an AI Swarm Pentest engagement. Real engagements include full technical evidence, executive narrative, and a remediation plan.

VALIDATED FINDING Evidence attached

Server-Side Request Forgery in webhooks/receiver

demo-project/sample-app · src/webhooks/receiver.py:42

SeverityHigh CVSS 3.18.6 Priority79 Confirmedvia replay

Untrusted caller-supplied URLs reach an internal egress without an allowlist. Replayed in a sandbox against a fresh authorized target — the same control was validated to fail twice.

REVIEWER-READY REMEDIATION Merge-ready PR

Validate the target URL against an allowlist of permitted hostnames. Reject private/internal IP ranges. Enforce HTTPS only.

plexicus/remediation/webhooks-ssrf 3 changed · 0 new files
42resp = requests.get(target_url)
42+if not is_allowed_host(target_url):
43+  raise WebhookRejected(target_url)
44+resp = requests.get(target_url, timeout=5)
Every engagement hands over:
  • Executive briefing
  • Validated findings list
  • Merge-ready PRs
  • Compliance mapping (NIS2 · DORA · CRA)