On October 6, 2025, OpenAI announced AgentKit—a coordinated set of tools meant to shrink the distance between a prototype agent and something your team trusts in production. If you have built agents on the Responses API or the open-source Agents SDK, you already know the pain: orchestration sprawl, bespoke chat UIs, connector whack-a-mole, and eval scripts that nobody runs until an incident.
AgentKit does not replace good engineering. It packages the boring parts so you can focus on tools, policies, and user outcomes. This walkthrough explains how a developer might ship a first production agent using Agent Builder, ChatKit, the Connector Registry, and Evals—and when rolling your own with LangChain or raw SDK calls still makes sense.
Architecture Before Canvas
Start with a narrow use case. "Internal IT help desk for SaaS credentials" beats "AI for everything." Define:
- Inputs users provide (ticket text, email, asset ID).
- Tools the agent may call (read-only directory lookup, password reset webhook, knowledge search).
- Hard stops (no production database writes without human approval).
Sketch a sequence diagram on paper: user message → classification → retrieval → tool call → response. Agent Builder will map to this diagram, but clarity upfront prevents drag-and-drop spaghetti.
Document your threat model alongside the happy path. What happens if a user asks the agent to reset someone else's password? What if retrieval returns a document from the wrong tenant? What if a tool endpoint times out mid-workflow? Production agents fail in predictable ways; designing for those failures before you open the canvas saves weeks of rework.
Step 1: Model the Workflow in Agent Builder
Agent Builder is a visual canvas for multi-step workflows with versioning. OpenAI positioned it for collaboration between engineers and subject-matter experts—legal, support leads, sales ops—who should not read Python to understand logic.
Create a new workflow from a blank canvas or template close to your domain (support bots are common starting points). Add nodes for:
- Intent routing — a lightweight classification step that sends password resets differently from FAQ queries.
- Retrieval — attach a knowledge source or MCP tool that searches approved docs.
- Tool execution — wire MCP endpoints or built-in connectors with explicit input schemas.
- Response synthesis — generate user-facing text with citations where possible.
Use preview runs inside Builder to step through transcripts before any UI exists. Invite a support lead to edit phrasing in prompts while you lock tool permissions—a division of labor that previously required pull requests for every wording tweak.
Version everything. Name versions like 0.3-add-okta-readonly. When Evals fail later, you need to roll back workflow graphs, not guess which prompt drifted.
Guardrails nodes belong in the graph early, not as an afterthought. OpenAI's Guardrails SDKs (Python and JavaScript) integrate with Builder to mask PII, flag jailbreak patterns, and block disallowed topics. Treat them as middleware with defaults stricter than you think you need; loosen with data.
For multi-agent patterns, resist the temptation to add agents for every subtask. A single well-instrumented workflow with clear tool boundaries often outperforms a graph of five agents that pass unstructured context to each other. Add specialized agents only when you have measured latency or quality gains that justify the complexity.
Step 2: Embed the Experience with ChatKit
Building chat UI sounds trivial until you need streaming tokens, tool call indicators, retry on disconnect, and brand-themed layouts. ChatKit is OpenAI's embeddable component layer for agentic chat in your app or site.
Integration outline:
- Stand up a backend route that exchanges your user session for an AgentKit/Responses API session token using your server-side API key—never expose platform keys in the browser.
- Install ChatKit per OpenAI docs and configure theme tokens to match your design system.
- Map ChatKit thread IDs to your internal user IDs for audit logs.
- Enable "thinking" or tool status indicators if your UX research shows they reduce abandonment during slow tool calls.
ChatKit handles presentation; your backend enforces authorization. An agent that can reset passwords must confirm the requester owns the account—ChatKit will not do that for you.
HubSpot's support agent was cited at DevDay as a ChatKit production example. Study third-party writeups for latency budgets: support users tolerate three to five seconds if progress is visible; internal agents may demand sub-second retrieval.
Consider accessibility from day one. Streaming text must work with screen readers. Tool status indicators need text alternatives. Error states should offer actionable next steps—not generic apologies. ChatKit provides primitives; your product team owns the experience standards.
Step 3: Govern Connectors with the Registry
The Connector Registry centralizes how data sources and MCP servers attach across ChatGPT Enterprise and API workspaces. It matters when you graduate from a single developer's API key to dozens of teams importing Dropbox, Google Drive, SharePoint, Teams, or custom MCP servers.
Admin checklist:
- Enable the Global Admin Console prerequisite OpenAI documented for Registry access.
- Register only approved MCP endpoints with ownership tags and rotation schedules.
- Separate read connectors used in customer agents from write connectors reserved for internal automation.
- Document which workflows depend on which connector IDs so deprecations do not surprise you.
Developers still implement MCP servers; Registry is the control plane. If you lack Enterprise admin access, simulate governance with environment-separated API keys and manual allowlists until Registry reaches your tenant.
When building custom MCP servers, treat them like microservices. Define explicit input and output schemas. Log every invocation with correlation IDs that match your agent traces. Implement timeouts and circuit breakers so a slow SharePoint query does not hang the entire workflow. The Registry governs who can connect; your server code governs what happens when they do.
Step 4: Prove Readiness with Evals
Shipping without evals is shipping hope. AgentKit's expanded Evals add datasets, trace grading, automated prompt optimization, and third-party model comparisons—features that mirror what mature agent teams built in-house.
Minimal eval loop:
- Dataset — collect fifty to two hundred real or sanitized transcripts labeled with expected outcomes (resolved, escalated, tool invoked).
- Trace grading — run the workflow end-to-end; automated graders score tool selection, citation accuracy, and forbidden actions.
- Regression gate — block deployment when success rate drops more than an agreed threshold versus the previous workflow version.
- Prompt optimization — use human annotations on failures to propose prompt revisions; review in Builder before merge.
Store eval artifacts next to workflow versions. When legal asks why the agent refused a request, you want traces—not vibes.
Expand your dataset over time with production failures. Every escalated ticket, every user complaint, every guardrail trigger is a candidate eval case. Teams that treat production incidents as eval inputs improve faster than teams that only test with synthetic scenarios.
Run adversarial evals deliberately. Red-team prompts should include jailbreak attempts, requests for other users' data, and instructions to bypass guardrails. If your eval suite only contains friendly queries, it will not protect you in production.
Step 5: Production Hardening Developers Should Not Skip
Authentication and tenancy. Map every ChatKit session to a tenant ID. Inject tenant scoping into retrieval tools so one customer never sees another's docs.
Idempotency for tools. Password resets and ticket creates must tolerate retries. Pass idempotency keys through MCP tool payloads.
Observability. Log model version, workflow version, tool latencies, and guardrail triggers. OpenAI provides pieces; you own dashboards.
Fallback paths. When tools fail, the agent should offer escalation to humans with context attached—not loop apologies.
Rate limits and cost caps. AgentKit inherits standard API model pricing. Set per-tenant budgets in your proxy layer.
Secrets management. Rotate API keys on a schedule. Never embed credentials in workflow graphs that non-engineers can export. Use short-lived tokens where possible.
Deployment strategy. Feature-flag new workflow versions. Run shadow traffic against the new graph before cutting over production sessions. Rollback should be a one-click revert to the previous workflow version, not a frantic prompt edit.
When AgentKit vs LangChain vs Raw Agents SDK
Choose AgentKit when ChatKit, visual workflow iteration, and integrated Evals shorten time-to-production and your org accepts OpenAI platform coupling.
Choose LangChain (or similar) when you need deep custom orchestration, multi-vendor model routing, or on-prem constraints AgentKit does not address. LangChain's ecosystem of integrations and community patterns remains valuable for teams that need flexibility over polish.
Choose raw Agents SDK when you want code-first control with minimal UI, already have a design system, and only need selective Kit components—say, Evals alone.
Hybrid patterns are normal: Agents SDK in backend services, ChatKit for UI, custom LangChain retrieval where you already invested. The goal is not purity—it is shipping something maintainable.
LangChain excels when your orchestration logic is complex enough that a visual canvas becomes harder to reason about than code. Long-running agents with custom memory strategies, dynamic tool registration, and multi-model routing often fit better in code. AgentKit excels when your workflow is stable enough to visualize and your bottleneck is UI, governance, or measurement—not orchestration expressiveness.
A Reference Week-by-Week Plan
Week 1: Scope use case, implement MCP read tools, prototype workflow in Agent Builder with preview runs.
Week 2: Embed ChatKit in staging, wire auth, add Guardrails for PII and jailbreak attempts.
Week 3: Build eval dataset from historical tickets, define graders, establish baseline scores.
Week 4: Load test tool endpoints, run red-team prompts, fix failures, promote workflow version 1.0 to production with feature flag.
Week 5+: Monitor traces, expand dataset monthly, tune prompts only through eval-backed changes.
Adjust the timeline for your organization's review cycles. Legal approval for connector access, security review for MCP servers, and procurement for API budget can add weeks. Start those conversations in Week 1, not Week 4.
Closing Perspective
AgentKit does not magically produce trustworthy agents. It removes excuses for skipping versioning, UI polish, connector governance, and measurement—the unglamorous infrastructure that separates demos from products.
DevDay 2025 marked OpenAI's bid to own the full agent lifecycle the same way it owns foundation models. Developers still decide whether that lifecycle fits their risk profile. If it does, start small, eval relentlessly, and treat Agent Builder as source code whose graph edits deserve the same review culture as a pull request.
Your first production agent will not be your best one. With AgentKit's components wired correctly, it can still be the first you are willing to support at 2 a.m.—which is the definition of production ready.
Further Reading
Discover more articles on similar topics across our network
Comments
Loading comments…