
We’re excited to partner with OpenAI as one of the first open-model inference providers in the OpenAI B2B Marketplace, including a native integration within Codex. OpenAI enterprise customers can now use existing OpenAI commitments for open models served by Baseten within Codex or via the Responses API. For enterprises, this means delivering even more intelligence per dollar through your existing OpenAI commitment.
We’ve written about the multi-model future, and this announcement shows how fast that future is becoming the present. Agentic coding is currently the most widely adopted AI use case, and OpenAI’s Codex and GPT models are two of the most popular choices for scaling code-generation workflows. With open models powered by Baseten now available to OpenAI customers, organizations can optimize agentic workflows across open and closed models and route each task to the best-fit model.
Agentic coding is the first multi-model workflow at global scale
Today, companies leading in AI are shifting to use a mixture of models at Pareto frontiers across intelligence, cost, latency, and capacity.
Top engineering teams are actively funneling tasks, through static or adaptive routers, to the model that delivers the best mix of cost, quality, and performance for the job. And with the cycle between new open and closed models down to weeks, no model stays the best tool for every job for long.
To stay ahead, companies need instant access to the latest models with fast, reliable inference that integrates seamlessly into their harnesses, gateways, and other tooling. Baseten was designed for exactly this, providing high-performance infrastructure primitives and developer tools that integrate natively into each organization’s unique estate.
New capabilities to scale multi-model intelligence for code generation
The best intelligence per dollar in production works at scale when models consistently perform as expected. You need performance, reliability, and flexibility to make a multi-model system technically functional. To scale within an enterprise, you need to add a streamlined developer experience, global governance capabilities, and front-line engineering support.
Full-stack performance optimization: Code generation workloads are particularly challenging at the inference layer. Multiple turns, long prompts over large repos, and huge context windows tax every part of the stack, so our performance work spans everything from tooling down to the engine level.
Multi-cloud capacity and reliability. Code generation runs in bursts across whole engineering organizations, and capacity is the constraint that bites first. Baseten runs on more than 90 clusters across 20+ clouds, and deployments run active-active, so losing a provider or a region reroutes traffic instead of stopping work. Relationships with 200+ compute providers make us the first call when new capacity comes online, so supply grows ahead of demand.
Day zero model access: When a new open model tops the coding benchmarks, it’s not acceptable to wait a quarter to evaluate it. We launch major open models on Model APIs the day they ship. Evaluating the newest frontier model the day it releases is as simple as updating a single line of code.
Developer (and agent) experience: Inference alone isn’t enough for coding workflows. Agents need somewhere to run the code they write, which is why Blaxel, now part of Baseten, gives every agent its own sandbox.
Enterprise governance: Code is sensitive by default, and Baseten’s inference runs on US-based infrastructure with zero data retention (ZDR) for all prompts. For even higher compliance requirements, you can pin deployments to specific regions, implement fine-grained AuthN and AuthZ, and view usage across any model, user, or key.
Applied research and embedded engineering support: Forward-deployed engineers tune deployments to your latency targets, and applied researchers run post-training with you. That's how LangChain trains custom models for LangSmith Engine with Baseten Loops.
Our partnership with OpenAI is another step toward providing the best multi-model agentic coding experience. With access to open frontier models on day zero, fast, reliable inference, and guaranteed capacity in our cloud or yours, we’ll keep building what you need to accelerate AI adoption and increase the per-dollar value of intelligence.
To access open models natively in Codex through your OpenAI commitment, get on the list. For everything else, talk to an engineer.