API key management keys
Automate key administration with a new org-scoped key type. A WORKSPACE_MANAGE_API_KEYS key creates team API keys of any permission level, as well as lists and revokes team and personal keys through...
Jul 22, 2026Observability APIs updates
Pull logs, metrics, and audit logs programmatically through the Management API. Endpoints return logs and metrics for any deployment or environment across your workspace, models, or chains.
Jul 21, 2026Workspace GPU usage
Organization admins can now see how many GPUs their workspace is using across every model and deployment, from the new GPU usage tab in Organization settings.
Jul 15, 2026Inkling available on Baseten
You can start sending requests to Inkling today through our Model APIs by calling the OpenAI-compatible endpoint with your Baseten API key. For larger workloads, dedicated deployments are available.
Jul 8, 2026Model API Deprecation (GLM 5.1, GLM 5, Kimi K2.5, Nemotron Super 120B)
Model API Deprecation (GLM 5.1, GLM 5, Kimi K2.5, Nemotron Super 120B)
Jul 7, 2026Personal API key visibility for admins
Organization admins can now view every member's personal API keys on the API keys page, alongside team keys.
Jul 7, 2026Events on Metrics and Logs graphs
Metrics and Logs now overlay platform events on your graphs, so you can line up a latency spike or scaling change with what caused it. Deployments, promotions, autoscaling and instance-type changes,...
Jul 1, 2026Try the new baseten CLI
We're building one CLI for the whole Baseten model workflow: deploy from local, call your models, stream and filter logs, check metrics, and manage deployments, environments, and secrets, all with...
Jun 30, 2026Connect coding agents to Baseten
Connect your coding agent to the Baseten MCP server and install the Baseten skill to manage your workspace from your agent. Your agent can deploy and promote models, tune autoscaling, pull logs, and...
Jun 29, 2026Configure scale-down rate
You can now cap how aggressively the autoscaler removes replicas when traffic drops. Set max_scale_down_rate between 1% and 50% (default 50%) to limit the share of excess replicas removed at each...