Try the new DeepSeek V4 Pro 0813 today. Frontier intelligence at a fraction of the cost. Here
Product

Model APIs made for products, not toys

On-demand frontier models running on the Baseten Inference Stack that won’t ruin launch day.

DJ Zappegos logo

With Baseten, we now support open-source models like DeepSeek and Llama in Retool, giving users more flexibility for what they can build. Our customers are creating AI apps and workflows, and Baseten's Model APIs deliver the enterprise-grade performance and reliability they need to ship to production.

DJ Zappegos
Engineering Manager, Retool
benefits

Build your product with pre-optimized frontier models

Baseten Model APIs are built for production first, with the performance and reliability that only the Baseten Inference Stack can enable.

Ship faster

Use our Model APIs as drop-in replacements for closed models with comprehensive observability, logging, and budgeting built in.

Scale further

Run leading open-source models on our optimized infra with the fastest runtime available, all on the latest-generation GPUs. 

Spend less

Spend 5-10x less than closed alternatives with our optimized multi-cloud infrastructure and efficient frontier open models.

Features

Fast inference that scales with you

 Try out new models, integrate them into your product, and launch to the top of Hacker News and Product Hunt—all in a single day.

OpenAI compatible

Migrate from closed models to open-source by swapping a URL. We’re fully OpenAI compatible with support for function calling and more.

Pre-optimized performance

We ship leading models optimized from the bottom up with the Baseten Inference Stack, so every Model API is ultra-fast out of the box.

Seamless scaling

Go from Model API to dedicated deployments on the hardware of your choosing in two clicks from the Baseten UI.

Four nines of uptime

We achieve reliability that only active-active redundancy can provide with our cloud-agnostic, multi-cluster autoscaling.

Secure and compliant

We take extensive security measures, never store inference inputs or outputs, and are SOC 2 Type II certified and HIPAA compliant.

Featureful inference

Structured outputs and tool use are baked into our Model APIs as part of the Baseten Inference Stack.

Instant access to leading models

Model library
DeepSeek V4 Pro 0813
Model API

DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813 is the latest 1.6T-parameter open frontier model from DeepSeek AI

Kimi K3
Model API

Kimi K3

The open frontier model. 2.8T parameters, 1M-token context, top benchmarks scores.

Kimi K2.6
Model API

Kimi K2.6

Kimi K2.6 builds on Kimi K2.5 with increased agentic capabilities

GLM-5.2 Fast
Model API

GLM-5.2 Fast

A speed-optimized GLM-5.2 Model API built for real-time workloads

GLM-5.2
Model API

GLM-5.2

GLM-5.2 is Z.AI's next-gen model for agentic engineering, with stronger coding and agentic capabilities and sustained execution on long-horizon tasks.

DeepSeek-V4-Flash-0731
Model API

DeepSeek-V4-Flash-0731

DeepSeek V4-Flash-0731 is an open-weight 284B MoE (13B active) with 1M context and selectable reasoning effort, tuned for coding, chat, and agent workflows.

Inkling
Model API

Inkling

Inkling is a multimodal MoE model (975B total, 41B active, 256k context) reasoning over text, image, and audio inputs.

Inkling-Small
Model API

Inkling-Small

An efficient, open-weight multimodal AI model. Enjoy faster inference, lower costs, and native text, image, and audio support at 1M tokens.

NVIDIA Nemotron 3 Ultra
Model API

NVIDIA Nemotron 3 Ultra

550B hybrid Mamba-Transformer MoE with 55B active params, latent MoE routing, multi-token prediction, and up to 1M token context

Pricing

Price per

1M tokens

Model

Input

Cache Input

Output

$3.00
$0.30
$15.00Try
$0.95
$0.16
$4.00Try
$2.10
$0.21
$6.60Try
$1.40
$0.14
$4.40Try
$1.74
$0.145
$3.48Try
$1.00
$0.17
$4.05Try
$0.50
$0.10
$1.20Try
$0.10
-
$0.50Try

Built for every stage in your inference journey

Explore resources
Dedicated

Get dedicated resources

Launch dedicated deployments as your scale grows. We’ll work with you to choose the best hardware for your use case.

Get started

Launch dedicated deployments as your scale grows. We’ll work with you to choose the best hardware for your use case.

Get started
Training

Fine-tune for any use case

Tailor any model on custom data with featureful training infra built for multi-node jobs, model caching, checkpointing, and more.

Learn more

Tailor any model on custom data with featureful training infra built for multi-node jobs, model caching, checkpointing, and more.

Learn more
Guide

Get the Baseten Inference Stack

Learn how we optimized inference infra and model performance from the ground up to build the fastest stack on the market.

Read more

Learn how we optimized inference infra and model performance from the ground up to build the fastest stack on the market.

Read more
Lily Clifford logo

Rime's state-of-the-art p99 latency and 100% uptime is driven by our shared laser focus on fundamentals, and we're excited to push the frontier even further with Baseten.

Lily Clifford
Co-founder and CEO, Rime