changelog / post
DeepSeek V4.1 Flash Fast available on Baseten
DeepSeek V4.1 Flash Fast is now available through Baseten Model APIs. It serves the same 552B-parameter multimodal model as DeepSeek V4.1 Flash at full precision, on dedicated Fast capacity for higher sustained throughput, and is purpose-built for long-running agentic coding sessions. It has a 1M-token context window, vision, tool calling, and reasoning on by default.
Send requests to deepseek-ai/DeepSeek-V4.1-Flash-Fast through our OpenAI-compatible endpoint with your Baseten API key:
curl https://inference.baseten.co/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $BASETEN_API_KEY" \
-d '{
"model": "deepseek-ai/DeepSeek-V4.1-Flash-Fast",
"messages": [
{"role": "user", "content": "Read the failing test, find the bug, and send me the patch."}
],
"reasoning_effort": "high"
}'For more information, see our docs.