Inference Engineering

Inference Engineering by Philip Kiely is your guide to becoming an expert in inference.

Inference is the most valuable category in the AI industry.

Inference engineering, on the other hand, is still in its infancy. Inference engineers work across the stack from CUDA to Kubernetes in pursuit of faster, less expensive, more reliable serving of generative AI models in production.

On November 30, 2022 – the day that ChatGPT was launched – there were perhaps a few hundred inference engineers in the world, though they didn’t call themselves that at the time. These specialists mostly worked at frontier labs like OpenAI, Midjourney, and Anthropic or big tech companies like Google and NVIDIA.

Back then, it looked like this might be the way of the AI industry. Perhaps training generative AI models would be so hard, and so expensive, that only a handful of companies would develop closed models and thus require inference engineering for production serving. In this alternate future, the rest of the world would be mere consumers of AI via APIs, renting intelligence a token at a time.

Three years later, it turns out that training generative AI models is hard. It is expensive. But it is neither so hard nor so expensive that it is limited to that handful of players.

Instead, a Cambrian explosion of open models – more than two million and counting on Hugging Face – means that every engineer can now deploy their own intelligence to power their AI products. Research labs around the world, from OpenAI and NVIDIA Nemotron in America to Mistral AI and Black Forest Labs in Europe to Alibaba Qwen, DeepSeek AI, Z AI, and Moonshot AI in China, regularly release open models of all modalities.

Figure P.1: There are well over two million open models on Hugging Face, 25 times more models than five years ago.
Figure P.1: There are well over two million open models on Hugging Face, 25 times more models than five years ago.

Despite closed models getting smarter and cheaper, the movement into open models is accelerating. Open models differ in the availability of their weights:

  • Closed model: A proprietary model where weights are unavailable to the public, like GPT-5 or Claude Sonnet.
  • Open model: A model where weights are publicly available, like Llama or DeepSeek, usually released under the MIT license or a similar permissive license (though some models restrict commercial use, always double-check license terms).

Until December 2024, there was a meaningful gap in intelligence between closed and open models. When DeepSeek V3 and R1 were released, that gap disappeared.

Today, new closed models are matched by open models within months if not weeks, with occasional open models like Kimi K2 Thinking even exceeding closed model capabilities for brief windows.

Even if open models are constantly chasing closed models on benchmarks, they still change the equation for AI product builders. As both types of models get better, closed and open models cross capability thresholds and power new classes of products.

Figure P.2: Open and closed models are both improving rapidly, unlocking and expanding access to new capabilities.
Figure P.2: Open and closed models are both improving rapidly, unlocking and expanding access to new capabilities.

In 2022, it was impossible to build the kinds of AI-native products that define the industry today.

Over time, closed models got smarter and new categories like customer service voice agents and AI-powered IDEs became possible. These early models were slow, expensive, and unreliable, but the capabilities were there and AI engineers began building companies around these capabilities.

As open models crossed the same capability thresholds, these product builders began using them to replace closed models. Many also began fine-tuning open models to cross capability thresholds faster and even exceed closed model quality for their specific product and domain.

Figure P.3: Customizing open models unlocks new capabilities while retaining control over latency, reliability, and economics.
Figure P.3: Customizing open models unlocks new capabilities while retaining control over latency, reliability, and economics.

Switching to open models unlocks the opportunity to use inference engineering to make the models powering AI products better in new dimensions:

  • Latency: Closed model APIs are built for throughput, but open models can be optimized for real-time applications.
  • Availability: While APIs for GPT and Claude are stuck at two nines of uptime, it’s possible to achieve four nines or better with dedicated deployments of open models.
  • Cost: Open models are often at least 80 percent less expensive at scale.

So where three years ago it looked like inference engineering might be a niche field, today every company aiming to build truly differentiated and competitive AI products needs an inference strategy.

AI-native startups like Cursor, Clay, Gamma, and Mercor are redefining hypergrowth building products that rely on open and in-house models. Leading digital native companies like Notion and Superhuman are thriving by deeply integrating AI capabilities into products that hundreds of millions already love.

And a new generation of blended research and engineering teams – World Labs, Writer, Mirage, and dozens of others – are building enormous businesses by training and productizing their own foundation models.

Adoption is even strong in enterprise and regulated industries, which historically have been slow to adapt to new technologies. Companies like OpenEvidence, Abridge, and Ambience are making generative AI ubiquitous in healthcare, while at the world’s largest companies, AI initiatives are moving past the pilot stage into massive user adoption.

I’ve been incredibly fortunate to have a front-row seat to the fastest-moving market in history over the last four years at Baseten, where we power mission-critical inference for the best AI products, including every company listed in the previous paragraphs.

The incredible market-wide demand for inference means that everyone from developers to executives has the opportunity to learn inference engineering and use it to advance their career and business.

You are early. While the potential and impact of inference are becoming clear, the space is young. There are relatively few people working on inference, and newcomers can become experts quickly. There are enormous opportunities to solve novel, interesting, and deeply technical problems at all levels of the stack.

Inference Engineering is your guide to becoming an expert in inference. It contains everything that I’ve learned in four years of working at Baseten. This book is based on interviews with dozens of experts from our engineering team; technical talks I’ve delivered at conferences like NVIDIA GTC, PyTorch Conference, AWS re:Invent, and AI Engineer World’s Fair; and countless conversations with customers and builders around the world.

Thank you for reading Inference Engineering and welcome to the early days of inference.

Philip Kiely
San Francisco, CA