Company overview
Magicare AI enables nursing facilities to make admissions decisions that previously took 30-60 minutes in under 1 minute. Historically, when a hospital discharges a patient, nursing facilities take 30-60 minutes to review a referral packet that can run hundreds of pages. Whoever accepts the patient first wins the business, but missed details can lead to hospital readmission, creating a poor patient experience and generating additional costs for the patient and the hospital.
Magicare AI parses 250 individual data points from each patient referral packet and presents them in easy-to-consume admission recommendations, enabling facilities to make accurate decisions up to 100x faster. Magicare AI supports hundreds of post-acute care organizations and is an essential part of their workflow.
Impact highlights
33% reduction in cost
60M TPM of traffic handled
2 days saved on average for every rate limit request
Real time document parsing
Challenge
When Magicare’s customers receive a patient referral packet, Magicare’s agentic pipeline scans the entire packet. Agents parse relevant patient data from the referral packet, cross references the facilities capabilities and generate an admission proposal based on whether the patient is a fit for the facility. This agentic pipeline produces hundreds to thousands of parallel requests in order to parse the long context window and cross-reference unique facility capabilities.
Magicare’s agentic workload requires low end-to-end latency, high uptime, and the ability to handle sporadic bursts of requests. These requirements enable nursing facilities to win business and accept patients quickly and with high conviction. Magicare AI previously ran inference with a specialized hardware provider. Given the bursty nature of Magicare’s workload, their tokens per minute consistently exceeded 30M. This level of load was difficult to serve reliably with their previous provider: rate limit requests took days, TPM unlocks required upfront payment, and reliability was inconsistent.
"What I really care about is latency, reliability, and knowing that my account will not arbitrarily get blocked if we go over some rate limit. I need confidence we can scale without intervention."
Solution
Magicare AI utilized Baseten Model APIs to immediately unblock customer traffic. Magicare started with 10M TPM and 7k RPM and utilized Baseten’s custom Eagle3 speculative decoding to improve decode throughput by 2x. As the workload scaled, it became clear Magicare was best served by a hybrid dedicated deployment and Model API solution to ensure their base traffic could be served reliably with stable performance, while also ensuring traffic bursts wouldn’t be rate-limited by falling back to Baseten’s model APIs.
By using both Dedicated and Model APIs for the reasoning portion of their agentic pipeline, Magicare can tune performance to the specific needs of their workload without losing the ability to serve spiky traffic. To get a sense of the burst scale, if just 10% of traffic sent to Baseten failed, Magicare would have to serve that 10% across 13 different providers. Baseten’s reliability ensures that this doesn’t need to happen. Dedicated deployments use multi-cloud capacity management to run workloads across 20 different clouds, ensuring they can fail over and scale across clouds.
"Baseten stayed up when 3 different clouds went down, 2 of them being some of the largest providers of GPUs. That tells me everything I need to know about who to build on."
Result
Magicare AI now serves its admissions workload across a dedicated deployment and Baseten's Model APIs, averaging around 15 million tokens per minute and spiking to 30-40 million during business hours, with peaks exceeding 60 million. Magicare’s end users shouldn’t have to think about the implications of request load, and Baseten’s flexible infrastructure ensures performance is stable even as requests spike. Moving from prepaid rate limits to usage-based pricing reduced costs while eliminating the days-long delays that stalled production.
"We tested a lot of vendors. A lot of people say they are performant and multi-cloud, but when we asked for the TPM we needed, we were told we could only unlock it with massive upfront compute reservations. With Baseten, we serve the majority of our traffic on one provider, pay as we go, and have a team who we can co-engineer with as we scale and add new models and workloads."