Loading memory…
Loading memory…
THE ROLE Every simulation we run touches multiple models (LLMs, speech-to-text, text-to-speech), and our Fortune 500 customers need hundreds, sometimes thousands of these running concurrently. Making that fast, reliable, and cost-efficient is the job. We've built the skeleton. Our team has done this before: running single-digit-percent-of-Google-scale compute at Waymo for massive workloads. The auto-scaling foundations, the queuing systems, the monitoring patterns are in place. But we're at an inflection point — demand is growing fast and there's a ton of low-hanging fruit: optimizing how many workloads run on a single machine, tuning scaling algorithms, deciding what to self-host versus what to keep as managed services. You'll own our model infrastructure end to end: • Scaling GPU and compute infrastructure. Architect and operate the auto-scaling systems that handle spikes of hundreds to thousands of concurrent simulations. Optimize how we provision, schedule, and monitor GPU instances. • Making the hosting decisions. We use a mix of closed-source hosted models and open-source self-hosted models today. You'll evaluate the tradeoffs (cost, latency, quality) and make the calls on wh