A complete application in a minute — tailored resume and cover letter, ready to send.
Modal Labs is building a platform that covers the entire life of an LLM—from training to deployment and observation. We seek a research‑leaning engineer with a systems background to own inference bets, optimize autoscaling, and ship ideas back into production.
Work in person in NYC or San Francisco, collaborate with customers and labs, and help turn frontier serving techniques into usable primitives and products with measurable impact on latency and quality.
Most of the value of owning a model shows up at serving time. We're building a platform that covers the whole life of an LLM – train it, deploy it, observe it – and inference is where teams feel the difference every day. We already run elastic inference, sandboxes, distributed volumes, and multi‑node training, and we control the infrastructure underneath, so the serving stack is ours to shape rather than something we resell.