Get more replies from employers
Send a job-specific resume in minutes.
Unknown Company in San Francisco, CA seeks a lead ML infrastructure engineer to own the distributed training and inference backbone for a foundation model trained from scratch. You will stand up clusters, build data pipelines at petabyte scale, and optimize GPU performance across model scales.
You will work with FSDP/DeepSpeed, NVIDIA GPUs, Linux, Python and C++, in a distributed cloud environment across GCP/AWS/Azure, with relocation supported.
San Francisco, CA · On-site (5 days/week) · Full-timeCompensation: $200K–$400K + competitive early-stage equity
Our client is a Series A AI research lab building large-scale foundation models for scientific and physical-AI domains. Backed by top-tier investors, they are pursuing a deliberately non-consensus technical thesis and are among the best-funded teams in their space. The founding team comes from self-driving, robotics, and scientific research, and they are scaling their research and engineering org significantly this year.
Founded 2024 · Small, fast-growing team · Industry: AI / foundation models / physical AI
You would own the distributed training and inference backbone for a foundation model trained from scratch — standing up clusters, building data and training pipelines at petabyte scale, and squeezing performance out of GPUs at a low level across model scales.
What you'll be doing
Tech stack: Distributed training frameworks (FSDP, DeepSpeed), NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and a major cloud platform (GCP, AWS, or Azure).