An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Jobtailor is seeking a highly capable ML Infrastructure Engineer to build and operate systems for large-scale model training and evaluation in our SF office. You will enhance reliability, throughput, and resource efficiency, develop shared inference platforms, optimize GPU scheduling, and collaborate with researchers and engineers to drive end-to-end performance.
Candidates should bring strong software fundamentals and experience with distributed systems; hybrid work from the US office is
Demonstrates expertise in building and operating large-scale distributed systems, with a focus on ML infrastructure and GPU performance optimization. Capable of diagnosing bottlenecks and improving system reliability and throughput through effective collaboration and ownership of projects.