Lila Sciences in San Francisco is looking for a Principal ML Engineer to design and scale the machine learning infrastructure for scientific applications. This role involves developing large-scale training pipelines, managing end-to-end ML systems, and collaborating across teams to bridge the gap between machine learning research and production. Candidates should have a Master's in Computer Science or a related field, 10+ years in production ML, and deep expertise in distributed training systems. A competitive compensation package is offered.
Qualifications
10+ years of hands-on experience building and operating production ML systems at scale.
Deep expertise in distributed training infrastructure with large-scale GPU clusters.
Strong fundamentals in system design, production-grade code, and CI/CD.
Responsibilities
Design, build, and optimize large-scale training pipelines for generative models.
Own production ML systems from deployment to monitoring.
Collaborate with AI scientists to bridge research and deployment.
Skills
Production ML systems
Distributed training infrastructure
ML frameworks (PyTorch, JAX, TensorFlow)
Software engineering fundamentals
Cross-functional collaboration
Education
Master's degree in Computer Science or related field
Job description
Lila Sciences in San Francisco is looking for a Principal ML Engineer to design and scale the machine learning infrastructure for scientific applications. This role involves developing large-scale training pipelines, managing end-to-end ML systems, and collaborating across teams to bridge the gap between machine learning research and production. Candidates should have a Master's in Computer Science or a related field, 10+ years in production ML, and deep expertise in distributed training systems. A competitive compensation package is offered.