A cutting-edge technology company is seeking a Senior Machine Learning Engineer to build and operate systems that power large-scale machine learning training. This role includes designing ML infrastructure, optimizing performance, and enhancing developer experiences. Candidates should have a Bachelor’s degree in Computer Science and expertise in SLURM, Kubernetes, and GPU workloads. The company offers competitive salary, stock options, full medical benefits, flexible PTO, and a mission-driven culture promoting inclusion.
Qualifications
Expertise supporting production ML systems using SLURM and Kubernetes.
Strong understanding of GPU-accelerated workloads and distributed systems concepts.
Solid Linux fundamentals and experience debugging infrastructure-level issues.
Ability to build automation and tooling.
Responsibilities
Design and improve ML infrastructure systems supporting distributed workloads.
Build workload execution and orchestration patterns across GPU environments.
Troubleshoot performance and scalability issues.
Partner with teams to improve developer experience.
Skills
Production ML systems
Performance optimization
Cluster operations
Workload orchestration
Automation tooling
Education
Bachelor of Science in Computer Science or related field
Tools
SLURM
Kubernetes
Python
Go
Job description
A cutting-edge technology company is seeking a Senior Machine Learning Engineer to build and operate systems that power large-scale machine learning training. This role includes designing ML infrastructure, optimizing performance, and enhancing developer experiences. Candidates should have a Bachelor’s degree in Computer Science and expertise in SLURM, Kubernetes, and GPU workloads. The company offers competitive salary, stock options, full medical benefits, flexible PTO, and a mission-driven culture promoting inclusion.