We’re looking for a senior, hands-on engineer to build and operate the infrastructure that enables machine learning teams to develop, deploy, and scale production systems. This role sits at the intersection of software engineering, cloud infrastructure, platform engineering, and ML, with a strong emphasis on reliability and operational excellence.
You’ll work closely with ML engineers, data scientists, and software engineers to create scalable infrastructure, streamline deployment processes, and ensure production ML workloads are reliable, observable, and cost-efficient.
Key Responsibilities
- Design, build, and maintain infrastructure supporting machine learning training, model deployment, and inference workloads.
- Develop and improve CI/CD pipelines, infrastructure-as-code, observability, and operational tooling.
- Own production infrastructure with a focus on availability, performance, scalability, and cost optimization.
- Build and operate cloud-native services and Kubernetes-based infrastructure supporting ML and data workloads.
- Collaborate with ML and data teams to productionize models and improve deployment and operational processes.
- Create automation, internal tools, and platform abstractions that improve developer productivity and engineering workflows.
- Monitor system health, troubleshoot production issues, and drive improvements to platform reliability.
- Participate in on-call rotations and take end-to-end ownership of critical infrastructure and platform components.
- Help establish best practices around infrastructure, deployment, monitoring, and production operations.
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
- 7+ years of experience in software engineering, platform engineering, DevOps, SRE, or a similar discipline.
- Strong backend/systems engineering fundamentals with significant experience operating production environments.
- Strong Python experience and Scala microservices.
- Proven experience building and maintaining backend, platform, or infrastructure services.
- Hands‑on experience with cloud infrastructure and modern DevOps tooling, such as AWS, Sagemaker, Docker, Kubernetes, Terraform, and CI/CD platforms.
- Experience designing, deploying, and supporting distributed systems in production.
- Experience with machine learning infrastructure, model serving, data platforms, or ML deployment workflows is highly desirable.
- Strong troubleshooting and problem-solving skills, with a proactive, operations-focused mindset.
- Ability to navigate complex technical challenges and work effectively in a fast-moving engineering environment.
- Strong communication and collaboration skills, with the ability to partner across engineering and data-focused teams.
Top 3 Resume Signals:
Experience building and operating real-time distributed systems at scale; ownership of production ML infrastructure and model-serving environments; and strong cloud-native engineering experience using AWS and Kubernetes.
What You’ll Bring
The ideal candidate is a strong backend/systems engineer who enjoys owning infrastructure from design through production. You’re comfortable diving into complex technical problems, automating repetitive processes, and improving the reliability of systems at scale. Experience supporting ML workloads is a major advantage, but strong platform and distributed-systems expertise is equally important.