A tech consulting firm in Austin is seeking a Lead Site Reliability Engineer specializing in AI/ML Platforms. The ideal candidate will have over 10 years of DevOps/SRE experience, including at least 3 years in a leadership role. Responsibilities include ensuring platform reliability and supporting production ML workloads. Strong skills in Python, automation, and incident management are essential for success in this role.
Qualifications
10+ years of DevOps/SRE experience required.
3+ years of experience in a lead role.
Experience supporting AI/ML platforms or production ML workloads.
Responsibilities
Focus on platform reliability in a large enterprise environment.
Develop automation for monitoring and incident management.
Support production ML workloads.
Skills
Python
DevOps
Site Reliability Engineering
Automation
Incident Management
Job description
A tech consulting firm in Austin is seeking a Lead Site Reliability Engineer specializing in AI/ML Platforms. The ideal candidate will have over 10 years of DevOps/SRE experience, including at least 3 years in a leadership role. Responsibilities include ensuring platform reliability and supporting production ML workloads. Strong skills in Python, automation, and incident management are essential for success in this role.