Staff Production Engineer: Operational Excellence & Reliability
Crusoe
San Francisco (CA)
On-site
USD 209,000 - 253,000
Full time
14 days+
Application generator
Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Get past ATS filters
Benefits offered by this job
Health insurance package options
401(k) with 100% match up to 4%
Generous paid time off
Job summary
A technology company in San Francisco is seeking a Staff Production Engineer focused on Operational Excellence. This role is vital for ensuring the reliability and performance of the company’s AI-optimized cloud platform. The ideal candidate will have 8+ years of experience in Production Engineering, familiarity with GPU workloads, and skills in incident management and automation. The role includes significant collaboration with various engineering teams to drive continuous improvements and support demanding AI workloads. Competitive compensation and benefits are offered.
Qualifications
8+ years of experience in Production Engineering or SRE.
Experience with GPU workloads and HPC environments.
Strong understanding of cloud infrastructure including Kubernetes.
Responsibilities
Lead efforts to improve availability metrics for the cloud platform.
Drive production incident response and resolution.
Design automation to reduce operational toil.
Skills
Production Engineering
Incident Management
Scripting
Collaboration
Linux/Unix Systems
Education
Bachelor's degree in Computer Science or Engineering
Tools
Prometheus
Grafana
Terraform
Ansible
Job description
A technology company in San Francisco is seeking a Staff Production Engineer focused on Operational Excellence. This role is vital for ensuring the reliability and performance of the company’s AI-optimized cloud platform. The ideal candidate will have 8+ years of experience in Production Engineering, familiarity with GPU workloads, and skills in incident management and automation. The role includes significant collaboration with various engineering teams to drive continuous improvements and support demanding AI workloads. Competitive compensation and benefits are offered.