Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Get past ATS filters
Benefits offered by this job
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
100% Parental Leave top-up
Personal enrichment benefits
Remote-flexible work
6 weeks of vacation
Job summary
A leading AI technology company in San Francisco is looking for a Site Reliability Engineer to build high-performance, scalable systems for model serving. You will collaborate with teams to deploy optimized NLP models and ensure high availability. Ideal candidates have 5+ years of experience, particularly in Kubernetes and large-scale infrastructure. The role is focused on automating services and maintaining system reliability in a fast-paced environment. The company values diversity and provides competitive benefits including 6 weeks of vacation.
Qualifications
5+ years of engineering experience running production infrastructure at a large scale.
Experience designing distributed systems with Kubernetes, and GPU workloads.
Strong understanding of distributed systems.
Responsibilities
Build self-service systems for deploying and managing services.
Automate environment observability and resilience.
Ensure defined SLOs are met, including on-call rotation.
Skills
Kubernetes
GCP
AWS
Collaboration
Troubleshooting
Golang
C++
Cloud Infrastructure
Job description
A leading AI technology company in San Francisco is looking for a Site Reliability Engineer to build high-performance, scalable systems for model serving. You will collaborate with teams to deploy optimized NLP models and ensure high availability. Ideal candidates have 5+ years of experience, particularly in Kubernetes and large-scale infrastructure. The role is focused on automating services and maintaining system reliability in a fast-paced environment. The company values diversity and provides competitive benefits including 6 weeks of vacation.