Site Reliability Engineer II: Cloud, Automation & Observability
Restaurant365
San Francisco (CA)
Hybrid
USD 98,583 - 138,016
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Comprehensive medical benefits, 100% paid for employee
401k + matching
Equity Option Grant
Unlimited PTO + Company holidays
Wellness initiatives
Job summary
Restaurant365 in San Francisco seeks a Site Reliability Engineer II to support and enhance its cloud infrastructure. Candidates should have 2-4 years of experience in site reliability engineering or DevOps, proficiency with cloud platforms like Azure or AWS, and automation experience with tools such as Terraform and Ansible. This role requires a collaborative approach to incident response and system reliability, offering competitive salary and unlimited PTO as part of a comprehensive benefits package.
Qualifications
2-4 years of experience in site reliability engineering, DevOps, or cloud operations.
Experience with cloud platforms (Azure or AWS), including services such as AKS, ECS, Functions/Lambda.
Proficiency with infrastructure‑as‑code and automation tools.
Responsibilities
Respond to production incidents and perform triage and troubleshooting.
Identify and automate manual processes to improve efficiency.
Enhance monitoring tools and platforms for better observability.
Skills
Incident response
System monitoring
Automation
Performance troubleshooting
Linux engineering
Scripting (Python, Bash, PowerShell)
Education
BS in Computer Science, Information Systems, or related field
Tools
Terraform
Ansible
Cloud platforms (Azure, AWS)
Monitoring tools (Prometheus, Grafana, ELK)
CI/CD tools (GitLab, Git)
Job description
Restaurant365 in San Francisco seeks a Site Reliability Engineer II to support and enhance its cloud infrastructure. Candidates should have 2-4 years of experience in site reliability engineering or DevOps, proficiency with cloud platforms like Azure or AWS, and automation experience with tools such as Terraform and Ansible. This role requires a collaborative approach to incident response and system reliability, offering competitive salary and unlimited PTO as part of a comprehensive benefits package.