Senior Site Reliability Engineer — Cloud, Resilience & Automation
Compunnel Inc.
Denton (TX)
Hybrid
USD 120,000 - 150,000
Full time
14 days+
Application generator
Turn this role into an interview — a resume and cover letter built around what this employer wants.
Get past ATS filters
Job summary
A leading technology firm is seeking a Site Reliability Engineer to support cloud-based infrastructure and applications. The ideal candidate will manage production environments, ensuring system availability and resilience through advanced engineering practices. Responsibilities include defining observability and reliability protocols, with hands-on experience in AWS, Azure, and Kubernetes being crucial. This position involves both technical leadership and a proactive approach to incident management, in a collaborative team setting.
Qualifications
5-8+ years of hands-on experience deploying and supporting distributed systems.
Exposure to OS-level scripting languages (Korn/Bash/JavaScript).
Experience with on-call management and incident response.
Responsibilities
Provide Cloud and Platform Engineering support for production environments.
Lead production support—availability and resiliency of critical applications.
Define practices in resiliency engineering, automation, and chaos testing.
Skills
Experience with public cloud environments (AWS and Azure)
Experience with container orchestration (Kubernetes)
Hands-on experience with observability tools (Prometheus, Grafana, Datadog)
Strong communication skills
Experience with infrastructure as code tools (Terraform, Chef)
Education
Bachelor's degree in Engineering or Computer Science
Tools
AWS
Azure
Kubernetes
Terraform
Datadog
Job description
A leading technology firm is seeking a Site Reliability Engineer to support cloud-based infrastructure and applications. The ideal candidate will manage production environments, ensuring system availability and resilience through advanced engineering practices. Responsibilities include defining observability and reliability protocols, with hands-on experience in AWS, Azure, and Kubernetes being crucial. This position involves both technical leadership and a proactive approach to incident management, in a collaborative team setting.