A complete application in a minute — tailored resume and cover letter, ready to send.
engineeringjobs.net, Inc. seeks a senior leader to drive production operations and SRE for mission-critical cloud services in a regulated environment. You will manage incident response, observability, and remediation while guiding secure AI-assisted development across teams.
Responsibilities include shaping AWS architecture, optimizing costs, ensuring resiliency, and implementing automation. Strong leadership during major incidents is essential.
Lead production operations and SRE practices for business-critical cloud services, including incident response, observability, reliability objectives, capacity planning, and durable remediation. Shape AWS architecture, cost optimization, resiliency, automation, and responsible AI-assisted engineering practices across teams.
Requires formal training or certification in software engineering concepts and at least five years of applied experience, along with deep AWS expertise and proven experience managing production services. Candidates should demonstrate strong SRE and observability skills, leadership in major incidents and remediation, and the ability to guide secure, compliant AI-assisted development in a regulated environment.
AWS Cloud Architecture, Site Reliability Engineering, Live Site Management, Production Operations, Incident Response, Root-Cause Analysis, Observability, SLI/SLO Definition, Distributed Tracing, Capacity Planning, Performance Engineering, Infrastructure as Code, Cloud Cost Optimization, Disaster Recovery, AI-Assisted Software Development, Executive Communication
Health Care Coverage, On-Site Health and Wellness Centers, Retirement Savings Plan, Backup Childcare, Tuition Reimbursement, Mental Health Support, Financial Coaching