- Apply software engineering, automation, and DevOps principles to improve how services are built, tested, deployed, observed, operated, and recovered
- Use data, evidence, experimentation, and rigorous engineering analysis to identify reliability risks, test assumptions, and guide technical decisions
- Define, implement, or improve SLIs, SLOs, error budgets, and service-health measures
- Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation
- Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience and operational readiness
- Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices
- Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle
- Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning
- Provide enterprise-level technical leadership for critical production incidents and influence engineering practices that improve incident response, escalation, service restoration, and sustainable on-call operations
- Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices
- Establish enterprise technical direction, develop senior technical leaders, and multiply the capability of the broader engineering organization
Requirements
- Typically 15+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture, or a comparable technical discipline
- Experience with software development or scripting using one or more modern programming languages
- Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level
- Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures
- Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role
- Preferred hands-on experience with AWS or comparable experience with Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI)
- Experience developing, deploying, operating, or improving highly available production software or distributed systems
- Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation
- Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness
- Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience
- Must independently possess eligibility to work in the United States at the date of hire
- Position is ineligible for employment Visa sponsorship
Core Competencies
Demonstrates extensive experience in Software Engineering, Site Reliability Engineering, and DevOps principles, with a strong focus on automation, observability, and incident management. Proficient in leveraging cloud technologies, particularly AWS, to enhance system reliability and operational readiness.
Highest-signal resume keywords
- 15+ Years Experience in Software Engineering
- Expertise in AWS or Comparable Cloud Technologies
- Proficient in CI/CD and Infrastructure as Code
- Experience with SLIs, SLOs, and Incident Management
- Strong Analytical and Problem-Solving Skills
Hard Skills
- Software Development
- Scripting
- Distributed Systems
- Automation
- Observability
- Incident Management
- Performance Analysis
- Capacity Management
- Disaster Recovery
- Resilience Testing
Soft Skills
- Analytical Skills
- Problem-Solving
- Communication
- Collaboration
Industry Keywords
- Site Reliability Engineering
- DevOps
- Infrastructure Engineering
- Cloud/Platform Engineering
- Software Engineering Principles
Tools & Technologies
- AWS
- Microsoft Azure
- Google Cloud Platform
- Oracle Cloud Infrastructure
- CI/CD Tools
- Monitoring Tools
- Logging Tools
- Dashboards
- Containers
- Orchestration