Engineering Manager, SRE needed to lead a Site Reliability Engineering team for a globally distributed technology platform. Requires hands-on technical depth in Kubernetes, AWS, PostgreSQL, CI/CD, observability, and infrastructure as code, plus proven people leadership experience. Fully remote, asynchronous environment with a salary range of USD $75,450–$169,700.
Technical (Must-have)
- Kubernetes
- AWS
- PostgreSQL
- CI/CD
- Observability
- Infrastructure as Code
- Terraform
- GitLab CI
- GitHub Actions
- Jenkins
- Docker
- Shell scripting
- DNS
- TLS
- AI infrastructure
Soft Skills
- Prioritization
- Written communication
- Documentation
- Relationship-building
- Collaboration
- Conflict resolution
- Stakeholder management
- Coaching
- Judgment
- Accountability
- Adaptability
- Curiosity
Technical (Nice-to-have)
- Elixir
- Java
- Clojure
- Node.js
- Python
- OpenTelemetry
- Distributed tracing
- Honeycomb
- Aurora
- Linux systems administration
- Security
- FinOps
- Cloud cost management
Key Responsibilities
- Lead and develop a Site Reliability Engineering team, owning the full career lifecycle of direct reports including onboarding, feedback, performance management, progression, coaching, and hiring.
- Establish a clear team direction and priorities aligned with broader company goals, balancing operational commitments with project delivery and protecting the team’s focus.
- Serve as the team's spokesperson across engineering and with senior leadership, communicating priorities, progress, risks, and technical challenges clearly.
- Own SRE delivery goals, deciding what the team commits to, how work is prioritized, and how operational responsibilities are managed.
- Design and maintain effective support rotations and on-call processes while strengthening incident response practices.
- Provide technical leadership across Kubernetes, AWS, PostgreSQL, DNS and TLS, CI infrastructure, and the broader infrastructure platform.
- Guide the development of reliability practices including SLOs, error budgets, observability, incident response, and post-incident improvements.
- Partner closely with Security on infrastructure threats, patching, controls, audits, and compliance obligations.
- Manage relationships with infrastructure and platform vendors, including renewals and commercial discussions with support from senior leadership.
- Remain hands-on enough to review technical work, challenge architectural decisions, participate credibly in incidents, and identify emerging reliability issues before they escalated.
- Build strong relationships across engineering and encourage teams to bring operational and reliability challenges forward early.
- Continuously improve team health, collaboration, conflict resolution, and retrospective practices.
Kubernetes, AWS, PostgreSQL, CI/CD, Observability, Infrastructure as Code, Terraform, GitLab CI, GitHub Actions, Jenkins, Docker, Shell scripting, DNS, TLS, AI infrastructure