Production SRE Lead: Reliability, Scale & Observability

Vultr

United States

Remote

USD 140,000 - 160,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Company-paid insurance
401(k) match
Professional development reimbursement
Paid holidays + PTO accrual
Remote office stipend
Internet reimbursement
Gym membership reimbursement

Job summary

Vultr is seeking a Manager of Production Site Reliability Engineering to build and lead a new team responsible for the availability, performance, and operability of our production control plane. This is the environment that runs our customer-facing web applications and the internal platform that supports them.

You will hire, grow, and lead a team of SREs and database engineers, set the operational culture, and partner closely with engineering teams across the organization to ensure the

Qualifications

  • 10+ years in SRE or infra with at least 2 years in leadership.
  • Experience operating production web stacks at scale and debugging performance.
  • Strong Linux, networking, systemd and security posture.
  • Proven ability to hire and build engineering teams.

Responsibilities

  • Build and lead the Production Site Reliability Engineering team.
  • Own availability, performance, and operability of the production web stack.
  • Lead incident response and drive postmortems for durable improvements.
  • Own configuration management for the production environment.
  • Partner with engineering to ensure new services are production-ready.
  • Drive the control plane re-architecture and cutover planning.
  • Set the operational roadmap including capacity planning and DR.
  • Establish monitoring, alerting, and observability across the stack.
  • Manage the production database and caching environment.
  • Foster a culture of operational excellence with runbooks and reviews.

Skills

Site Reliability Engineering
Team Leadership
Linux Systems
Observability
Incident Response
Database Replication
Migrations / Cutovers
PHP Basics
Cloudflare / CDN

Tools

Puppet
Ansible
Chef
HAProxy
Keepalived
Redis

Job description

Vultr is seeking a Manager of Production Site Reliability Engineering to build and lead a new team responsible for the availability, performance, and operability of our production control plane. This is the environment that runs our customer-facing web applications and the internal platform that supports them.

You will hire, grow, and lead a team of SREs and database engineers, set the operational culture, and partner closely with engineering teams across the organization to ensure the

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Production Site Reliability Engineering Lead
Production Site Reliability Engineering Lead

WebHosting • Northern (KY)

Hybrid
USD 140,000 - 160,000
401(k) match up to 4%
Paid holidays & PTO
Professional development reimbursement
+1
Senior SRE: Production Reliability & Observability
Senior SRE: Production Reliability & Observability

Stradit LLC • Dallas (TX), Northern (KY)

Hybrid
USD 140,000 - 190,000
Staff SRE: Scale, Observability & Automation Leader
Staff SRE: Scale, Observability & Automation Leader

Replit • Northern (KY)

Hybrid
USD 180,000 - 260,000
Salary & equity
401(k) matching
Health, dental, vision, life
+9
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Staff SRE: Observability & Reliability Platform Lead
Staff SRE: Observability & Reliability Platform Lead

CVS Health • Scottsdale (AZ)

On-site
USD 118,000 - 261,000
Production Engineering Leader: SRE & Reliability
Production Engineering Leader: SRE & Reliability

CoreWeave • Bellevue (WA)

On-site
USD 207,000 - 275,000
Medical insurance
Dental insurance
Vision insurance
+7
Staff SRE: Scale, Observability & Kubernetes
Staff SRE: Scale, Observability & Kubernetes

Replit • Foster City (CA)

On-site
USD 180,000 - 260,000
Competitive Salary & Equity
401(k) with 4% match
Health, Dental, Vision and Life Ins.
+2
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Staff SRE - Remote-Optional, Scale & Reliability Leader
Staff SRE - Remote-Optional, Scale & Reliability Leader

Pivotal Health • New York (NY)

Hybrid
USD 180,000 - 240,000
Competitive compensation with equity
Full health, dental, vision coverage
401(k) retirement savings plan
+2