Lead SRE: Build Reliable, Low-Latency Platforms (Remote)

T. Rowe Price

Juneau (AK)

Hybrid

USD 140,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Flexible remote work
Health care benefits
Tuition assistance
Wellness programs
Employee stock purchase plan
Competitive retirement plan

Job summary

T. Rowe Price is seeking a Lead Site Reliability Engineer within the CDO Technology Group to ensure availability, latency, performance, and stability of critical infrastructure supporting data platforms, applications, and services.

The role involves proactive monitoring, automated alerts, capacity planning, and collaboration with development teams to optimize resources and deployments in a fast-paced financial environment. Remote-friendly with up to three days of remote work per week.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, or related field.
  • 8+ years as a Site Reliability Engineer or equivalent.
  • Proven experience monitoring, analyzing, and optimizing large-scale distributed systems.
  • Expertise in Linux systems administration, managing servers, OS, network configurations.
  • Strong scripting and automation skills (Bash, Python, etc.).
  • Familiarity with AWS.
  • Experience with DevOps tools and practices (GitLab CI/CD, Docker).
  • Excellent troubleshooting and problem-solving abilities.
  • Strong communication skills for technical and non-technical stakeholders.
  • Passion for maintaining high availability, performance, and reliability in a fast-paced financial environment.

Responsibilities

  • Proactively monitor and identify issues impacting system availability.
  • Implement automated alerts for outages or performance degradation.
  • Collaborate with development teams to design and implement resilience solutions.
  • Analyze performance metrics to identify latency bottlenecks and optimize.
  • Develop and maintain metrics dashboards for KPIs and detect anomalies.
  • Optimize resource utilization and collaborate with teams to allocate resources efficiently.
  • Participate in release planning and develop automated deployment and rollback procedures.
  • Design, implement, and maintain comprehensive monitoring infrastructure and alerts.
  • Respond to incidents, analyze root causes, and document lessons learned, contributing to capacity planning.

Skills

Scripting & automation
Monitoring & performance analysis
Troubleshooting
Communication skills
DevOps practices
AWS familiarity

Education

Bachelor's degree in Computer Science/IT or related field

Tools

AWS
GitLab CI/CD
Docker

Job description

T. Rowe Price is seeking a Lead Site Reliability Engineer within the CDO Technology Group to ensure availability, latency, performance, and stability of critical infrastructure supporting data platforms, applications, and services.

The role involves proactive monitoring, automated alerts, capacity planning, and collaboration with development teams to optimize resources and deployments in a fast-paced financial environment. Remote-friendly with up to three days of remote work per week.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability
Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

T. Rowe Price • Juneau (AK)

Hybrid
USD 140,000 - 190,000
Flexible remote work
Health care benefits
Tuition assistance
+3
Remote Senior Site Reliability Engineer — Reliability Lead
Remote Senior Site Reliability Engineer — Reliability Lead

Priority Technology Holdings, Inc. • Alpharetta (GA)

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
Medical, dental, and vision coverage
+1
Lead SRE: Scalable, High-Availability Systems
Lead SRE: Scalable, High-Availability Systems

Socket.dev • North Carolina

On-site
USD 150,000 - 210,000
Remote Lead Site Reliability Engineer: Performance & Scaling
Remote Lead Site Reliability Engineer: Performance & Scaling

Techholding • United States

Remote
USD 140,000 - 190,000
Lead SRE Engineer
Lead SRE Engineer

SIMARN Solutions • Charlotte (NC)

On-site
USD 120,000 - 160,000
Lead SRE: Observability, Dynatrace & Automation
Lead SRE: Observability, Dynatrace & Automation

CRC Group • Charlotte (NC)

On-site
USD 120,000 - 190,000
Medical, dental, vision insurance
401(k) plan with company match
Paid time off
+1
Principal SRE — Observability & Cloud Reliability
Principal SRE — Observability & Cloud Reliability

T. Rowe Price • Washington

Hybrid
USD 159,000 - 339,000
Competitive compensation
Annual bonus eligibility
Hybrid work schedule
+2
Lead SRE / DevOps Engineer: Observability & Resilience
Lead SRE / DevOps Engineer: Observability & Resilience

Synechron • Dallas (TX)

Hybrid
USD 125,000 - 135,000
Remote SRE Engineering Manager — Lead Reliability
Remote SRE Engineering Manager — Lead Reliability

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000