Lead Site Reliability Engineer (SRE)

T. Rowe Price

Juneau (AK)

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible remote work
Health care benefits
Tuition assistance
Wellness programs
Employee stock purchase plan
Competitive retirement plan

Job summary

T. Rowe Price is seeking a Lead Site Reliability Engineer within the CDO Technology Group to ensure availability, latency, performance, and stability of critical infrastructure supporting data platforms, applications, and services.

The role involves proactive monitoring, automated alerts, capacity planning, and collaboration with development teams to optimize resources and deployments in a fast-paced financial environment. Remote-friendly with up to three days of remote work per week.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, or related field.
  • 8+ years as a Site Reliability Engineer or equivalent.
  • Proven experience monitoring, analyzing, and optimizing large-scale distributed systems.
  • Expertise in Linux systems administration, managing servers, OS, network configurations.
  • Strong scripting and automation skills (Bash, Python, etc.).
  • Familiarity with AWS.
  • Experience with DevOps tools and practices (GitLab CI/CD, Docker).
  • Excellent troubleshooting and problem-solving abilities.
  • Strong communication skills for technical and non-technical stakeholders.
  • Passion for maintaining high availability, performance, and reliability in a fast-paced financial environment.

Responsibilities

  • Proactively monitor and identify issues impacting system availability.
  • Implement automated alerts for outages or performance degradation.
  • Collaborate with development teams to design and implement resilience solutions.
  • Analyze performance metrics to identify latency bottlenecks and optimize.
  • Develop and maintain metrics dashboards for KPIs and detect anomalies.
  • Optimize resource utilization and collaborate with teams to allocate resources efficiently.
  • Participate in release planning and develop automated deployment and rollback procedures.
  • Design, implement, and maintain comprehensive monitoring infrastructure and alerts.
  • Respond to incidents, analyze root causes, and document lessons learned, contributing to capacity planning.

Skills

Scripting & automation
Monitoring & performance analysis
Troubleshooting
Communication skills
DevOps practices
AWS familiarity

Education

Bachelor's degree in Computer Science/IT or related field

Tools

AWS
GitLab CI/CD
Docker

Job description

Summary

Lead Site Reliability Engineer (SRE) in the CDO Technology Group ensuring availability, latency, performance, and stability of critical infrastructure supporting data platforms, applications, and services.

Responsibilities
  • Availability
    • Proactively monitor and identify potential issues impacting system availability.
    • Implement automated alerts for outages or performance degradation.
    • Collaborate with development teams to design and implement solutions that enhance resilience and reduce downtime.
  • Latency
    • Analyze performance metrics to identify and resolve latency bottlenecks.
    • Implement performance optimization techniques and tools.
    • Ensure new features and code changes do not introduce performance regressions.
  • Performance
    • Develop and maintain metrics dashboards for KPIs.
    • Identify trends and anomalies indicating potential issues.
    • Recommend and implement optimization strategies.
  • Efficiency
    • Optimize resource utilization and minimize unnecessary expenditure.
    • Collaborate with development teams to optimize resource allocation for new applications.
  • Release Management
    • Participate in release planning to ensure smooth deployments.
    • Develop automated deployment and rollback procedures.
    • Monitor new releases and address issues promptly.
  • Monitoring
    • Design, implement, and maintain comprehensive monitoring infrastructure.
    • Analyze data to identify potential issues and troubleshoot proactively.
    • Develop alerts and notifications for critical events.
  • Emergency Response
    • Respond promptly to incidents and collaborate on resolution.
    • Analyze root causes and implement preventive measures.
    • Document incident responses and lessons learned.
    • Participate in capacity planning to anticipate future workloads.
    • Stay abreast of emerging technologies, trends, and best practices.
    • Review architecture design for high availability and disaster recovery.
    • Collaborate with reliability and infrastructure teams on observability, tracing, and alerting tooling.
Qualifications
  • Bachelor's degree in Computer Science, Information Technology, or related field.
  • 8+ years as a Site Reliability Engineer or equivalent.
  • Proven experience monitoring, analyzing, and optimizing large‑scale distributed systems.
  • Expertise in Linux systems administration, managing servers, OS, network configurations.
  • Strong scripting and automation skills (Bash, Python, etc.).
  • Familiarity with AWS.
  • Experience with DevOps tools and practices (GitLab CI/CD, Docker).
  • Excellent troubleshooting and problem‑solving abilities.
  • Strong communication skills for technical and non‑technical stakeholders.
  • Passion for maintaining high availability, performance, and reliability in a fast‑paced financial environment.
Benefits
  • Competitive salary and comprehensive benefits package.
  • Opportunity to work with cutting‑edge technologies and innovate solutions.
  • Collaborative and supportive work environment with continuous learning.
  • Competitive pay and bonuses, generous retirement plan, employee stock purchase plan with matching contributions.
  • Flexible and remote work opportunities.
  • Health care benefits (medical, dental, vision).
  • Tuition assistance.
  • Wellness programs (fitness reimbursement, Employee Assistance Program).
FINRA Requirements

FINRA licenses are not required and will not be supported for this role.

Work Flexibility

Eligible for remote work up to three days a week.

Commitment to Diversity, Equity, and Inclusion

We strive for equity, equality, and opportunity for all associates, fostering an environment where people can bring their authentic selves to work and create belonging.

Equal Opportunity Employer

T. Rowe Price is an equal‑opportunity employer and values diversity of thought, gender, and race. We prohibit discrimination on the basis of race, religion, creed, color, national origin, sex, gender, age, disability, marital status, sexual orientation, gender identity or expression, citizenship status, military or veteran status, pregnancy, or any other classification protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

T. Rowe Price • Washington

Hybrid
USD 175,000 - 299,000
Hybrid work schedule
Competitive compensation
Annual bonus eligibility
+4
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

T. Rowe Price • Owings Mills (MD)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Health and wellness benefits
Family care resources
Sr. Software Engineer
Sr. Software Engineer

Relha LLC • Owings Mills (MD), Northern (KY)

Hybrid
USD 121,000 - 206,000
Competitive compensation
Health and wellness benefits, incl. 2n
Paid time off for vacation, illness, &
+1
Senior Software Engineer, Developer Services Engineering
Senior Software Engineer, Developer Services Engineering

T. Rowe Price • Owings Mills (MD)

Hybrid
USD 121,000 - 206,000
Competitive compensation
Annual bonus eligibility
Hybrid work schedule
+1
Principal Cloud Reliability Engineer
Principal Cloud Reliability Engineer

T. Rowe Price • Owings Mills (MD)

Hybrid
USD 159,000 - 272,000
Competitive compensation
Annual bonus eligibility
Generous retirement plan
+2
Sr .Net Software Engineer (Hybrid- Owings Mills, MD)
Sr .Net Software Engineer (Hybrid- Owings Mills, MD)

Relha LLC • Owings Mills (MD), Northern (KY)

Hybrid
USD 121,000 - 206,000
Hybrid work model
Sr .Net Software Engineer (Hybrid- Owings Mills, MD)
Sr .Net Software Engineer (Hybrid- Owings Mills, MD)

T. Rowe Price • Owings Mills (MD)

Hybrid
USD 121,000 - 206,000
Competitive compensation
Annual bonus eligibility
Generous retirement plan
+5
Senior Data Security Engineer (Remote Opportunity)
Senior Data Security Engineer (Remote Opportunity)

Relha LLC • Owings Mills (MD), Northern (KY)

Hybrid
USD 97,000 - 164,000
Sr. Technical Business Analyst
Sr. Technical Business Analyst

T. Rowe Price • Seattle (WA)

Hybrid
USD 120,000 - 160,000
Remote work opportunities
Tuition assistance
Health care benefits
Sr. SRE - Site Reliability Engineer
Sr. SRE - Site Reliability Engineer

Charles Schwab • Austin (TX)

On-site
USD 140,000 - 190,000