Senior SRE Lead: Scalable, Reliable Systems

Navy Federal Credit Union

Vienna (VA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Navy Federal Credit Union is seeking an experienced Site Reliability Engineer to maintain and enhance reliability, availability, and performance of systems. You will design fault-tolerant architectures, automate tasks, and collaborate with development teams to ensure stable operations.

Ideal candidates have 7–10 years in SRE, strong programming skills (Python/Java/Go), and expertise in monitoring and incident response.

Qualifications

  • Master's degree or equivalent in CS, engineering or related field.
  • 7–10 years of site reliability engineering experience.
  • Subject matter expert in the area with cross-disciplinary understanding.
  • Advanced knowledge of system monitoring, incident response and automation tools.
  • Significant experience with SRE principles, practices and software development.

Responsibilities

  • Design and implement reliable, scalable, and highly available systems.
  • Analyze and evaluate complex factors to develop optimal reliability solutions.
  • Automate tasks and processes to improve efficiency and reduce human error.
  • Ensure systems operate within performance and reliability metrics.
  • Monitor, maintain and improve system performance, availability and reliability.
  • Troubleshoot complex production issues.
  • Lead collaboration with cross-functional teams to embed reliability best practices.
  • Develop and enhance standard operating procedures and instructions.
  • Build strong working relationships with team members and leadership.
  • Lead medium to large projects and mentor junior staff.

Skills

Python
Java
Go
Incident response
SRE knowledge

Education

Master's degree in computer science or engineering

Tools

Monitoring tools
Automation tools

Job description

Navy Federal Credit Union is seeking an experienced Site Reliability Engineer to maintain and enhance reliability, availability, and performance of systems. You will design fault-tolerant architectures, automate tasks, and collaborate with development teams to ensure stable operations.

Ideal candidates have 7–10 years in SRE, strong programming skills (Python/Java/Go), and expertise in monitoring and incident response.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer — Resilience Leader
Principal Site Reliability Engineer — Resilience Leader

Navy Federal Credit Union • Pensacola (FL)

On-site
USD 110,000 - 140,000
Senior Site Reliability Engineer - Lead, Automate & Scale
Senior Site Reliability Engineer - Lead, Automate & Scale

Navy Federal Credit Union • Winchester (VA)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer – Reliability at Scale
Senior Site Reliability Engineer – Reliability at Scale

Socket.dev • Vienna (VA)

On-site
USD 90,000 - 150,000
Senior SRE Lead: Reliability, Observability & Automation
Senior SRE Lead: Reliability, Observability & Automation

Bank of America • Chandler (AZ)

On-site
USD 180,000 - 240,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Navy Federal Credit Union • Pensacola (FL)

On-site
USD 110,000 - 140,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Navy Federal Credit Union • Winchester (VA)

On-site
USD 120,000 - 180,000
Senior SRE — Flexible, AI-Driven Reliability
Senior SRE — Flexible, AI-Driven Reliability

Salesforce, Inc. • San Francisco (CA)

Hybrid
USD 148,000 - 224,000
Senior SRE — AWS, Automation, 24/7 Uptime, Flexible
Senior SRE — AWS, Automation, 24/7 Uptime, Flexible

Fearless • Baltimore (MD)

On-site
USD 114,000 - 184,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Vice President, SRE Lead (Incident Management), Application Production Services & Engineering
Vice President, SRE Lead (Incident Management), Application Production Services & Engineering

Bank of America • United States

On-site
USD 110,000 - 140,000