Lead Site Reliability Engineer

Federal Reserve Bank of San Francisco

Richmond (VA)

On-site

USD 147,000 - 234,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Federal Reserve Bank of San Francisco is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and performance of critical systems. You will shape cloud architecture, automation, and security practices while mentoring a growing SRE team onsite in San Francisco and Richmond locations.

As a senior technical leader, you will own incident response, disaster recovery planning, and collaboration with software teams to implement best practices across the stack.

Qualifications

  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles.
  • 3+ years in a lead or senior technical position.
  • Proven track record of managing large-scale production systems.
  • Experience with on-call rotations and incident management.
  • GenAI based Applications: Working knowledge of LLMs and agentic applications a plus.

Responsibilities

  • Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure.
  • Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance.
  • Lead incident response, conduct root cause analysis, and implement preventive measures.
  • Develop and maintain disaster recovery and business continuity plans.
  • Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform).
  • Automate deployment pipelines, monitoring, and operational workflows.
  • Build and maintain internal tools and services to improve operational efficiency.
  • Collaborate with development teams to implement reliability best practices.
  • Conduct code reviews and provide technical guidance on system design.
  • Develop monitoring solutions, alerting systems, and observability frameworks.
  • Integrate security practices into CI/CD pipelines (SAST/DAST).
  • Ensure compliance with industry standards and regulatory requirements.
  • Mentor junior SRE team members and promote SRE culture across the organization.
  • Partner with software engineering teams to improve system reliability and drive technical initiatives.

Skills

Java
Python
Node.js
Microservices
Distributed systems
Cloud
AWS
Terraform
Kubernetes
Docker
CI/CD
GitLab
Security best practices

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

Terraform
GitLab CI/CD
Docker
Kubernetes
AWS

Job description

Company Federal Reserve Bank of San FranciscoWhen you join the Federal Reserve-the nation's central bank-you'll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we're building a dynamic and diverse team for our future.

We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.

Responsibilities
  • System Reliability & Performance
    • Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
    • Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
    • Lead incident response, conduct root cause analysis, and implement preventive measures
    • Develop and maintain disaster recovery and business continuity plans
  • Infrastructure & Automation
    • Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
    • Automate deployment pipelines, monitoring, and operational workflows
    • Optimize cloud resource utilization and cost management
  • Engineering & Development
    • Build and maintain internal tools and services to improve operational efficiency
    • Collaborate with development teams to implement reliability best practices
    • Conduct code reviews and provide technical guidance on system design
    • Develop monitoring solutions, alerting systems, and observability frameworks
  • Security & Compliance
    • Integrate security practices into CI/CD pipelines (SAST/DAST)
    • Implement and maintain security controls across infrastructure and applications
    • Ensure compliance with industry standards and regulatory requirements
    • Conduct security assessments and vulnerability management
  • Leadership & Collaboration
    • Mentor junior SRE team members and promote SRE culture across the organization
    • Partner with software engineering teams to improve system reliabilityDrive technical initiatives and contribute to architectural decisions
    • Document processes, runbooks, and technical specifications
Software Engineering
  • Strong proficiency inJava,Python, andNode.js
  • Experience with microservices architecture and distributed systems
  • Solid understanding of data structures, algorithms, and design patterns
  • Proficiency in writing clean, maintainable, and testable code
Cloud Infrastructure (AWS)
  • Extensive experience with AWS services including:
    • Compute:Lambda, ECS, EC2, Fargate
    • Storage:S3, EBS, EFS
    • Database:RDS, DynamoDB, Aurora
    • Networking:VPC, Route53, CloudFront, API Gateway
    • Monitoring:CloudWatch, X-Ray
  • AWS certifications (Solutions Architect, DevOps Engineer) preferred
DevOps & CI/CD
  • Expert-level knowledge ofGitLab(CI/CD pipelines, runners, GitOps)
  • AdvancedTerraformskills for infrastructure provisioning and management
  • Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
  • Proficiency with configuration management tools
Security
  • Hands-on experience withSAST(Static Application Security Testing) tools
  • Knowledge ofDAST(Dynamic Application Security Testing) methodologies
  • Understanding of security best practices, OWASP Top 10, and compliance frameworks
  • Experience with secrets management and identity access management (IAM)
Monitoring & Observability
  • Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
  • Log aggregation and analysis (CloudWatch Logs, Splunk)
  • Distributed tracing with aws X-Ray
Qualifications
  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
  • 3+ years in a lead or senior technical position
  • Proven track record of managing large-scale production systems
  • Experience with on-call rotations and incident management
  • GenAI based Applications:Working knowledge of LLMs and agentic applications a plus
  • Experience with serverless architectures and event-driven systems
  • Familiarity with chaos engineering principles and practices
  • Background in Agile/Scrum methodologies
  • Experience with multi-cloud or hybrid cloud environments

The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.

  • Eligible Locations for Hire: Richmond, VA, San Francisco, CA
  • The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA

Base Salary Range: Min: $146,700Mid: $190,500Max: $234,300 (Location: San Francisco)

The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate's qualifications, internal alignment considerations, district assignment, and geographic location.

The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at sf.hr.recruitment@sf.frb.org.

Full Time / Part Time Full timeRegular / TemporaryRegularJob Exempt (Yes / No)YesJob CategoryInformation Technology Family GroupWork ShiftFirst (United States of America)

The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.

Always verify and apply to jobs on Federal Reserve System Careers (https://rb.wd5.myworkdayjobs.com/FRS) or through verified Federal Reserve Bank social media channels.

Privacy Notice

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Federal Reserve Bank of San Francisco • San Francisco (CA)

On-site
USD 147,000 - 234,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Federal Reserve Bank of New York • San Francisco (CA)

On-site
USD 147,000 - 234,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

rb • San Francisco (CA)

On-site
USD 147,000 - 234,000
Senior SRE Lead — Cloud Reliability & Automation
Senior SRE Lead — Cloud Reliability & Automation

rb • San Francisco (CA)

On-site
USD 147,000 - 234,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Federal Reserve Bank of Boston • Boston (MA)

On-site
USD 90,000 - 140,000
Lead Site Reliability Engineer – Cloud & Automation
Lead Site Reliability Engineer – Cloud & Automation

Federal Reserve Bank of San Francisco • San Francisco (CA)

On-site
USD 147,000 - 234,000
Senior Network Automation Engineer, Full-Stack
Senior Network Automation Engineer, Full-Stack

Federal Reserve Bank of Richmond • Richmond (VA)

On-site
USD 93,000 - 156,000
Great medical benefits
Pension and 401(k) with employer match
Paid time off
+2
Lead SRE: Cloud Reliability & DevOps Leadership
Lead SRE: Cloud Reliability & DevOps Leadership

Federal Reserve Bank of San Francisco • Richmond (VA)

On-site
USD 147,000 - 234,000
Sr. Executive Assistant
Sr. Executive Assistant

Federal Reserve Bank of San Francisco • San Francisco (CA)

On-site
USD 83,000 - 132,000
Medical, dental, vision
401(k)/Pension
Paid time off
+4
Cloud AWS Support Reliability Engineer (SRE)
Cloud AWS Support Reliability Engineer (SRE)

Federal Reserve Bank of New York • New York (NY)

On-site
USD 160,000 - 230,000
Educational assistance
Onsite Health & Wellness Center