Senior Site Reliability Engineer

Fidelity Investments

Durham (NC)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fidelity Investments is seeking an experienced Site Reliability Engineer to join our Enterprise Infrastructure. You will help scale production systems with automation, resilience, and observability, focusing on reliability and performance.

The role blends software and systems engineering, with opportunities to broaden expertise in cloud, CI/CD, and chaos testing. The candidate should have 5+ years building or operating distributed systems in AWS, strong scripting abilities, and solid

Qualifications

  • Bachelor’s degree or higher in a technology related field; master’s is a plus.
  • 5+ years deploying and supporting highly distributed multi-tiered systems.
  • 3+ years in AWS cloud development and migration; building resilient platforms.
  • 3–5 years software development in Python/NodeJS/Java with SDLC focus.
  • Strong observability and monitoring experience with multiple tools.
  • Performance testing experience with K6/JMeter and related tools.
  • Experience with Infrastructure as Code (Terraform, IAM, ARM, Chef).
  • Solid understanding of Cloud and DevOps concepts, CI/CD pipelines.
  • Proven ability to scale and maintain high availability in production.

Skills

Cloud AWS
Python/Java/Go
CI/CD & DevOps
Observability tooling
Performance testing
Chaos engineering
Container orchestration
Distributed systems
Data analysis/queries
Communication skills

Education

Bachelor’s degree or higher in a technology related field
Master’s degree is a plus

Tools

Kubernetes
Terraform
IAM
OpenTelemetry
Datadog
Splunk
Kibana
Prometheus
Grafana
ELK/OpenSearch
K6
JMeter

Job description

Job Description

Note: Fidelity will not provide immigration sponsorship for this position

The Role

Our Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high availability with resilience by using automation and Infrastructure Code. We build reliability into our ecosystem by applying best practices in Resiliency Engineering, Automation, Observability, Performance testing and Chaos testing. The team comes from diverse technical backgrounds, and the responsibilities provide the opportunity for a variety of challenges. Ideal candidates will have a background in either software engineering or systems engineering with a desire to learn the other or previous experience as an SRE. We are looking for a Systems Thinking, SRE Engineer who has helped teams scale through production insights, operational automation, developer guidance, real‑time metrics, automation.

The Expertise And Skills You Bring
  • Bachelor’s degree or higher in a technology related field (Engineering, Computer Science, etc.) required, master’s degree is a plus.
  • Minimum 5 years of hands‑on experience deploying and/or supporting highly distributed multi‑tiered systems at a scale.
  • 3+ years of experience in Cloud development (AWS) and migration skills; experience with building and operating highly resilient platforms in AWS cloud environments.
  • 3‑5 years of experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation.
  • Ensure platforms meet high availability, scalability, fault tolerance, and disaster recovery requirements.
  • Hands‑on experience with one or more observability tools (Datadog, Splunk, Kibana, Prometheus, Grafana, ELK/OpenSearch, Open Telemetry).
  • Hands‑on experience in designing, developing, and executing performance tests using K6/JMeter and other performance testing tools to ensure comprehensive performance testing.
  • Define Performance Test Strategy Document: set approach, metrics, benchmarks, baseline, user response requirements environments, technical environment and data conditions, and toolsets to use in executing the performance testing.
  • Experience in performance testing types: Load testing, Stress testing, Scalability testing, Spike testing, Volume testing, Chaos testing, Endurance/Soak testing.
  • Hands‑on experience with container orchestration, preferably with Kubernetes.
  • Experience identifying memory leakage, connection issues and throughput bottlenecks in various technologies such as web application(s), infrastructure, and Cloud.
  • Strong knowledge of CI/CD pipelines and DevOps practices.
  • Familiarity with chaos engineering and resilience testing tools (e.g., Chaos Monkey, Gremlin).
  • Experience working in high‑availability, large‑scale production environments.
  • Strong programming/scripting skills in one or more: Python, Java, Go, or Bash.
  • Expertise in automation frameworks and tools for performance validation.
  • Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef).
  • Solid understanding of Cloud Computing and DevOps concepts including CI/CD pipelines.
  • Experienced in instrumentation with systems skills on building and operating, monitoring, logging, alerting services of distributed systems at scale.
  • Proven experience in maintaining scalability and resiliency of complex environments.
  • Proven experience in implementing advanced observability practices and techniques at scale.
  • Ability to triage, execute root cause analysis, and be decisive under pressure.
  • Experience managing and interpreting large datasets using query languages and visualization tools.
  • Proficient communication skills with an ability to reach both technical and non‑technical audience.
  • Ability to work with a variety of individuals and groups, both in person and virtually, in a constructive and collaborative manner and build and maintain effective relationships.
  • Experience in designing, implementing, and maintaining performance test frameworks, which will validate to a high degree of confidence, the production readiness of software applications and infrastructure for stability and performance. Solid understanding of AWS services and experience setting up test environments on AWS (S3, EC2, RDS, etc.).
Fidelity’s Onsite Working Model

Fidelity is transitioning to a full‑time onsite working model through a phased rollout across regions and roles. Currently, some roles and locations require 100% onsite presence, while others require less. Onsite expectations are likely to evolve as the rollout continues. This transition does not apply to fully remote roles.

Certifications
  • Information Technology

Please be advised that Fidelity’s business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement‑related financial activities and the rules and regulations of numerous self‑regulatory organizations, including FINRA, among others. Those laws and regulations may restrict Fidelity from hiring and/or associating with individuals with certain Criminal Histories.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Soteria Reinsurance Ltd. • Town of Texas (WI), Northern (KY)

Hybrid
USD 150,000 - 210,000
Director, Full Stack Engineering
Director, Full Stack Engineering

Worky • Merrimack (NH)

On-site
USD 180,000 - 240,000
null
Senior Software Engineer/Developer
Senior Software Engineer/Developer

Fidelity Investments • North Carolina

On-site
USD 95,000 - 140,000
Principal Systems Engineer
Principal Systems Engineer

Fidelity Investments • North Carolina

On-site
USD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

TalentAlly • Roanoke (TX)

On-site
USD 120,000 - 180,000
Principal Full Stack Engineer
Principal Full Stack Engineer

Soteria Reinsurance Ltd. • Merrimack (NH), Northern (KY)

Hybrid
USD 150,000 - 210,000
Principal Full Stack Engineer
Principal Full Stack Engineer

Worky • Durham (NC)

On-site
USD 120,000 - 180,000
Principal Software Engineer/Developer
Principal Software Engineer/Developer

Worky • Merrimack (NH)

On-site
USD 150,000 - 190,000
Senior Software Engineer
Senior Software Engineer

Fidelity Investments • Roanoke (TX)

Hybrid
USD 110,000 - 140,000
Principal Full Stack Engineer
Principal Full Stack Engineer

Fidelity Investments • North Carolina

On-site
USD 130,000 - 180,000