Principal Site Reliability Engineer

Habitat For Humanity Of Durham

Durham (NC)

On-site

USD 140,000 - 170,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

On-site health centers
Fully paid parental leave

Job summary

Fidelity Investments in Durham, NC seeks a Principal Site Reliability Engineer to lead enterprise reliability strategies and architect resilient systems. The role focuses on high availability, multi-region AWS deployments, and proactive observability through Datadog, Grafana, and ELK stacks.

The candidate will drive CI/CD pipelines, mentor engineers, and define SLOs/SLIs while improving performance and incident response. Onsite presence is expected with phased flexibility.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related field with 5+ years in an SRE/principal role.
  • Alternative: Master’s degree with 3+ years in SRE/principal role in financial services environment.

Responsibilities

  • Define and lead enterprise reliability strategies.
  • Architect resilient systems and infrastructure.
  • Publish performance test results with improvement recommendations.
  • Maintain scalability and resiliency of complex environment.

Skills

SRE practices
Distributed systems
Observability
AWS
CI/CD

Education

Bachelor's degree in CS/Engineering
Master's degree in CS/Engineering

Tools

Jenkins
Kubernetes
Python
Terraform
Datadog
ELK stack

Job description

Principal Site Reliability Engineer
Location
  • Durham, NC
Team
  • Technology
Experience Level

Senior Manager

Job Description:

Note: Fidelity will not provide immigration sponsorship for this position.

Position Description:

Deploys and supports distributed, multi-tiered systems at scale while ensuring high availability and fault tolerance across multiple environments. Builds and operates resilient platforms in Amazon Web Services (AWS) using Elastic Compute Cloud (EC2), Simple Storage Service (S3), and Auto Scaling Groups for dynamic resource management. Designs, develops, and executes performance tests using Java-based frameworks, Apache JMeter, k6, and Rush-hour to validate system behavior under day-to-day traffic patterns. Defines and implements observability practices to monitor system health, latency, and error rates through metrics, logs, and distributed tracing using Datadog, Grafana, Splunk, and the Elasticsearch, Logstash, and Kibana (ELK) stack. Automates operational workflows with Python and Shell scripting to enhance efficiency and reduce manual tasks. Supports consistent build, deployment, and orchestration processes using cloud computing and DevOps technologies -- Continuous Integration and Continuous Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets, and implementing proactive monitoring and incident response strategies. Builds and refines methodologies for performance, load, stress, and chaos testing and develops analytics and reports aligned with business needs to improve system resilience and optimization.

Primary Responsibilities:

  • Defines and leads enterprise-level reliability strategies.
  • Architects resilient systems and infrastructure.
  • Creates and publishes performance test results report with recommendations on quality improvement.
  • Maintains scalability and resiliency of complex environment.
  • Implements advanced observability practices and techniques at scale.
  • Manages and interprets large datasets using query languages and visualization tools.
  • Advises senior leadership on reliability engineering best practices.
  • Mentors junior engineers.
  • Performs independent and complex technical and functional analysis for multiple divisional initiatives.
  • Develops innovative solutions to improve system availability, scalability, and performance.
  • Designs, implements, and maintains performance test frameworks.

Education and Experience:

Bachelor’s degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment.

Or, alternatively, Master’s degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and three (3) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment.

Skills and Knowledge:

Candidate must also possess:

  • Demonstrated Expertise ('DE') performing software performance benchmarking and engineering for online financial web applications, Application Programming Interfaces (APIs), and mobile transactions according to DevOps practices, using performance benchmarking tools Rushhour, Locust, K6, and JMeter; and configuring CI/CD and test automation, using Jenkins, Sonar, Ant, Maven, Artifactory, and Terraform in AWS.
  • 'DE' solutioning, designing, architecting, and building scalable and resilient enterprise-grade software platforms using cloud-based architecture and AWS services (EC2, Elastic Container Service (ECS), Lambda, Elastic MapReduce (EMR), and CloudFormation); developing microservices on Elastic Kubernetes Service (EKS), implementing CI/CD pipelines using DevOps tools (Bitbucket, GitHub, Artifactory, Sonar, Veracode, and Helm), and adhering to DevOps practices along with leveraging Java, Python, Spring Boot, Docker, EKS, and AWS.
  • 'DE' analyzing and monitoring system and application performance across Apache, NGINX, Java, and Node.js platforms, and Linux and Windows environments, using Splunk, Datadog, Kibana, Grafana, and AWS CloudWatch; diagnosing performance bottlenecks, recommending tuning strategies, reducing Mean Time to Detect (MTTD) and Mean Time to Repair (MTTR), using Application Performance Monitoring (APM) tools -- Dynatrace, New Relic, Splunk, and Datadog; and performing capacity planning to optimize Central Processing Unit (CPU), memory, and process configurations.
  • 'DE' instrumenting advanced observability practices at scale across cloud-native and hybrid environments; defining and tracking SLOs and SLIs to ensure reliability and performance metrics, using Python automation, Infrastructure as Code (IaC) methodologies, and observability tools (Datadog, Splunk, Dynatrace, Grafana, and the ELK stack; developing custom dashboards, alerting rules, and automated incident response workflows to proactively detect and resolve performance degradations, using Datadog, Catchpoint, Grafana, ELK stack, and Cloudwatch; and enabling actionable insights through trace-level correlation of end-to-end (E2E) user journeys and system behaviors, using Dynatrace, Splunk, Draw.io, and Miro.

#PE1M2

#LI-DNI

Fidelity’s Onsite Working Model Fidelity is transitioning to a full-time onsite working model through a phased rollout across regions and roles. Currently, some roles and locations require 100% onsite presence, while others require less. Onsite expectations are likely to evolve as the rollout continues. This transition does not apply to fully remote roles.

Certifications:
Category:

Information Technology

Please be advised that Fidelity’s business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement-related financial activities and the rules and regulations of numerous self-regulatory organizations, including FINRA, among others. Those laws and regulations may restrict Fidelity from hiring and/or associating with individuals with certain Criminal Histories.

Benefits that balance life and work

From our fully paid parent leave to our on-site health and wellness centers, our benefits support the belief that more balance you have, the better you can achieve your goals.

Company overview

At Fidelity, we are passionate about making our financial expertise broadly accessible and effective in helping people live the lives they want. We are a privately held company that places a high degree of value in creating and nurturing a work environment that attracts the best talent and reflects our commitment to our associates. We are proud of our diverse and inclusive workplace where we respect and value our associates for their unique perspectives and experience.

Reasonable accommodations

Fidelity will reasonably accommodate applicants with disabilities who need adjustments to participate in the application or interview process. To initiate a request for an accommodation contact the HR Accommodation Team by sending an email to accommodations@fmr.com, or by calling 800-835-5099, prompt 2, option 3.

Equal opportunity employer

Fidelity Investments is an equal opportunity employer. We believe that the most effective way to attract, develop, and retain a diverse workforce is to build an enduring culture of inclusion and belonging.

Applicant screening

At Fidelity, we value honesty, integrity, and the safety of our associates and customers within a heavily regulated industry. Certain roles may require candidates to go through a preliminary credit check during the screening process. Candidates who are presented with a Fidelity offer will need to go through a background investigation and may be asked to provide additional documentation as requested. This investigation includes but is not limited to a criminal, civil litigations and regulatory review, employment, education, and credit review (role dependent). These investigations will account for 7 years or more of history, depending on the role. Where permitted by federal or state law, Fidelity will also conduct a pre-employment drug screen, which will review for the following substances: Amphetamines, THC (marijuana), cocaine, opiates, phencyclidine.

AI Guidelines

Learn about our guidelines for use of AI when applying for a Fidelity job

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director, Technology Management - Software Engineering
Director, Technology Management - Software Engineering

Habitat For Humanity Of Durham • Durham (NC)

On-site
USD 150,000 - 190,000
Onsite health centers
Fully paid parent leave
Principal Digital Assets Engineer
Principal Digital Assets Engineer

Habitat For Humanity Of Durham • Durham (NC)

On-site
USD 150,000 - 190,000
Senior Manager, Advanced Data Analytics and Insights
Senior Manager, Advanced Data Analytics and Insights

Fidelity • Boston (MA)

On-site
USD 127,000 - 166,000
Fully paid parental leave
On-site health and wellness centers
Senior Manager, Data Analytics and Insights
Senior Manager, Data Analytics and Insights

Fidelity • Boston (MA)

On-site
USD 131,000 - 166,000
Senior Systems Engineer
Senior Systems Engineer

Habitat For Humanity Of Durham • Durham (NC)

On-site
USD 110,000 - 160,000
Principal Cloud Engineer
Principal Cloud Engineer

Habitat For Humanity Of Durham • Durham (NC)

On-site
USD 150,000 - 200,000
Fully paid parental leave
On-site health and wellness centers
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Socket.dev • North Carolina

On-site
USD 150,000 - 210,000
Principal Full Stack Engineer - AI Technology
Principal Full Stack Engineer - AI Technology

Habitat For Humanity Of Durham • Durham (NC)

On-site
USD 170,000 - 250,000
Onsite health & wellness centers
Fully paid parental leave
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Worky • Durham (NC)

On-site
USD 150,000 - 210,000
Manager, Quant Data Analytics and Insights
Manager, Quant Data Analytics and Insights

Fidelity • Boston (MA)

On-site
USD 126,000 - 141,000
On-site health centers
Fully paid parental leave
Work-life balance programs