Senior Site Reliability Engineer

Mission Staffing

New York (NY)

Hybrid

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology-driven investment firm is seeking a Senior Site Reliability Engineer to establish innovative SRE practices across its infrastructure. This key position in a hybrid setting involves collaborating with engineering teams to enhance service reliability and performance while maintaining high availability for critical production systems. The ideal candidate will have over 8 years of experience and expertise in observability tools like Prometheus and Grafana, ensuring robust operational performance in a complex tech ecosystem.

Qualifications

  • 8+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering roles.
  • Experience operating large-scale distributed systems in production environments.
  • Strong expertise with observability and monitoring platforms.
  • Strong observability and monitoring expertise (Prometheus, Grafana, Loki, Tempo, OpenTelemetry).
  • Deep understanding of containerization/orchestration (Docker, Kubernetes).
  • Experience across cloud (AWS) and on-prem environments.
  • Strong scripting/automation (Python, Bash, Go).
  • Experience building/maintaining CI/CD pipelines and modern DevOps workflows.

Responsibilities

  • Establish and evolve Site Reliability Engineering practices and standards.
  • Design and scale observability and monitoring platforms.
  • Define reliability standards for applications running in Kubernetes environments.
  • Define reliability standards for Kubernetes-based apps for performance and resiliency.
  • Build automation to improve deployment, health monitoring, and recovery.
  • Collaborate with engineers to improve service stability and scalability.
  • Promote SRE best practices like SLOs, incident reviews, blameless post-mortems.

Skills

Site Reliability Engineering
Observability and monitoring platforms
Containerization and orchestration
Scripting and automation with Python/Bash/Go
Docker & Kubernetes
Scripting (Python, Bash, Go)
CI/CD & DevOps

Tools

Prometheus
Grafana
Loki
Docker
Kubernetes
Docker
Kubernetes
AWS

Job description

Location: New York City or Chicago (Hybrid)

A technology-driven investment firm is expanding its Platform Engineering organization and is seeking an experienced Senior Site Reliability Engineer to help shape reliability practices across its infrastructure and production environments. This role offers the opportunity to build and scale SRE practices from the ground up, partnering closely with platform, DevOps, and cloud engineering teams to drive reliability, performance, and operational maturity across a complex technology ecosystem. You will work across both cloud and on-premise environments, supporting highly critical production systems including trading and data platforms. The role combines hands-on engineering with strategic influence, helping define reliability standards and operational frameworks across the organization.

What You’ll Do
  • Help establish and evolve Site Reliability Engineering practices, standards, and operational processes across engineering teams
  • Design and scale observability and monitoring platforms using tools such as Prometheus, Grafana, Loki, Tempo, and OpenTelemetry
  • Participate in a team-based on-call rotation (approximately one week per month) supporting critical production systems
  • Define reliability standards for applications running in Kubernetes environments, ensuring optimal configuration for performance, cost, and resiliency
  • Build automation and tooling to improve deployment pipelines, system health monitoring, and recovery processes
  • Partner with engineering teams to improve service stability, scalability, and fault tolerance
  • Promote SRE best practices such as service level objectives (SLOs), incident reviews, and blameless post-mortems
What You Bring
  • 8+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering roles
  • Experience operating large-scale distributed systems in production environments
  • Strong expertise with observability and monitoring platforms, including Prometheus, Grafana, Loki, Tempo, and OpenTelemetry
  • Deep understanding of containerization and orchestration technologies, including Docker and Kubernetes
  • Experience working across cloud infrastructure (AWS preferred) and on-premise environments
  • Strong scripting and automation skills using Python, Bash, or Go
  • Experience building and maintaining CI/CD pipelines and modern DevOps workflows
What Makes You Stand Out
  • Passion for building reliable, scalable infrastructure and improving operational maturity
  • Ability to translate complex reliability concepts into practical engineering solutions
  • Strong collaboration skills when working across engineering, platform, and infrastructure teams
  • A mindset focused on automation, observability, and continuous improvement
Why This Role
  • Opportunity to define and build SRE practices from the ground up
  • Work on mission-critical infrastructure supporting high-performance systems
  • Collaborate with platform, cloud, and engineering teams building modern infrastructure at scale
  • High-impact role within a technology-focused financial environment
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Equity or bonus opportunities
Health benefits
Paid time off
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
SRE
SRE

Benton Partners • Chicago (IL), New York (NY)

On-site
USD 175,000 - 225,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Site Reliability Engineer - Banking & Finance
Senior Site Reliability Engineer - Banking & Finance

Hamilton Barnes Associates Limited • New York (NY)

Hybrid
USD 360,000 - 440,000
Strong compensation and bonus potential
Collaborative engineering culture
Work on mission-critical systems
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

Hybrid
USD 130,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hamilton Barnes ? • New York (NY)

Hybrid
USD 340,000 - 400,000
Strong compensation and bonus potential
Hybrid working environment
Collaborative engineering culture