Site Reliability Engineer (SRE)

New York Technology Partners

Chicago (IL)

On-site

USD 120,000 - 160,000

Full time

11 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

New York Technology Partners is seeking highly capable Site Reliability Engineers to build, operate, and continuously improve the cloud platform powering our products. You will serve as an engineering first-responder for infrastructure incidents, restoring service quickly, identifying root causes, and engineering permanent improvements to reliability and scalability.

You’ll own the operational health of infrastructure spanning cloud platforms, Kubernetes and container orchestration, service

Responsibilities

  • Serve as a primary engineering responder for service incidents.
  • Diagnose issues across cloud infrastructure, Kubernetes, networking, storage, and CI/CD.
  • Restore service quickly while balancing immediate mitigation with long-term reliability.
  • Define SLIs/SLOs and drive post-incident improvements.
  • Build and maintain resilient cloud infrastructure and Kubernetes clusters.
  • Automate deployment pipelines and GitOps workflows.
  • Improve monitoring, logging, and observability.

Skills

Kubernetes
Terraform
CI/CD
Observability
Automation

Tools

GitOps tooling
Containers

Job description

We are looking for highly capable Site Reliability Engineers to build, operate, and continuously improve the cloud platform that powers our products. Every SRE serves as an engineering first-responder for infrastructure incidents, restoring service quickly, identifying root causes, and engineering permanent improvements that strengthen the reliability, scalability, and operational excellence of our platform.

You’ll own the operational health of infrastructure spanning cloud platforms, Kubernetes and container orchestration, service meshes, networking, storage, databases, Infrastructure as Code, CI/CD and GitOps pipelines, and the core platform services that power our applications. Through automation, observability, and continuous improvement, you’ll reduce operational risk, eliminate repetitive work, and help engineering teams deliver software with confidence.

As you progress from Site Reliability Engineer to Senior Site Reliability Engineer and Lead Site Reliability Engineer, your responsibilities expand from independently operating production systems to leading complex incident response, mentoring engineers, improving operational practices, and helping elevate the effectiveness of the SRE team.

This is a high-trust, high-impact role for engineers who enjoy solving difficult operational problems, automating away operational toil, and building resilient platforms that enable the business to scale.

Key Responsibilities
Primary: Reliability Engineering and Incident Response
  • Serve as a primary engineering responder for service incidents.
  • Diagnose issues across cloud infrastructure, Kubernetes, networking, storage, databases, CI/CD systems, and platform services.
  • Restore service quickly while balancing immediate mitigation with long-term reliability.
  • Perform root cause analysis and drive corrective actions through completion.
  • Clearly communicate incident status, customer impact, and recovery progress during production events.
  • Participate in post-incident reviews and continuously improve operational practices.
  • Ensure services are observable through metrics, logs, traces, health checks, and actionable alerting.
  • Contribute to the implementation of telemetry across our backend systems to improve observability and operational insight.
  • Define service level indicators (SLIs), service level objectives (SLOs), and error budgets that align platform reliability with business priorities.
  • Communicate complex technical issues clearly to both technical and non-technical stakeholders.
  • Design, deploy, and maintain resilient cloud infrastructure.
  • Build and support Kubernetes clusters and platform services.
  • Develop and maintain Infrastructure as Code using Terraform or similar tools.
  • Engineer improvements that increase platform reliability, scalability, resiliency, and operational efficiency.
  • Strengthen disaster recovery, backup, and business continuity capabilities.
  • Continuously improve monitoring, alerting, logging, and observability.
  • Perform capacity planning and infrastructure forecasting to support future growth and maintain service reliability.
  • Continuously identify and eliminate operational toil through automation, tooling, and platform improvements.
Supporting: Automation & Platform Operations
  • Automate operational processes to reduce manual effort and operational risk.
  • Improve deployment pipelines and GitOps workflows.
  • Build internal tooling that improves engineering productivity.
  • Partner with Product Engineering, Production Engineering, and Security to improve production readiness and platform reliability.
  • Improve deployment reliability, release processes, and the developer experience through automation and platform engineering.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Senior Engineer - Site Reliability Engineering
Senior Engineer - Site Reliability Engineering

LSEG (London Stock Exchange Group) • Allen (TX)

On-site
USD 140,000 - 190,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000