Site Reliability Engineer

Harrison Clarke

New York (NY)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Harrison Clarke is seeking a Senior Site Reliability Engineer in New York to own and optimize cloud infrastructure across AWS and GCP. You'll maintain and secure distributed systems, partner with engineering teams to streamline CI/CD processes, and monitor high-availability environments. Candidates should have over 5 years of experience in cloud operations and a strong grasp of SRE methodologies. The ideal candidate thrives in collaborative settings and brings a low-ego working style, especially under pressure.

Qualifications

  • 5+ years in cloud-based systems operations as an SRE or DevOps engineer.
  • Hands-on experience with infrastructure as code and configuration management.
  • Strong command of SRE methodologies including SLOs/SLIs and incident management.

Responsibilities

  • Maintain, improve, and secure cloud infrastructure across AWS and GCP.
  • Partner with engineering teams for deployment and troubleshooting.
  • Drive CI/CD maturity using Jenkins.

Skills

Cloud-based systems operations
Infrastructure as code
SRE methodologies
Networking fundamentals
Production workloads management
Incident management
Scripting or programming language
Collaboration under pressure

Tools

AWS
GCP
Jenkins
Kubernetes
Prometheus
ELK

Job description

Harrison Clarke partners exclusively with venture-backed technology companies building category-defining products. We are currently conducting a confidential retained search on behalf of one of our flagship portfolio companies, a well-capitalised, mission-driven technology firm that has been operating at scale since 2014, with millions of global users and a reputation for rigorous engineering.

The Role

As Senior Site Reliability Engineer, you will own the infrastructure foundation that the entire engineering organization depends on. This isn't a support function, it's a strategic one. You'll work at the intersection of reliability, scalability, and developer experience, ensuring that a high-availability distributed system stays fast, resilient, and ready to grow.

What You'll Be Doing
  • Maintaining, improving, and securing cloud infrastructure across AWS and GCP, alongside Linux systems at scale
  • Partnering directly with engineering teams to streamline deployment, packaging, and troubleshooting of complex distributed applications
  • Driving CI/CD maturity using Jenkins and adjacent tooling
  • Building, operating, and continuously improving Kubernetes clusters in production; serving as the internal authority on container orchestration
  • Leading application migration efforts onto Kubernetes in close collaboration with development squads
  • Owning internal platform services including Prometheus and ELK
  • Monitoring high-availability environments and responding to incidents with urgency and rigour; conducting thorough, blameless post-mortems
  • Participating in architecture and code reviews, setting the bar for infrastructure best practice
  • Evaluating emerging technologies and making pragmatic decisions on adoption
  • Identifying and eliminating toil through intelligent automation
What You Bring
  • 5+ years in cloud-based systems operations as an SRE or DevOps engineer
  • Hands-on experience with infrastructure as code and configuration management
  • Strong command of SRE methodologies: SLOs/SLIs/error budgets, capacity planning, disaster recovery testing
  • Deep understanding of networking fundamentals
  • Proven track record managing production workloads with sophisticated monitoring and alerting
  • Comfort with on-call responsibilities and a systematic approach to incident management
  • Proficiency in at least one scripting or programming language
  • A collaborative, low-ego working style — you raise the team up, especially under pressure
Bonus Points
  • Ability to read and reason about Go, Rust, C++, or TypeScript
  • Experience applying AI-driven approaches to operational automation
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

FORT • United States

Hybrid
USD 150,000 - 180,000
Healthcare benefits
Flexible work environment
Large-scale cloud platform project
+1
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
SRE (Site Realiability Engineer)
SRE (Site Realiability Engineer)

STRATIS Cloud Tech Solutions INC • Arkansas

On-site
USD 100,000 - 130,000
Competitive salary and benefits
Growth and learning opportunities
Friendly and collaborative team environment
Site Reliability Engineer
Site Reliability Engineer

Triwill Group • United States

Remote
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000