Senior Site Reliability Engineer - Cloud Platform | Energy Trading & Infrastructure Firm

Techfellow Limited

New York (NY)

Hybrid

USD 270,000 - 330,000

Full time

43 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Techfellow Limited seeks an experienced Platform Infrastructure Engineer in New York. The role centers on reliability engineering, AWS architecture, Kubernetes, IaC and observability, moving toward resilient cloud-native patterns.

Responsibilities include ownership of SRE practices, OpenTelemetry adoption, DR design, incident response and automated deployment pipelines, while staying deeply technical and influencing platform design across the firm.

Qualifications

  • 6+ years hands-on experience across SRE, Platform Engineering, DevOps or production infra.
  • Own business-critical systems rather than just supporting them.
  • Design and operate highly available AWS infrastructure with resilient patterns.
  • Build and evolve Kubernetes platform infra using Kubernetes, Docker, Terraform and Git.
  • Own IaC and delivery automation; improve CI/CD, testing and environment consistency.
  • Develop observability across metrics, logs and tracing with OpenTelemetry and Datadog.
  • Establish reliability practices: SLOs/SLIs, error budgets, capacity and health.
  • Strengthen incident management with alerting, runbooks and RCA.

Responsibilities

  • Design and operate disaster recovery capabilities with RTO/RPO targets.
  • Improve deployment and upgrade processes; ensure repeatable patterns.
  • Strengthen monitoring and incident response across platforms.
  • Automate infrastructure tasks using Python, Go or Bash.
  • Lead platform decisions on cloud-native patterns and security controls.
  • Collaborate with product teams to keep services resilient in production.

Skills

SRE
Platform Engineering
DevOps
AWS
Kubernetes
Docker
Terraform
Git
CI/CD
OpenTelemetry
Datadog
Python
Go
Bash
Argo CD
Helm
Karpenter
Crossplane
Serverless

Tools

Kubernetes
Docker
Terraform
Git
OpenTelemetry

Job description

[Up to c. $300k Comp Package | Hybrid Working - 4 Days In Office]

Role Overview

We’re representing a global commodities trading and investment organisation expanding the infrastructure capability behind its firmwide investment, trading and analytics technology platform. The environment combines cloud infrastructure, container platforms, internal applications, APIs and data-driven tooling used across business-critical workflows.

This hire will sit within a small Platform Infrastructure function and take substantial hands-on ownership across reliability engineering, AWS architecture, Kubernetes, infrastructure automation and production observability. A major part of the mandate is moving the platform towards more resilient, repeatable cloud-native operating patterns while improving how services are deployed, monitored and recovered. Alongside core SRE responsibilities, the role offers meaningful ownership of the organisation’s OpenTelemetry adoption, Kubernetes evolution and disaster recovery capability. It suits an experienced engineer who wants to remain deeply technical while influencing how critical infrastructure is designed and operated across the firm...

Role Snapshot

  • Bring 6+ years of hands-on experience across Site Reliability Engineering, Platform Engineering, DevOps or production infrastructure, with evidence of owning business-critical systems rather than primarily supporting them
  • Design and operate highly available AWS infrastructure, applying resilient patterns across compute, storage, networking, scaling, load balancing, backup and recovery
  • Build and evolve Kubernetes-based platform infrastructure using strong practical experience with Kubernetes, Docker, Terraform and Git
  • Own Infrastructure as Code and delivery automation, creating repeatable deployment patterns and improving CI/CD, testing, upgrades, patching and environment consistency
  • Develop the observability platform across metrics, logs and distributed tracing, including continued adoption of OpenTelemetry, alongside tooling such as Datadog
  • Establish measurable reliability practices covering SLOs, SLIs, error budgets, capacity, operational health and reduction of recurring production failure
  • Strengthen incident management through effective alerting, troubleshooting, runbooks, root-cause analysis and engineering changes that prevent repeat incidents
  • Design and validate high-availability and disaster recovery capabilities, including RTO/RPO targets, recovery procedures, failover testing and dependency planning
  • Use Python, Go, Bash or comparable engineering automation to reduce manual infrastructure work, supported by strong Linux, networking and production troubleshooting fundamentals
  • (Preferred) Experience with Argo CD, Helm, Karpenter, Crossplane, progressive delivery, serverless architectures, additional cloud platforms or security controls within regulated or financial-services environments
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

On-site
USD 140,000 - 200,000
Senior DevOps Engineer | Dynamic Asset Management Leader
Senior DevOps Engineer | Dynamic Asset Management Leader

Techfellow Limited • New York (NY)

Hybrid
USD 350,000 - 425,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

On-site
USD 140,000 - 170,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

On-site
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000