Lead Site Reliability Engineer — Scalable Infra for ML Pipelines

Treeswift Inc

New York (NY)

Hybrid

USD 160,000 - 215,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology company is seeking a full-time SRE/infrastructure engineer to enhance and scale their data processing platform. This hybrid role is based in Lower Manhattan, requiring two days in-office weekly. Responsibilities include designing reliability for data pipelines using technologies like AWS and Kubernetes. Ideal candidates will have 7-10 years' experience, strong debugging skills, and be passionate about operational excellence. Competitive salary range is $160,000 to $215,000, commensurate with experience.

Qualifications

  • 7-10 years of experience in software engineering focused on SRE, infrastructure, or DevOps.
  • Hands-on knowledge of cloud-native solutions and container orchestration.
  • Strong debugging skills in Linux environments.

Responsibilities

  • Partner with data and engineering teams for pipeline execution.
  • Design and implement reliability for data pipeline operations.
  • Lead improvements in infrastructure for platform performance.

Skills

Observability
Systems engineering
Cloud infrastructure
Kubernetes
Communication

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
Apache Airflow
AWS

Job description

A technology company is seeking a full-time SRE/infrastructure engineer to enhance and scale their data processing platform. This hybrid role is based in Lower Manhattan, requiring two days in-office weekly. Responsibilities include designing reliability for data pipelines using technologies like AWS and Kubernetes. Ideal candidates will have 7-10 years' experience, strong debugging skills, and be passionate about operational excellence. Competitive salary range is $160,000 to $215,000, commensurate with experience.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
SRE: Scalable ML Infra & CI/CD Architect
SRE: Scalable ML Infra & CI/CD Architect

Baseten • San Francisco (CA)

On-site
USD 165,000 - 330,000
Senior Site Reliability Engineer — NYC, Equity & Impact
Senior Site Reliability Engineer — NYC, Equity & Impact

Menlo Ventures • New York (NY)

On-site
Confidential
Medical, Dental & Vision benefits
Generous parental leave
401(K) with company match
+1
Site Reliability Engineer — Scale, Incidents & Infra
Site Reliability Engineer — Scale, Incidents & Infra

General Intuition & Medal • New York (NY)

On-site
USD 120,000 - 160,000
Competitive salary and meaningful equity
Comprehensive medical, dental, and vision coverage
401(k)
+5
Senior SRE: ML Infra at Scale, Multi-Cloud K8s
Senior SRE: ML Infra at Scale, Multi-Cloud K8s

The Consensus • New York (NY)

On-site
USD 110,000 - 140,000
Competitive compensation
100% insurance coverage
Flexible PTO policy
+3
Senior Site Reliability Engineer, Data Platform & Cloud
Senior Site Reliability Engineer, Data Platform & Cloud

Optomi • Town of Florida (NY)

On-site
USD 145,000 - 160,000
Staff Site Reliability Engineer — NYC On-site, Equity
Staff Site Reliability Engineer — NYC On-site, Equity

Menlo Ventures • New York (NY)

On-site
Confidential
Medical, Dental & Vision benefits
Generous parental leave
401(K) with generous company match
+1
Lead SRE—AI/ML Platform Reliability & Automation
Lead SRE—AI/ML Platform Reliability & Automation

Compunnel, Inc. • Austin (TX)

On-site
USD 120,000 - 160,000
Senior SRE & Software Engineer — Scalable Infra
Senior SRE & Software Engineer — Scalable Infra

Harvey • San Francisco (CA)

On-site
USD 200,000 - 260,000
SRE: AI Inference Platform & ML Systems
SRE: AI Inference Platform & ML Systems

Cohere • San Francisco (CA)

Hybrid
USD 120,000 - 150,000