Director Site Reliability Engineering

techchaintalent

New York (NY)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Stellar Development Foundation seeks a Director of Site Reliability Engineering to lead a team of 4 SREs and shape how engineering teams own, operate, and improve production services. This hands-on leadership role reports to the CTO and will set vision, operating model, and culture for SRE across core infrastructure.

The ideal candidate combines strong technical judgment with pragmatic leadership to enable reliable, scalable services and high developer productivity across a distributed

Qualifications

  • 10+ years in SRE or related engineering roles.
  • 5+ years leading infrastructure or reliability teams.
  • 3+ years cloud infra experience in AWS or GCP.
  • 3+ years Kubernetes, IaC, CI/CD, and deployment safety.

Responsibilities

  • Lead a distributed SRE team of 4 with a clear vision and priorities.
  • Define and roll out a Service Ownership and Maturity Framework across engineering.
  • Own and improve core infrastructure services, including cloud foundations, Kubernetes, CI/CD, observability, and automation.
  • Help engineering teams become stronger owners and operators of their services.
  • Make reliability and developer productivity measurable via metrics and operational intelligence.
  • Improve deployment automation, resilience, and disaster recovery readiness.
  • Mature incident response, escalation, and postmortems across a distributed team.
  • Build self-service infrastructure to reduce toil and speed up delivery.
  • Collaborate with Security, Compliance, Legal, Finance, Procurement, and Corporate IT.

Skills

Leadership
Coaching
Strategic thinking
Communication

Tools

Kubernetes
CI/CD
GitHub
AWS
GCP
Terraform

Job description

About Stellar

Stellar is a decentralised, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient than most blockchain-based systems. Since 2014, the Stellar Development Foundation has helped fuel the tremendous growth of the Stellar blockchain network, an open-source platform that operates at high scale today.

About the Role

SDF is hiring a Director of Site Reliability Engineering to lead a team of 4 SREs and shape how engineering teams own, operate, and improve production services. This is a backfill reporting directly to the CTO.

The Director will set the vision, operating model, and culture for SRE while owning the core infrastructure services that help SDF engineering teams build, deploy, observe, and operate software with confidence. Engineering teams at SDF own the services they build; SRE provides the frameworks, standards, shared infrastructure, tooling, observability practices, and enablement model that make strong service ownership possible across engineering.

This is a hands-on leadership role. The ideal candidate brings strong technical judgment, pragmatic leadership, and the ability to influence through trust, clarity, and execution.

Key Responsibilities
  • Lead, coach, and develop a distributed SRE team of 4, setting a clear vision, charter, operating model, priorities, and success measures
  • Define and roll out a Service Ownership and Maturity Framework across engineering, with expectations that vary appropriately by service criticality
  • Own and improve core engineering infrastructure services, including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation
  • Help engineering teams become stronger owners and operators of their services through better standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices
  • Make reliability, operational maturity, infrastructure health, and developer productivity more measurable through trusted metrics and practical operational intelligence
  • Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability based on actual impact and risk
  • Mature incident response, escalation, postmortems, and on-call health across a geographically distributed team
  • Build paved paths and self-service infrastructure that reduce toil, lower cognitive load, and help engineering teams move faster
  • Partner closely with Security, Compliance, Legal, Finance, Procurement, and Corporate IT where infrastructure, access management, cloud operations, vendor review, or controls intersect with engineering
  • Pragmatically evaluate AI-assisted and agentic workflows where they can improve infrastructure operations, service ownership, developer workflows, or toil reduction
Requirements
  • 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles
  • 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers
  • 3+ years of experience with modern cloud infrastructure in AWS, GCP, or similar environments
  • 3+ years of experience with Kubernetes, container orchestration, infrastructure-as-code, declarative systems, CI/CD, and deployment safety
Bonus Skills
  • Experience leading SRE, infrastructure, or platform work in a lean, high-agency organisation
  • Experience supporting globally distributed teams or 24/7 operational coverage
  • Experience improving developer productivity through paved paths, self-service infrastructure, automation, and reduced toil
  • Experience with infrastructure security fundamentals, secrets management, access controls, cloud security practices, or compliance-related infrastructure controls
  • Experience in financial services, regulated environments, blockchain, crypto, or other high-reliability technical ecosystems
  • Experience evaluating vendors and infrastructure platforms with scepticism, technical rigour, and cost discipline
  • Practical experience applying AI-assisted or agentic workflows to infrastructure, reliability, operations, observability, or developer productivity
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director of Site Reliability Engineering
Director of Site Reliability Engineering

techchaintalent • San Francisco (CA)

On-site
USD 190,000 - 270,000
Director Site Reliability Engineering
Director Site Reliability Engineering

TechChain Talent • New York (NY)

On-site
USD 180,000 - 250,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Stellar • San Francisco (CA)

On-site
USD 180,000 - 260,000
Health coverage
Flexible time off
Parental leave
+5
Director, Site Reliability Engineering
Director, Site Reliability Engineering

P2P • New York (NY)

On-site
USD 205,000 - 305,000
Health insurance
Flexible time off
Gym reimbursement
+2
Director, Site Reliability Engineering
Director, Site Reliability Engineering

P2P • San Francisco (CA)

Hybrid
USD 205,000 - 305,000
Comprehensive health coverage
Flexible time off
Generous paid parental leave
+3
Director of SRE & Platform Reliability
Director of SRE & Platform Reliability

Stellar • San Francisco (CA)

On-site
USD 180,000 - 260,000
Health coverage
Flexible time off
Parental leave
+5
Director of SRE & Platform Reliability
Director of SRE & Platform Reliability

P2P • New York (NY)

Hybrid
USD 205,000 - 305,000
Health insurance
Flexible time off
Gym reimbursement
+2
Director of SRE & Platform Reliability
Director of SRE & Platform Reliability

P2P • San Francisco (CA)

Hybrid
USD 205,000 - 305,000
Comprehensive health coverage
Flexible time off
Generous paid parental leave
+3
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000