Director, Site Reliability Engineering

Stellar

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health coverage
Flexible time off
Parental leave
Gym reimbursement
401K
Commuter benefits
HSA
Company retreats

Job summary

Stellar Development Foundation (SDF) seeks a Director of Site Reliability Engineering to lead a distributed SRE team and shape how engineering owns, operates, and improves production services.

You will set vision, operating model, and culture for SRE while owning core infrastructure services like cloud foundations, Kubernetes, CI/CD, observability, and automation. Strong leadership and pragmatic judgment are essential.

Qualifications

  • 10+ years in SRE, infra, or related roles.
  • 5+ years leading or developing infrastructure/reliability engineers.
  • Deep technical judgment across cloud infra, distributed systems, risk, and automation.
  • 3+ years in modern cloud infra (AWS/GCP) and Kubernetes.

Responsibilities

  • Lead and develop a distributed SRE team with a clear vision and priorities.
  • Roll out a Service Ownership & Maturity Framework across engineering.
  • Own core infra services: cloud foundations, Kubernetes, CI/CD, observability, secrets.
  • Improve reliability, runbooks, dashboards, and on-call practices.
  • Mature incident response, postmortems, and cross‑org collaboration.

Skills

SRE leadership
Cloud infrastructure
Kubernetes
CI/CD
Observability
Incident response
Executive communication

Tools

GitHub workflows

Job description

Interested in working on cutting-edge blockchain technology and creating equitable access to the global financial system? Since 2014, the mission-driven team at the Stellar Development Foundation (SDF) has helped fuel the tremendous growth of the Stellar blockchain network, an open-source platform that operates at high-scale today. Developers and companies around the world build on it, and the SDF team is expanding to support the rapidly growing and changing Stellar ecosystem.

SDF is looking for a Director of Site Reliability Engineering to lead a small, high-leverage SRE team and help shape how engineering teams own, operate, and improve production services.

This is a senior engineering leadership role reporting to the CTO. You will set the vision, operating model, and culture for SRE while owning the core infrastructure services that help SDF engineering teams build, deploy, observe, and operate software with confidence.

Engineering teams at SDF own the services they build. SRE provides the frameworks, standards, shared infrastructure, tooling, observability practices, and enablement model that make strong service ownership possible across engineering.

You will be successful here if you bring strong technical judgment, pragmatic leadership, and the ability to influence through trust, clarity, and execution. SDF is a small, mission-driven foundation with a broad technical surface area, so this role requires leverage, ownership, and a bias toward solving the right problems over creating processes for its own sake.

In this role, you will:
  • Lead, coach, and develop a distributed SRE team, setting a clear vision, charter, operating model, priorities, and success measures.
  • Define and roll out a Service Ownership & Maturity Framework across engineering, with expectations that vary appropriately by service criticality.
  • Own and improve core engineering infrastructure services, including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
  • Help engineering teams become stronger owners and operators of their services through better standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices.
  • Make reliability, operational maturity, infrastructure health, and developer productivity more measurable through trusted metrics and practical operational intelligence.
  • Improve deployment automation, resilience, self‑healing patterns, disaster recovery readiness, and service reliability based on actual impact and risk.
  • Mature incident response, escalation, postmortems, and on‑call health across a geographically distributed team.
  • Build paved paths and self‑service infrastructure that reduce toil, lower cognitive load, and help engineering teams move faster while strengthening ownership and reliability.
  • Partner closely with Security, Compliance, Legal, Finance, Procurement, and Corporate IT where infrastructure, access management, cloud operations, vendor review, or controls intersect with engineering.
  • Pragmatically evaluate AI‑assisted and agentic workflows where they can improve infrastructure operations, service ownership, developer workflows, or toil reduction.
You have:
  • 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles.
  • 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers.
  • Strong experience defining team charters, operating models, roadmaps, success measures, and engineering practices for infrastructure or reliability teams.
  • Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability tradeoffs, automation, and operational risk.
  • 3+ years of experience with modern cloud infrastructure in AWS, GCP, or similar environments.
  • 3+ years of experience with Kubernetes, container orchestration, infrastructure‑as‑code, declarative systems, CI/CD, and deployment safety.
  • Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on‑call practices.
  • Experience helping product or application engineering teams improve service ownership, operational readiness, and production accountability.
  • A pragmatic approach to tooling: you understand when to build, buy, adapt, simplify, or retire systems based on the actual engineering problem.
  • The ability to operate effectively in a small or mid‑size engineering organization where influence comes from credibility, judgment, and outcomes rather than bureaucracy.
  • Clear executive communication skills and the ability to partner directly with a CTO and senior engineering leaders.
Bonus Points if:
  • Experience leading SRE, infrastructure, or platform work in a lean, high‑agency organization.
  • Experience supporting globally distributed teams or 24/7 operational coverage.
  • Experience improving developer productivity through paved paths, self‑service infrastructure, automation, and reduced toil.
  • Experience with infrastructure security fundamentals, secrets management, access controls, cloud security practices, or compliance‑related infrastructure controls.
  • Experience in financial services, regulated environments, blockchain, crypto, Web3, or other high‑reliability technical ecosystems.
  • Experience evaluating vendors and infrastructure platforms with skepticism, technical rigor, and cost discipline.
  • Practical experience applying AI‑assisted or agentic workflows to infrastructure, reliability, operations, observability, or developer productivity.
USA Benefits/Perks:
  • Competitive health, dental & vision coverage with most plans covered at 100% for the employee + any dependents.
  • Flexible time off + 15 company holidays including a company‑wide holiday break.
  • Generous paid parental leave for all parents, plus paid pregnancy disability leave for birthing parents.
  • Gym reimbursement ($80 per month).
  • Life & AD&D (up to $50K).
  • Short & Long term disability.
  • 401K with 4% match.
  • Health & Dependent Care FSA Accounts.
  • Commuter benefits with $250/month employer contribution.
  • Health Savings Account (HSA) with monthly employer contribution.
  • Family building benefits through Kindbody.
  • Wellbeing benefits (One Medical, Rightway, Headspace).
  • L&D budget of $1,500/year.
  • Daily lunch and snacks in office.
  • Company retreats.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director, Site Reliability Engineering
Director, Site Reliability Engineering

Stellar • New York (NY)

On-site
USD 180,000 - 240,000
Health, dental, vision insurance
Flexible time off
Parental leave
+5
Director, Site Reliability Engineering
Director, Site Reliability Engineering

P2P • New York (NY)

Hybrid
USD 205,000 - 305,000
Health insurance
Flexible time off
Gym reimbursement
+2
Director, Site Reliability Engineering
Director, Site Reliability Engineering

P2P • San Francisco (CA)

Hybrid
USD 205,000 - 305,000
Comprehensive health coverage
Flexible time off
Generous paid parental leave
+3
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Stellar Development Foundation • United States

Hybrid
USD 205,000 - 305,000
Competitive health, dental & vision coverage
Flexible time off + 15 company holidays
Generous paid parental leave
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

techchaintalent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

TechChain Talent • San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

TechChain Talent • New York (NY)

On-site
USD 120,000 - 150,000
Director of SRE & Platform Reliability
Director of SRE & Platform Reliability

Stellar • San Francisco (CA)

On-site
USD 180,000 - 260,000
Health coverage
Flexible time off
Parental leave
+5
Director of SRE & Platform Reliability
Director of SRE & Platform Reliability

P2P • San Francisco (CA)

Hybrid
USD 205,000 - 305,000
Comprehensive health coverage
Flexible time off
Generous paid parental leave
+3
Senior Software Engineer, Core
Senior Software Engineer, Core

Stellar Development Foundation • New York (NY)

On-site
USD 180,000 - 290,000
Health, dental & vision coverage
Flexible time off
Paid parental leave
+6