Senior Site Reliability Engineer

Block

San Francisco (CA)

On-site

USD 160,700 - 283,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare coverage
Health Savings Account
Retirement Plans
Employee Stock Purchase Program
Wellness programs
Paid parental leave
Paid time off
Learning and Development resources

Job summary

Block in San Francisco is looking for a Site Reliability Engineer who will enhance platform reliability and support critical services. You will work in a dynamic environment utilizing AI-driven tooling to improve system observability and incident response, ensuring that our products remain stable and accessible to all.

This role involves participating in primary oncall rotations, managing high-severity incidents, and driving overall reliability improvements across the organization. Candidates should possess strong incident management skills and 5+ years of software development experience.

Qualifications

  • 5+ years of software development experience.
  • Experience running production oncall for high-availability systems.
  • Familiarity with AI-driven tooling for observability, incident analysis, or automation.

Responsibilities

  • Build and extend platforms to improve system reliability.
  • Standardize reliability tools across multiple platforms.
  • Lead stabilization of sev 0–1 incidents.

Skills

AI-driven tooling
Incident management
CI/CD pipelines
Monitoring & observability
Software development

Tools

Kotlin
Amazon Web Services
Terraform
DataDog
Kubernetes

Job description

Block builds simple, powerful tools that make progress towards an economy that’s truly open to all.

Each of our brands unlocks different aspects of the economy for more people. Square makes commerce and financial services accessible to sellers. Cash App is the easy way to spend, send, and store money. Afterpay is transforming the way customers manage their spending over time. TIDAL is a music platform that empowers artists to thrive as entrepreneurs. Bitkey is a simple self-custody wallet built for bitcoin. Proto is a suite of bitcoin mining products and services. Together, we’re helping build a financial system that is open to everyone.

The Role

As a member of the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics-driven, systems-oriented, and focused on building distributed platforms that enable safe, scalable product development.

You will leverage and continuously improve AI-driven tooling and automation to enhance observability, accelerate incident detection and response, and reduce operational toil. This includes applying AI to incident analysis, alert tuning, and operational workflows.

You will participate in primary platform oncall (12 hours per day, one week every few weeks, depending on team size), supporting Block's most critical (Tier 0) services. In this role, you will lead incident command, coordinate mitigation, and drive effective escalation during high-severity events.

You Will
  • Build and extend platforms to improve system reliability
  • Work on team goals that encompass reliability for the entire company
  • Standardize reliability tools across multiple platforms and organizations
  • Triage, coordinate, and lead stabilization of sev 0–1 incidents
  • Serve as primary oncall, maintaining structured escalation paths and exercising leadership escalation
  • Drive platform-wide reliability improvements, shared operational tooling, and deploy‑safety patterns
  • Use AI‑driven systems to improve signal detection, reduce noise, and accelerate root cause analysis
  • Design and implement safe deployment patterns (progressive delivery, automated rollback, guardrails)
You Have
  • Drive to root cause systems with many moving parts and take the necessary steps to fix them
  • Demonstrated technical initiative and leadership on previous projects, especially those with a backend/platform focus
  • Familiarity with AI‑driven tooling for observability, incident analysis, or automation
  • A mindset that naturally reaches for AI to accelerate problem‑solving and reduce toil
  • Experience running production oncall for high‑availability systems
  • Strong incident management skills — structured triage, mitigation under pressure, blameless postmortems
  • Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation
  • Monitoring & observability expertise — building/tuning alerts for uptime, error rates, latency regression, and resource exhaustion
  • Ability to create and maintain evidence‑based maturity assessments using trailing 90‑day data windows
  • Comfort with vendor/dependency management — maintaining validated escalation contacts reachable within 5 minutes
  • Boundless curiosity, autonomy, and a strong sense of accountability
  • A strong desire to perform and grow as an engineer
  • 5+ years of software development experience
Technologies We Use and Teach
  • Kotlin, Modern Java (11+)
  • HTTP, JSON, gRPC, and Protocol Buffers
  • MySQL / Vitess / DynamoDB
  • Event driven architectures
  • DataDog
  • LaunchDarkly
  • Terraform, Kubernetes, Istio/Envoy
  • Amazon Web Services

This program shifts Block from reactive incident handling to repeatable, system‑wide reliability gains — fewer customer‑visible incidents, faster response, higher product velocity, and lower burnout across the organization.

We’re working to build a more inclusive economy where our customers have equal access to opportunity, and we strive to live by these same values in building our workplace.

Block is a proud equal opportunity employer. We work hard to evaluate all employees and job applicants consistently, based solely on the core competencies required of the role at hand, and without regard to any legally protected class. We believe in being fair, and are committed to an inclusive interview experience, including providing reasonable accommodations to disabled applicants throughout the recruitment process. We encourage applicants to share any needed accommodations with their recruiter, who will treat these requests as confidentially as possible.

Full‑time employee benefits
  • Healthcare coverage (Medical, Vision and Dental insurance)
  • Health Savings Account and Flexible Spending Account
  • Retirement Plans including company match
  • Employee Stock Purchase Program
  • Wellness programs, including access to mental health, 1:1 financial planners, and a monthly wellness allowance
  • Paid parental and caregiving leave
  • Paid time off (including 12 paid holidays)
  • Paid sick leave (1 hour per 26 hours worked (max 80 hours per calendar year to the extent legally permissible) for non‑exempt employees and covered by our Flexible Time Off policy for exempt employees)
  • Learning and Development resources
  • Paid Life insurance, AD&D, and disability benefits

These benefits are further detailed in Block's policies. This role is also eligible to participate in Block's equity plan subject to the terms of the applicable plans and policies, and may be eligible for a sign‑on bonus. Sales roles may be eligible to participate in a commission plan subject to the terms of the applicable plans and policies. Pay and benefits are subject to change at any time, consistent with the terms of any applicable compensation or benefit plans.

Block takes a market‑based approach to pay, and pay may vary depending on your location. U.S. locations are categorized into one of four zones based on a cost of labor index for that geographic area. The successful candidate's starting pay will be determined based on job‑related skills, experience, qualifications, work location, and market conditions. These ranges may be modified in the future.

Zone A: USD $189,000 - USD $283,600

Zone B: USD $179,600 - USD $269,400

Zone C: USD $170,100 - USD $255,100

Zone D: USD $160,700 - USD $241,100

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Block • New York (NY)

On-site
USD 170,000 - 284,000
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
Software Engineer, Finance Applications
Software Engineer, Finance Applications

Block • San Francisco (CA)

On-site
USD 218,000 - 327,000
Sr. People Tech Analyst
Sr. People Tech Analyst

Block • San Francisco (CA)

Hybrid
USD 159,000 - 258,000
Remote work
Medical insurance
Flexible time off
+2
Manager, Mid-Market Sales
Manager, Mid-Market Sales

Socket.dev • Los Angeles (CA)

On-site
USD 234,000 - 351,000
Healthcare coverage
401(k) retirement plan
Employee stock purchase program
+6
Senior Software Engineer, Data Enablement
Senior Software Engineer, Data Enablement

Block • San Francisco (CA)

Hybrid
USD 217,800 - 326,800
Remote work
Medical insurance
Flexible time off
+1
Manager, Mid-Market Sales
Manager, Mid-Market Sales

Square • Los Angeles (CA)

On-site
USD 252,000 - 377,000
Healthcare coverage
Health Savings Account
Retirement Plans with company match
+7
Staff Applied Machine Learning Engineer - Intelligent Data, Signals & Systems
Staff Applied Machine Learning Engineer - Intelligent Data, Signals & Systems

Block • San Francisco (CA)

On-site
USD 276,000 - 416,000
Remote work
Medical insurance
Flexible time off
+2
Principal Security Engineer
Principal Security Engineer

Block • Austin (TX)

On-site
USD 319,000 - 479,000
Remote work
Flexible time off
Medical insurance
+1
Senior Software Engineer, Ledgering
Senior Software Engineer, Ledgering

Block • Oakland (CA)

On-site
USD 217,000 - 327,000
Remote work
Medical insurance
Flexible time off
+2
Senior Software Engineer, Ledgering
Senior Software Engineer, Ledgering

Block • San Francisco (CA)

Remote
USD 217,800 - 326,800
Remote work
Medical insurance
Flexible time off
+2