Senior Site Reliability Engineer for Global, Scalable Systems

United States Digital Space LLC

New York (NY)

On-site

USD 183,000 - 247,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking a Senior Site Reliability Engineer to partner with product and platform teams to keep large distributed systems reliable and scalable. You will design, implement, and operate high-availability services with a focus on automation, incident response, and performance engineering.

The role emphasizes collaboration across engineering teams, hands-on debugging, and driving improvements to reduce toil while delivering resilient, scalable infrastructure for

Qualifications

  • 5+ years in site reliability engineering/DevOps for a product with millions of users.
  • Experience diagnosing issues in large-scale distributed systems.
  • Proficiency in Java, Kotlin, Python or Go.

Responsibilities

  • Collaborate with internal teams to identify sources of instability in distributed systems and drive operational excellence.
  • Support core infrastructure (i.e., understand, diagnose, and debug these systems in production).
  • Provide system design consulting, develop software platforms/frameworks, and conduct launch reviews and root cause analysis.
  • Maintain and document sustainable postmortem/incident response practices.
  • Advocate for and implement changes that improve reliability, scalability, and velocity.
  • Reduce the burden of toil with iterative development of tooling and automation.
  • Collaborate with engineering teams to release new features and become an authority on our services.

Skills

DevOps
Distributed systems
Java
Kotlin
Python
Go
Container orchestration
Incident response
Automation tooling
Databases (DynamoDB/MySQL/PostgreSQL)
Docker
Mesos
Kubernetes
Nomad

Tools

Docker
Mesos
Kubernetes
Nomad

Job description

United States Digital Space LLC is seeking a Senior Site Reliability Engineer to partner with product and platform teams to keep large distributed systems reliable and scalable. You will design, implement, and operate high-availability services with a focus on automation, incident response, and performance engineering.

The role emphasizes collaboration across engineering teams, hands-on debugging, and driving improvements to reduce toil while delivering resilient, scalable infrastructure for

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer for Global Education
Senior Site Reliability Engineer for Global Education

United States Digital Space LLC • Pittsburgh

On-site
USD 183,000 - 247,000
Equity compensation
Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior Network Reliability Engineer
Senior Network Reliability Engineer

United States Digital Space LLC • United States

Hybrid
USD 76,000 - 105,000
Senior Database Reliability Engineer - Scale & Safeguards
Senior Database Reliability Engineer - Scale & Safeguards

United States Digital Space LLC • United States

Remote
USD 64,000 - 106,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
SRE - Manufacturing Infra for Scalable Reliability
SRE - Manufacturing Infra for Scalable Reliability

United States Digital Space LLC • Bastrop (TX)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer - Remote
Senior Site Reliability Engineer - Remote

Bright-Vision-Technologies • United States

Remote
USD 100,000 - 150,000
Senior Staff Engineer, Global Support Platform (Remote)
Senior Staff Engineer, Global Support Platform (Remote)

United States Digital Space LLC • United States

Remote
USD 170,000 - 230,000
The latest technology and equipment
Potential to work remotely, includings
28 days vacation + up to 5 sick days
+3
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Remote Senior Site Reliability Engineer - Cloud & Automation
Remote Senior Site Reliability Engineer - Cloud & Automation

Multi Media LLC • United States

On-site
USD 169,000 - 215,000
Fully Remote
Health Insurance
Vision Insurance
+10