Senior Site Reliability Engineer

United States Digital Space LLC

New York (NY)

On-site

USD 183,000 - 247,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking a Senior Site Reliability Engineer to partner with product and platform teams to keep large distributed systems reliable and scalable. You will design, implement, and operate high-availability services with a focus on automation, incident response, and performance engineering.

The role emphasizes collaboration across engineering teams, hands-on debugging, and driving improvements to reduce toil while delivering resilient, scalable infrastructure for

Qualifications

  • 5+ years in site reliability engineering/DevOps for a product with millions of users.
  • Experience diagnosing issues in large-scale distributed systems.
  • Proficiency in Java, Kotlin, Python or Go.

Responsibilities

  • Collaborate with internal teams to identify sources of instability in distributed systems and drive operational excellence.
  • Support core infrastructure (i.e., understand, diagnose, and debug these systems in production).
  • Provide system design consulting, develop software platforms/frameworks, and conduct launch reviews and root cause analysis.
  • Maintain and document sustainable postmortem/incident response practices.
  • Advocate for and implement changes that improve reliability, scalability, and velocity.
  • Reduce the burden of toil with iterative development of tooling and automation.
  • Collaborate with engineering teams to release new features and become an authority on our services.

Skills

DevOps
Distributed systems
Java
Kotlin
Python
Go
Container orchestration
Incident response
Automation tooling
Databases (DynamoDB/MySQL/PostgreSQL)
Docker
Mesos
Kubernetes
Nomad

Tools

Docker
Mesos
Kubernetes
Nomad

Job description

Our mission at the company is to develop the best education in the world and make it universally available. It’s a big mission, and that’s where you come in!

At the company, you’ll join a team that cares about finding innovative solutions to complex technical problems, running countless experiments (300+ at a time!) with our massive user base to make data-driven decisions, and educating our users and employees alike. You’ll have limitless learning opportunities, mentorship and collaboration with world-class minds, and a variety of projects with large scopes — while doing work that’s both fun and meaningful.

Join our life-changing mission to develop education for our half a billion (and growing!) learners around the world.

About the role...

As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure the company’s sophisticated distributed systems and products are built and maintained with extraordinary quality, and operated in measurable and scalable ways.

You will...
  • Collaborate with internal teams to identify sources of instability in distributed systems and drive operational excellence
  • Support core infrastructure (i.e understand, diagnose, and debug these systems in production)
  • Provide system design consulting, develop software platforms/frameworks, and conduct launch reviews and root cause analysis
  • Maintain and document sustainable postmortem/incident response practices
  • Advocate for and implement changes that improve reliability, scalability, and velocity
  • Reduce the burden of toil with iterative development of tooling and automation
  • Collaborate with engineering teams to release new features and become an authority on our services
You have...
  • 5+ years of experience within site reliability engineering/DevOps of a product with millions of users
  • Experience identifying and solving issues in large-scale distributed systems
  • Experience with Java, Kotlin, Python or Go
  • An understanding of containerization toolsets and container orchestration technologies (Docker, Mesos, Kubernetes, Nomad, etc)
Exceptional candidates will have...
  • Experience in improving automation and tooling to reduce service maintenance toil
  • Proven experience driving improvements to incident response processes
  • Experience assessing reliability and troubleshooting issues in Dynamo, MySQL, and/or PostgreSQL databases

*The offered salary is dependent upon several factors, including work experience, skills, and internal peer comparisons. The posted range is subject to change in the future. For this role, base salary is supplemented by equity compensation. We encourage you to talk with your recruiter for more information related to compensation for this role!*

Salary Range:

$182,800—$247,300 USD

Accommodations:

We will do everything we can within reason to make sure that your interview takes place in an environment that fairly and accurately assesses your skills. If you need assistance or accommodation, please contact hr@unitedstatesdigital.space.

Equal Employment Opportunity:

the company is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics.

Fraud Warning:

Unfortunately, there is a rise in scammers pretending to be real the company employees. the company and our employees will never ask for your Social Security number, bank details, or passport info, and we’ll never ask you to deposit a check, purchase equipment, or exchange money during the interview process. Real the company employees always use an email that ends in @the company.com or @recruiting.the company.com. Stay alert and double-check these details before sharing any information.

By applying for this position your data will be processed as per the the company Applicant Privacy Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Pittsburgh

On-site
USD 183,000 - 247,000
Equity compensation
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Socket.dev • Pittsburgh, Northern (KY)

Hybrid
USD 183,000 - 247,000
Software Engineer
Software Engineer

United States Digital Space LLC • United States

On-site
USD 127,000 - 176,000
Learning allowance
Equity compensation
Medical, dental, vision coverage
+3
Manager, Software Engineering - Data Platform
Manager, Software Engineering - Data Platform

United States Digital Space LLC • United States

Remote
USD 258,000 - 376,000
Equity compensation
Health, dental, and vision coverage
Retirement benefits with company contr
Site Reliability Engineer I
Site Reliability Engineer I

Talanto • Atlanta (GA), Northern (KY)

Hybrid
USD 98,000 - 149,000
Company equity
ESPP (Employee Stock Purchase Program)
Generous paid vacation time
+2
Sales Engineer - HRIS/ATS
Sales Engineer - HRIS/ATS

United States Digital Space LLC • Denver (CO)

On-site
USD 116,000 - 136,000
Medical coverage
Dental coverage
Vision coverage
+6
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Site Reliability Engineer Intern — Summer 2027
Site Reliability Engineer Intern — Summer 2027

Talanto • San Francisco (CA), Northern (KY)

Hybrid
USD 62,000 - 69,000
Senior Engineering Manager, Site Reliability
Senior Engineering Manager, Site Reliability

Horizon3.ai • United States

On-site
USD 260,000 - 280,000
Hybrid & Remote Work
Competitive Compensation
Equity package
Senior Enterprise Engineer
Senior Enterprise Engineer

Duolingo • Pittsburgh

On-site
USD 162,000 - 220,000