Lead Site Reliability Engineer

United States Digital Space LLC

New York (NY)

Hybrid

USD 179,000 - 226,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Unlimited PTO
Employee stock options
Medical, dental, vision with HSA
401k with match
Parental leave
Home office stipend
Learning & Development stipend
Well-being benefits
Hybrid work in Union Square

Job summary

United States Digital Space LLC is seeking an experienced Infrastructure/SRE engineer to design and automate large-scale infrastructure. You will own provisioning, migrations, and reliable, self-service platforms for other engineers across Kubernetes clusters, databases, and services.

You will reduce toil, implement scalable deployment tooling, and contribute to architecture decisions around reliability and security. Strong coding, on-call experience, and collaboration are essential.

Qualifications

  • 10+ years of experience in infrastructure, SRE, or software engineering roles.
  • Strong software engineering skills and ability to build systems, not just scripts.
  • Experience managing production infrastructure at scale (cloud + containerized systems).
  • Experience with Infrastructure as Code (e.g., Terraform).
  • Experience running and troubleshooting distributed systems (Docker/Kubernetes).
  • Experience with observability/debugging tools (Datadog, CloudWatch, ELK/EFK).
  • Proficiency in at least one programming language (Python, Go, JavaScript, etc.).
  • Experience in on-call rotations and improving systems based on incidents.
  • Strong communication and collaboration skills.

Responsibilities

  • Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations)
  • Reduce operational toil by turning manual processes into reliable, repeatable workflows
  • Build internal tooling and platforms that enable safe self-service changes for other engineers
  • Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services)
  • Implement and evolve systems for deploying and running applications in Kubernetes
  • Contribute to architecture decisions across infrastructure, reliability, and security
  • Write and review production-quality code
  • Participate in on-call rotations while building systems to prevent incidents

Skills

Terraform
Docker
Kubernetes
Datadog
CloudWatch
ELK/EFK
Python
Go
JavaScript
Observability
Automation
SRE
On-call rotations

Tools

Terraform
Docker
Kubernetes
Datadog
CloudWatch
ELK/EFK
AWS

Job description

the company is where you belong!

the company helps solve the identity risk problem for companies that offer financial products by enabling them to outpace fraud and confidently serve more people around the world. Over 800 of the world's largest financial institutions and fintechs turn to the company to take control of fraud, credit, and compliance risk, and grow with the clearest picture of their customers.

Through our values: Be Bold, Get Scrappy, Collaborate, and Celebrate Our Differences, we are creating a workplace where you can grow, thrive, and belong. See how we've been continuously recognized and named one ofInc. Magazine's Best Workplaces,Forbes America's Best Startup Employers,Best Fintech to Work for by American Banker, year after year.

Check out our investors and read more about us here.

About the team

the company's Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure footprint: 15+ Kubernetes clusters, 100+ databases, dozens of services, and complex data organization.

Our challenge isn't just scale—it’s making that scale reliable, secure, and operable with less manual work.

We're looking for engineers who enjoy turning complex, fragile systems into automated, self-service platforms with strong safety guarantees.

What you'll be doing

Reporting to the Engineering Manager of Infrastructure, you'll:

  • Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations)
  • Reduce operational toil by turning manual processes into reliable, repeatable workflows
  • Build internal tooling and platforms that enable safe self-service changes for other engineers
  • Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services)
  • Implement and evolve systems for deploying and running applications in Kubernetes
  • Contribute to architecture decisions across infrastructure, reliability, and security
  • Write and review production-quality code
  • Participate in on-call rotations—but focus on building systems that prevent incidents, not just respond to them
Who we're looking for
  • 10+ years of experience in infrastructure, SRE, or software engineering roles
  • Strong software engineering skills—you build systems, not just scripts
  • Experience managing production infrastructure at scale (cloud + containerized systems)
  • Experience with Infrastructure as Code (e.g., Terraform)
  • Experience running and troubleshooting distributed systems (Docker/Kubernetes)
  • Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.)
  • Proficiency in at least one programming language (Python, Go, JavaScript, etc.)
  • Experience participating in on-call rotations and improving systems based on incidents
  • Strong communication and collaboration skills
You might be a great fit if you
  • Default to automation over manual processes
  • See repetitive work and immediately want to eliminate it
  • Think in terms of systems, failure modes, and long-term scalability
  • Care about building infrastructure that other engineers can use safely and confidently
  • Enjoy working in a small team with high ownership and impact
Nice to have
  • Experience running Kubernetes in production at scale
  • Deep familiarity with AWS
  • Experience building internal platforms or developer tooling
  • Background in distributed systems or large-scale data systems

We're a lean team, so your impact will be felt immediately, and opportunities will grow as the company scales up. If this all sounds like a good fit for you, why not join us?

the company is committed to fair and equitable compensation practices. Below is the anticipated starting base compensation range for this role; however, pay may vary depending on job-related knowledge, in-demand skills, relevant experience, and/or geography. In addition to a competitive base salary, this position is also eligible for equity awards in the form of stock options (ISOs) as well as a competitive total benefits package.

*This position has a salary range of $179,000 to $226,000.*

Benefits and Perks
  • Unlimited PTO and flexible work policy
  • Employee stock options
  • Medical, dental, vision plans with HSA (monthly employer contribution) and FSA options
  • 401k with 100% match up to 4% of annual employee compensation
  • Eligible new parents receive 16 weeks of paid parental leave
  • Home office stipend for new employees
  • Annual Learning & Development annual stipend
  • Well-being benefits include access to ClassPass, OneMedical, UrbanSitter, and Spring Health
  • Hybrid work environment: employees are expected to work Tuesdays through Thursdays from our HQ in Union Square, Manhattan. Tasty lunches catered from a variety of local restaurants and frequent employee-organized cultural events contribute to our positive office energy. On Monday/Friday most employees Zoom into work from home while some take advantage of the quieter office.

the company is proud to be an equal-opportunity workplace and employer. We're committed to equal opportunity regardless of race, color, ancestry, religion, gender, gender identity, parental or pregnancy status, national origin, sexual orientation, age, citizenship, marital status, disability, or veteran status. We are committed to an inclusive interview experience and provide reasonable accommodations to applicants with visible and invisible disabilities. We encourage applicants to share needed accommodations with your recruiter.

All the company jobs are listed on our careers page. Any communication during the recruitment process, including interview requests or job offers, will come directly from a recruiting team member with an the company.com email address. We do not use outside applications or automated text messaging in our recruiting process. We will not ask for any sensitive financial or identification information during the recruiting process. If you're ever unsure, please contact us directly via our website before sharing personal information.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer (FedRAMP)
Staff Site Reliability Engineer (FedRAMP)

United States Digital Space LLC • Washington

On-site
USD 174,000 - 239,000
Health insurance
Dental insurance
Vision insurance
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Bellevue (CA)

On-site
USD 180,000 - 230,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion
Senior Manager, Site Reliability Engineering - Infrastructure Platform
Senior Manager, Site Reliability Engineering - Infrastructure Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 232,000 - 319,000
Equity
Bonus
Health insurance
+2
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Capitolis • New York (NY)

Hybrid
USD 151,000 - 191,000
Unlimited PTO
Employee stock options
401k with 100% match
+4
Staff TDI Site Reliability Engineer, Okta Federal
Staff TDI Site Reliability Engineer, Okta Federal

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 174,000 - 239,000
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 204,000 - 306,000
Equity
Bonus
Health insurance
+6
Staff Software Engineer, Core Infrastructure
Staff Software Engineer, Core Infrastructure

United States Digital Space LLC • San Francisco (CA)

On-site
USD 194,000 - 267,000
Equity
Bonus
Health insurance
+6
Senior Software Engineer II, Frontend Platform
Senior Software Engineer II, Frontend Platform

United States Digital Space LLC • New York (NY)

Hybrid
USD 198,000 - 250,000
Unlimited PTO
Employee stock options
Health, dental, vision with HSA and FS
+7
Founding Forward Deployed Engineer
Founding Forward Deployed Engineer

United States Digital Space LLC • New York (NY)

Hybrid
USD 166,000 - 250,000
Unlimited PTO
Stock options
Health plans
+6