Senior Staff Site Reliability Engineer

United States Digital Space LLC

Michigan

Hybrid

USD 150,000 - 190,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking a high-performance Site Reliability Engineer in Oakland, CA, with hybrid work options. You’ll own reliability of the data platform, oversee deployment pipelines, and drive incident response with cross‑functional teams to ensure scalable, resilient infrastructure.

You will partner with engineering, product, support, and sales to maintain 100% availability and rapid incident resolution, leveraging cloud networking and automation to improve the data stack.

Qualifications

  • Experienced Site Reliability Engineer focused on reliability and performance of production infrastructure.
  • Ability to work with engineering, product managers, support, and sales engineers to build reliable systems.
  • Strong knowledge of cloud platforms and container orchestration in production environments.

Responsibilities

  • Monitor availability, capacity, and throughput of production systems.
  • Suggest improvements to product roadmap to improve reliability.
  • Coordinate bug fixes for critical issues for support or sales needs.
  • Advise on production infrastructure to achieve 100% availability.
  • Ensure deployment artifacts are scalable across environments via automation.
  • Collaborate with security to remediate infrastructure vulnerabilities.

Skills

Kubernetes
Cloud platforms
Python
Linux
Cloud networking
PostgreSQL

Tools

Terraform
Ansible
Buildkite
Pulumi
ArgoCD

Job description

From the company's founding until now, our mission has remained the same: to make access to data as simple and reliable as electricity. With the company, customer data arrives in their warehouses, canonical and ready to query, with no engineering or maintenance required. We're proud that more organizations continue to leverage our technology every day to become truly data-driven.

About Us

the company and dbt Labs are bringing together two industry-leading companies with a shared mission: helping organizations unlock the full value of their data.

Together, we're delivering the data infrastructure layer that helps organizations move, transform, and trust their data - from the moment data moves, through every transformation, to the context teams and AI systems rely on.

the company helps organizations automate data movement across the systems, clouds, engines, and tools they rely on. dbt Labs pioneered analytics engineering, helping teams transform data into reliable, governed insights. Together, we support thousands of organizations as they build a trusted foundation for analytics, AI, and better business decisions.

As we bring our teams and technology together, we're building on the strengths of both companies while continuing to deliver the products and experiences our customers know and trust. It's an exciting time to join us: we're creating a company with the scale, talent, and technology to help more organizations put their data to work with greater speed, confidence, and impact.

During this transition period, you may see references to both the company and dbt Labs throughout our recruiting process as we integrate our teams, systems, and career sites.

About the Role

the company and dbt Labs is building data pipelines to power the modern data stack for thousands of companies.

the company is looking for a high-performance, experienced engineer to be a part of a team of Site Reliability Engineers. You will be working closely with engineering teams, product managers, as well as support and sales engineers to build the future of the the company Data Platform Reliability.

As a member of the Site Reliability Engineering team, you will take ownership of the overall performance and reliability of the company's infrastructure, the robustness of the deployment pipeline, as well as timely and effective incident response and resolution. You will take responsibility for the growth and stability of the company's infrastructure, and be a key player driving effective incident response and overall issue avoidance.

This is a full-time position based out of ourOakland, CA office. Our hybrid work model offers a blend of remote flexibility and in-person collaboration, including two days in the office each week to connect and build as a team.

What You'll Do
  • Responsible for ongoing reliability and robustness of the company's production infrastructure by monitoring availability, capacity, and throughput.
  • Evolve systems by adding reliability into our product roadmap
  • Coordinate the re-prioritize or fix critical bugs for support or sales requirements as needed
  • Make recommendations to production infrastructure by interfacing with engineering to ensure 100% availability
  • Ensure scalable artifacts deployment to all environments by automation scripts
  • Constantly monitor infrastructure vulnerabilities and remedy them by working with the security team
Skills We're Looking For
  • Expertise in managed Kubernetes (EKS, AKS and GKE)
  • Expertise of Cloud Platforms and related tooling: AWS, Azure, GCP, Terraform, Ansible, Buildkite, Pulumi and ArgoCD
  • Expertise in Python Programming.
  • Expertise with Linux operating systems internals and administration
  • Expertise with cloud networking like VPNs, Privatelinks, and Private Service connect (GCP)
  • Experience with databases such as PostgreSQL
Bonus Skills
  • Bonus if you have Java, GO Programing

#LI-HYBRID #LI-EM1

The compensation range displayed on this job posting reflects the minimum and maximum target for new hire compensation for the target position and level, and may include sales incentives or target bonuses depending on the role. Our compensation ranges are dete

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

FORT • United States

Hybrid
USD 150,000 - 180,000
Healthcare benefits
Flexible work environment
Large-scale cloud platform project
+1
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

Fivetran • Michigan

Hybrid
USD 232,000 - 290,000
Employer-paid medical insurance
Generous PTO and parental leave
RSU stock grants
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LeanData Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Jobot • Akron (OH)

Remote
USD 100,000 - 150,000
Comprehensive health insurance
Vision insurance
Dental insurance
+3
Site Reliability Engineer
Site Reliability Engineer

Skill • Southlake (TX)

On-site
USD 66,000 - 73,000
Health insurance
Vision insurance
Dental insurance
+2
Senior Site Reliability Engineer, Forward Deployed - Remote USA ONLY
Senior Site Reliability Engineer, Forward Deployed - Remote USA ONLY

Ardan Labs • United States

Remote
USD 140,000 - 210,000
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
Senior Site Reliability Engineer - Data Infrastructure (San Jose)

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600