Senior Site Reliability Engineer

Umbra

Arlington (TX)

On-site

USD 150,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible Time Off
Sick Leave (family & medical)
Medical, Dental, Vision, Life, LTD,STD
Pet Insurance (employee funded)
401k with 3% company contribution
Stock Options
Free Parking
Free lunch in office

Job summary

Umbra, an American space technology company, is seeking a Senior Site Reliability Engineer to design, build, operate, and scale mission-critical infrastructure powering Umbra’s systems. You will work with engineering teams to improve processes, evaluate new technologies, and drive reliability and performance improvements across platforms.

The role requires deep expertise in distributed systems, cloud infrastructure, and automation, with a strong focus on reliability and scalability.

Qualifications

  • Bachelor’s degree in Computer Science or a related technical field.
  • 5-8+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform with distributed systems experience.
  • Extensive experience with AWS services (EC2, S3, Lambda, VPC) and cloud infrastructure best practices.
  • Proficiency running, optimizing, and scaling Kubernetes clusters in production.
  • Experience using Terraform to architect and manage production infrastructure.
  • IaC and GitOps practices to increase reliability and reduce manual tasks.
  • Agile/Scrum methodologies and leadership in projects or teams.
  • Experience in building monitoring and alerting strategies for large-scale systems.

Responsibilities

  • Ensure reliability and scalability of critical systems, meeting SLAs through proactive monitoring and incident response.
  • Develop and promote new technologies to enhance team capabilities with proofs of concept.
  • Lead by example in fostering a culture of excellence and reliability.
  • Continuously evaluate and improve processes to increase efficiency.
  • Collaborate with product managers and stakeholders to align on technical strategy and provide guidance.
  • Participate in on-call rotations and resolve complex technical issues.

Skills

Distributed systems
Site Reliability Engineering
Cloud infrastructure
GitOps
Agile/Scrum
Performance monitoring

Education

Bachelor's degree in Computer Science or related field

Tools

AWS
Kubernetes
Terraform

Job description

Umbra is an American space technology company delivering advanced systems, from sensors to spacecraft, that empower customers worldwide with unmatched access to critical information from space. Our mission is simple and ambitious: redefine space—for people, systems, and missions in every domain. Umbra’s ecosystem operates through three business units: Remote Sensing (the data), Space Systems (the components), and Mission Solutions (the platforms).Together, our teams develop capabilities that deliver persistent access, resilient performance, and mission-ready solutions, advancing U.S. space leadership while keeping the world safe and informed.

About the Team

Remote Sensing – The Data

Remote Sensing is where Umbra got its start, and our agile Synthetic Aperture Radar (SAR) constellation remains the most capable on the market. We transform satellite data into real-world, actionable insights that strengthen U.S. national security and intelligence, support disaster response, and advance scientific discovery. Our team delivers data at scale with unmatched quality, persistence, and the speed and responsiveness our partners demand.

If you want to work on cutting-edge space technology that’s redefining what’s possible in remote sensing, you belong here at Umbra.

About the Job

We are seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's systems. In this role, you will leverage a deep understanding of modern infrastructure, distributed systems, and the broader technology stack to drive technical excellence, make thoughtful architectural decisions, and balance long-term scalability with operational reliability.

You'll partner closely with engineering teams to improve processes, champion best practices, evaluate emerging technologies, and implement solutions that enhance the performance, resilience, and efficiency of our platforms. The ideal candidate is a collaborative technical leader who communicates effectively across technical and non-technical teams and drives meaningful improvements that have a lasting impact across the organization.

This position is based on-site in either our Arlington, VA office, Reston, VA office or Santa Barbara/Goleta, CA office.

Key Responsibilities

  • Ensure the reliability and scalability of critical systems, meeting SLAs through proactive monitoring and effective incident response.
  • Develop and promote new technologies and tools, conducting research and creating proofs of concept to introduce solutions that enhance the team's capabilities.
  • Lead by example in fostering a culture of excellence and reliability.
  • Continuously evaluate and improve team processes and workflows to increase efficiency and reduce complexity.
  • Collaborate closely with cross-functional teams, product managers, and stakeholders to align on technical strategy and provide expert guidance.
  • Participate in on-call rotations, providing support and resolving complex technical issues.
Required Qualifications
  • Bachelor’s degree in Computer Science or a related technical field.
  • 5-8+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems.
  • Extensive experience with AWS services (EC2, S3, Lambda, VPC Networking) and deep knowledge of cloud infrastructure, networking, and security best practices.
  • Proficiency running, optimizing, and scaling Kubernetes clusters in production environments.
  • Experience using and writing Terraform to architect and manage production infrastructure.
  • Ability to create and utilize Infrastructure-as-code (IaC), GitOps practices, and automation tools to increase reliability and reduce manual tasks.
  • Proven success in leading teams or projects using Agile/Scrum methodologies.
  • Expertise in infrastructure and software architecture, capable of designing and implementing large-scale, reliable systems with minimal guidance.
  • Experience developing and managing comprehensive infrastructure monitoring and alerting strategies.
Desired Qualifications
  • 10+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems.
  • Advanced understanding of cloud and application security, identity management, and compliance.
  • Expertise in service mesh and service registration technologies, focusing on performance and reliability.
  • Experience in the aerospace industry.
  • Flexible Time Off, Sick, Family & Medical Leave
  • Medical, Dental, Vision, Life, LTD, STD (employer funded)
  • Vol Life, Critical Illness, Accidental, Hospital Indemnity, Pet Insurance (employee funded)
  • 401k with 3% non-elective company contribution
  • Stock Options
  • Free Parking
  • Free lunch daily in office

Umbra is an Equal Opportunity Employer. We do not discriminate in hiring on the basis of sex, gender identity, sexual orientation, race, color, religious creed, national origin, physical or mental disability, protected veteran status, or any other characteristic protected by federal, state, or local law.

Employment Eligibility Verification

In compliance with federal laws, all hired persons will be required to verify their identity and eligibility to work in the United States by completing the required Employment Eligibility Verification Form (I-9 Form) upon hire.

ITAR/EAR Requirements

This position may include access to technology and/or data that is subject to U.S. export controls pursuant to ITAR and EAR. To comply with federal export controls, all persons hired must be a U.S. citizen, U.S. national, U.S. lawful permanent resident, refugee or asylee as defined by 8 U.S.C. § 1324b(a)(3), or must otherwise be eligible to obtain the required authorizations from the U.S. Department of State and/or U.S. Department of Commerce as applicable.

Pay Transparency
This job posting may cover multiple career levels. To ensure greater transparency, we provide base salary ranges for all roles, regardless of location. Our standard pay ranges are based on the role’s function and level, benchmarked against similar growth‑stage companies. Compensation may vary based on geographical location, as certain regions may have different cost‑of‑living factors. The final offer will also be influenced by the candidate's skills, responsibilities, and relevant experience.

Compensation Range

The Compensation Range for this role is $150,000 - $180,000 DOE.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Umbra • Arlington (VA)

On-site
USD 150,000 - 180,000
Flexible Time Off
Stock Options
Free Parking
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Umbra • Santa Barbara (CA)

On-site
USD 150,000 - 180,000
Flexible Time Off
Medical Benefits
401k
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Umbra • Virginia (MN)

On-site
USD 150,000 - 180,000
Flexible Time Off
Medical, Dental, Vision
401k with company contribution
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Umbra • Reston (VA)

On-site
USD 150,000 - 180,000
Free Parking
Free lunch daily in office
Stock Options
+2
Senior Software Engineer (Remote Sensing)
Senior Software Engineer (Remote Sensing)

Umbra Lab, Inc. • Santa Barbara (CA)

On-site
USD 155,000 - 185,000
Flexible Time Off
Medical, Dental, Vision, Life, LTD,STD
401k with company contribution
+3
Senior Software Engineer
Senior Software Engineer

Umbra • Reston (VA)

On-site
USD 155,000 - 185,000
Flexible Time Off
Medical, Dental, Vision, Life
401k with company contribution
+3
Senior Software Engineer (Remote Sensing)
Senior Software Engineer (Remote Sensing)

Umbra • Reston (VA)

On-site
USD 155,000 - 185,000
Free parking
Free lunch
Stock options
+2
Senior Software Engineer (Order and Delivery)
Senior Software Engineer (Order and Delivery)

Umbra • Arlington (VA)

On-site
USD 180,000 - 220,000
Flexible Time Off
Medical Insurance
Stock Options
+2
Senior Software Engineer (Order and Delivery)
Senior Software Engineer (Order and Delivery)

Umbra • Santa Barbara (CA)

On-site
USD 180,000 - 220,000
Flexible Time Off
Medical, Dental, Vision, Life
401k with company contribution
+2
Engineering Manager
Engineering Manager

Umbra • Arlington (VA)

On-site
USD 195,000 - 235,000
Flexible Time Off
Medical Insurance
Dental Insurance
+4