Senior Site Reliability Engineer

InvestedintheMission

El Segundo (CA)

On-site

USD 153,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) matching
Unlimited PTO
Health insurance including Vision and Dental
Lunch and snacks provided every day
Dinners provided twice a week
Equity in a fully funded startup

Job summary

Varda Space Industries is seeking a Senior Site Reliability Engineer to build and maintain critical infrastructure for their operations in El Segundo, CA. This role demands expertise in Kubernetes and DevOps principles to enhance system reliability and scalability. Successful candidates will have a strong background in scripting, cloud infrastructure, and automation. Benefits include equity, health coverage, and unlimited PTO, contributing to a creative and ambitious team focused on space development.

Qualifications

  • 5+ years of Site Reliability Engineering experience or 7+ years in DevOps/SRE in lieu of a degree.
  • Experience operating Kubernetes in production environments.
  • Knowledge of software-defined networking (VPC, Subnets, Firewalls).

Responsibilities

  • Deploy and maintain mission-critical applications and infrastructure.
  • Build Infrastructure as Code (IaC) frameworks using Terraform.
  • Implement and operate observability systems and actionable alerting.

Skills

Kubernetes
Infrastructure as Code (IaC)
Python
DevOps
SRE

Education

Bachelor’s degree in computer science or related STEM field

Tools

Terraform
Prometheus
Grafana

Job description

About Varda

Low Earth orbit is open for business. Varda is accelerating the development of commercial space infrastructure, from in-orbit pharmaceutical processing to reliable and economical reentry capsules.

From life-saving pharmaceuticals to more powerful fiber optics, there is a world of products used on Earth today that can only be manufactured in space. Varda is accelerating innovation in the orbital economy by creating both the products and infrastructure needed so space can directly benefit life on Earth. Our mission is to expand the economic bounds of humankind.

Our team is uniquely suited to accomplishing this goal, with leadership and staff comprised of veterans from SpaceX, Blue Origin, major pharmaceutical companies and Silicon Valley. Varda was founded in January 2021 by Will Bruey and Delian Asparouhov with significant backing from world class investors including Khosla Ventures, Lux Capital, Founders Fund, Caffeinated Capital, General Catalyst, and Also Capital.

Varda is headquartered in El Segundo, California, where we have offices and a production facility where our vehicles, equipment, and materials are built, integrated, and tested. Varda also has offices in Washington, DC and Huntsville, AL.

Join Varda, and work to create a bustling in-space ecosystem.

About This Role

At Varda Space Industries, we\'re pushing the boundaries of what\'s possible in space and materials science — and we’re looking for bold engineers to help us get there. As a Senior Site Reliability Engineer, you\'ll be critical in building, scaling, and maintaining the infrastructure that powers our systems on Earth, in orbit, and everything in between.

We are looking for an experienced engineer with deep working knowledge of Kubernetes and containerized technologies. You are a hands-on operator and builder who applies first-principles thinking to both software delivery (DevOps) and production reliability (SRE), and thrives in complex, mission-critical environments.

In this role, you will:

  • Solve challenging technical problems across a wide range of modern technologies.
  • Apply a software engineering mindset to automate operations and improve system reliability, scalability, and resilience
  • Design and build infrastructure that enables rapid development — from cloud-based services to embedded software running on spacecraft.
  • Shape Varda’s infrastructure strategy and drive operational excellence across containerized and modernized environments.
Responsibilities
  • Deploy, maintain, and operate mission-critical applications and infrastructure supporting spacecraft and company-wide systems.
  • Build and evolve Infrastructure as Code (IaC) frameworks using tools such as Terraform
  • Implement and operate observability systems (metrics, logging, tracing) and actionable alerting.
  • Build and maintain CI/CD pipelines to enable safe, repeatable, and rapid deployments.
  • Partner with software and hardware engineers to deliver highly operable, reliable, and scalable systems and pipelines, ensuring they have the tools and infrastructure needed for rapid iteration.
  • Identify, analyze, and resolve system bottlenecks and reliability risks; perform performance tuning and implement long-term stability improvements.
  • Respond to and resolve production incidents; perform root cause analysis and drive corrective actions through blameless postmortems.
  • Rotate through the team’s on-call schedule to keep critical systems healthy and responsive.
  • Must be willing to work extended hours and weekends as needed
  • Occasionally travel to customer sites and other Varda locations to troubleshoot, deploy, or test critical infrastructure.
Basic Qualifications
  • Bachelor’s degree in computer science, engineering, or related STEM field with 5+ years of Site Reliability Engineering experience, or 7+ years of progressive experience in DevOps, SRE, or Systems Engineering in lieu of a degree.
  • Experience with Infrastructure as Code (IaC) using tools like Terraform to automate server provisioning and configuration management
  • Experience operating Kubernetes or similar container orchestration platforms in production environments.
  • Experience with Prometheus, Grafana, InfluxDB, or similar technologies.
  • Knowledge of software-defined networking (VPC, Subnets, Firewalls, VPNs, etc.)
  • Python, Bash, PowerShell (or similar) scripting experience
  • Positive and strong communication skills, both written and oral
Preferred Skills and Experience
  • Experience in provisioning and managing scalable Azure cloud infrastructure using native tools and best practices
  • Experience implementing configuration management, provisioning, and workflow automation solutions via Infrastructure as Code, CI/CD, and GitOps (e.g., Ansible, Salt, ArgoCD, etc).
  • Strong understanding of Linux systems and container runtimes (e.g., containerd, Docker)
  • Experience with GPU workloads or high-throughput computing.
  • Hands-on experience operating and optimizing High Performance Computing (HPC) environments, including workload schedulers such as Slurm (e.g., queue/partition design, fair-share scheduling, and cluster resource management).
  • Experience with hybrid environments (cloud + on-prem or edge systems)
  • Experience debugging distributed systems at scale (network, storage, latency)
  • Experience with databases and data modeling
Pay Range
  • Site Reliability Engineer: $153,000.00 - $185,000.00/per year
  • This role is on-sitein El Segundo, CA
  • Leveling and base salary are determined by job-related skills, education level, experience level, and job performance
  • You will be eligible for long-term incentives in the form of stock options and/or long-term cash awards
ITAR Requirements
  • Varda, like all employers, must ensure that its employees working in the United States are lawfully authorized to work in the U.S. Additionally, our employees are exposed to and have access to certain export-controlled items. At present, some of our technology to which employees have access requires a license to be exported to individuals other than “U.S. Persons” as defined in U.S. export regulations. Because our employees are provided access to export-controlled items, our current policy is to only hire “U.S. persons” who are permitted to have access to our technology without an export license.
Benefits
  • Exciting team of professionals at the top of their field working by your side
  • Equity in a fully funded space startup with potential for significant growth
  • 401(k) matching
  • Unlimited PTO
  • Health insurance, including Vision and Dental
  • Lunch and snacks provided on site every day. Dinners provided twice a week.
  • Maternity / Paternity leave

Varda Space Industries is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. Candidates and employees are always evaluated based on merit, qualifications, and performance. We will never discriminate on the basis of race, color, gender, national origin, ethnicity, veteran status, disability status, age, sexual orientation, gender identity, martial status, mental or physical disability, or any other legally protected status.

E-Verify Statement

Varda Space Industries, Inc. participates in the U.S. Department of Homeland Security E-Verify program. The E-Verify program is an Internet-based employment eligibility verification system operated by the U.S. Citizenship and Immigration Services. Learn more about the E-Verify program.

E-Verify Notice Right To Work Notice

Read more

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Varda • El Segundo (CA)

On-site
USD 153,000 - 185,000
Flexible PTO
Medical/Dental/Vision insurance
Relocation assistance
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Vardaspace • El Segundo (CA)

On-site
USD 153,000 - 185,000
Flexible PTO policy
100% company-paid Medical, Dental, Vision insurance
401(k) retirement plan with 6% employer match
Site Reliability Engineer II
Site Reliability Engineer II

Vardaspace • El Segundo (CA)

On-site
USD 133,000 - 170,000
Stock options
Relocation assistance
Health insurance
+1
Site Reliability Internship - Spring 2027
Site Reliability Internship - Spring 2027

Varda Space Industries • El Segundo (CA)

On-site
USD 38,000 - 53,000
Housing stipend
Commuter benefit
Relocation support
Senior Build Reliability Engineer
Senior Build Reliability Engineer

Vardaspace • El Segundo (CA)

On-site
USD 140,000 - 170,000
Flexible PTO and holidays
100% company-paid health insurance
Wellness reimbursement and gym expenses
Space Build Reliability Engineer
Space Build Reliability Engineer

Vardaspace • El Segundo (CA)

On-site
USD 109,000 - 152,000
Stock options
Relocation assistance
On-site in El Segundo
+1
Senior Embedded Linux Engineer, C++
Senior Embedded Linux Engineer, C++

Varda Space Industries • El Segundo (CA)

On-site
USD 169,000 - 216,000
Flexible PTO + 12 holidays
100% company-paid Medical, Dental, and
Senior Spacecraft Firmware Engineer, C++
Senior Spacecraft Firmware Engineer, C++

Socket.dev • El Segundo (CA)

On-site
USD 169,000 - 216,000
Stock options
Health insurance (employee + dependent
401(k) employer match
+1
Senior Spacecraft Embedded Software Engineer, C++
Senior Spacecraft Embedded Software Engineer, C++

Socket.dev • El Segundo (CA)

On-site
USD 169,000 - 216,000
Stock options
Relocation support
Comprehensive benefits
Senior Spacecraft Build Reliability Engineer, FMEA
Senior Spacecraft Build Reliability Engineer, FMEA

Varda • El Segundo (CA)

On-site
USD 140,000 - 179,000
Health & Wellness
12 paid holidays
Company equity opportunities
+1