Site Reliability Engineer (Manufacturing Infrastructure)

United States Digital Space LLC

Bastrop (TX)

On-site

USD 120,000 - 180,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking a Site Reliability Engineer to support manufacturing infrastructure across critical programs. The role emphasizes reliability, scalability, and proactive maintenance of compute, storage, and networking for factory environments.

Ideal candidates have strong software engineering fundamentals and can own complex installations while coordinating with site teams. On-call, travel to multiple sites, and strict on-site requirements apply.

Qualifications

  • Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 3+ years of professional experience in SRE/DevOps.
  • 1+ years of software development experience
  • Experience with Linux operating systems

Responsibilities

  • Deploy, upgrade, operate, maintain, and scale compute, storage, and networking for manufacturing systems across Starship, Starlink, Starshield, and Terafab
  • Manage infrastructure as code and use observability to provide a complete picture of platform health
  • Design for reliability, stability, and scale; find and remove bottlenecks with measurement and engineering
  • Practice proactive maintenance: capacity planning, lifecycle management, and reducing toil before it becomes an incident
  • Partner with software engineers, manufacturing stakeholders, and site teams to build operable, maintainable systems
  • Improve the full lifecycle—from design through deployment, operation, and continuous refinement
  • Practice sustainable incident response and blameless postmortems
  • Provide high-quality support to manufacturing and engineering users
  • Communicate clearly with stakeholders and teammates
  • Participate in on-call and travel to sites as needed for deployments, incidents, and cross-site reliability

Skills

SRE/DevOps
Software development
Linux
Stakeholder communication

Education

Bachelor’s degree in CS/IS/Engineering
3+ years SRE/DevOps experience

Tools

Terraform
Ansible
Docker
Kubernetes
Puppet
vSphere
QEMU/KVM

Job description

the company was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today the company is actively developing the technologies to make this possible, with the ultimate goal ofenabling human life on Mars.

SITE RELIABILITY ENGINEER (MANUFACTURING INFRASTRUCTURE)

The application software team is the central nervous system of the company. Manufacturing is how the company turns designs into hardware. The compute, storage, and networking that run our factories must be as reliable as the products we build. This team owns infrastructure supporting Starship, Starlink, Starshield, and Terafab. This position will have a direct impact on factory uptime, throughput, and production scale across programs.

The ideal candidate has strong software engineering fundamentals and a passion for infrastructure: reliability, stability, proactive maintenance, and scalability. You understand the system before you change it, solve hard problems, communicate clearly with stakeholders and teammates, and take ownership of work that manufacturing depends on.

Aerospace experience is not required. We value smart, motivated, collaborative engineers who treat teammates with fairness, respect, and support, and who want to take full ownership of challenging problems to help make humanity multi-planetary.

RESPONSIBILITIES:
  • Deploy, upgrade, operate, maintain, and scale compute, storage, and networking for manufacturing systems across Starship, Starlink, Starshield, and Terafab
  • Manage infrastructure as code and use observability to provide a complete picture of platform health
  • Design for reliability, stability, and scale; find and remove bottlenecks with measurement and engineering
  • Practice proactive maintenance: capacity planning, lifecycle management, and reducing toil before it becomes an incident
  • Partner with software engineers, manufacturing stakeholders, and site teams to build operable, maintainable systems
  • Improve the full lifecycle—from design through deployment, operation, and continuous refinement
  • Practice sustainable incident response and blameless postmortems
  • Provide high-quality support to manufacturing and engineering users
  • Communicate clearly with stakeholders and teammates
  • Participate in on-call and travel to sites as needed for deployments, incidents, and cross-site reliability
BASIC QUALIFICATIONS:
  • Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 3+ years of professional experience in SRE or DevOps in lieu of a degree
  • 1+ years of software development experience
  • Experience with Linux operating systems
PREFERRED SKILLS AND EXPERIENCE:
  • Experience with compute, storage, and/or networking infrastructure in production
  • Infrastructure as code (Terraform, Ansible, Puppet, or similar)
  • Containers and virtualization (Docker, Kubernetes, vSphere, QEMU, KVM, etc.)
  • Databases and data modeling (Postgres, Clickhouse, etc.)
  • Ability to translate high-level requirements into implementations from first principles
  • Comfort with mission-critical systems and appropriate urgency and care
  • Skillful communication with customers, peers, and management
  • Comfort operating across multiple sites and manufacturing programs
ADDITIONAL REQUIREMENTS:
  • Must be able to work extended hours and weekends as needed
  • Must be able to travel to different sites (Hawthorne, CA; Redmond, WA; Cape Canaveral, FL; Starbase, TX)
  • Ability to pass Air Force background check for Cape Canaveral
  • This role requires you to be onsite. Remote and/or hybrid work will not be considered
ITAR REQUIREMENTS:
  • To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.

the company is an Equal Opportunity Employer; employment with the company is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.

Applicants wishing to view a copy of the company’s Aff… or applicants requiring reasonable accommodation to the application/interview process should reach out tohr@unitedstatesdigital.space*.*

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (Manufacturing Infrastructure)
Site Reliability Engineer (Manufacturing Infrastructure)

SpaceX • Bastrop (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (Manufacturing Infrastructure)
Site Reliability Engineer (Manufacturing Infrastructure)

SPACE EXPLORATION TECHNOLOGIES CORP • Bastrop (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer (Manufacturing Infrastructure)
Site Reliability Engineer (Manufacturing Infrastructure)

SpaceX • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Software Engineer (Application Software)
Software Engineer (Application Software)

United States Digital Space LLC • Horn Lake (MS)

On-site
USD 70,000 - 100,000
Application Software Engineer - Memphis
Application Software Engineer - Memphis

United States Digital Space LLC • Horn Lake (MS)

On-site
USD 95,000 - 125,000
Production Engineer, Site Reliability (Application Software)
Production Engineer, Site Reliability (Application Software)

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, and dental insurance
Full Stack Software Engineer, Internal Systems - Memphis
Full Stack Software Engineer, Internal Systems - Memphis

United States Digital Space LLC • Horn Lake (MS)

On-site
USD 110,000 - 165,000
Site Reliability Engineer (Application Software)
Site Reliability Engineer (Application Software)

SpaceX • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) retirement plan
Medical, vision, dental coverage
+1
Full Stack Engineer (Application Software)
Full Stack Engineer (Application Software)

United States Digital Space LLC • El Segundo (CA)

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives (stock)
Discretionary bonuses
+6
Production Engineer, Site Reliability (Application Software)
Production Engineer, Site Reliability (Application Software)

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
Long-term cash awards
Health insurance
+6