Senior Site Reliability Engineer

The Recruiting Guy

Arlington (VA)

On-site

USD 175,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A pioneering open source technology company is seeking a Senior Cloud Infrastructure Engineer to design, maintain, and optimize large-scale systems for AI workloads. The role involves managing GPU environments, collaborating on infrastructure design, and driving best practices within a fast-growing team. Candidates should have substantial experience with Python, Kubernetes, Terraform, and Ansible.

Qualifications

  • 5+ years experience as an Infrastructure Engineer or Site Reliability Engineer building and operating large-scale distributed systems.
  • Skilled in Python and comfortable with Terraform and Ansible.
  • Familiar with Kubernetes and related tools like FluxCD, Prometheus, and Grafana.

Responsibilities

  • Design, build, and maintain core infrastructure for AI workloads.
  • Manage and automate GPU compute clusters using Python, Kubernetes, Terraform, and Ansible.
  • Collaborate with core engineers on infrastructure for new features.

Skills

Python
Kubernetes
Terraform
Ansible
FluxCD
Prometheus
Grafana

Job description

Senior Cloud Infrastructure Engineer

Location: San Francisco, CA (On‑site only) — must live within commuting distance or be willing to relocate.

Compensation: $175,000 – $250,000 per year (Salaried W2 Full‑Time).

About the Company

We are a pioneering open source technology company transforming how creators interact with generative AI. Our node‑based visual interface enables artists, developers, and innovators to design, control, and customize AI workflows with complete flexibility.

About the Role

You will lead the design, deployment, and maintenance of large‑scale distributed systems that power AI workloads. Collaborate closely with core engineers to shape the company’s long‑term infrastructure vision while ensuring scalability, performance, and reliability across environments.

What You’ll Do
  • Design, build, and maintain the core infrastructure that powers AI workloads at scale.
  • Manage and automate GPU compute clusters using tools such as Python, Kubernetes, Terraform, and Ansible.
  • Architect and operate systems for orchestration, observability, distributed storage, and networking.
  • Ensure reliability, scalability, and performance across production environments.
  • Collaborate closely with core engineers to design infrastructure for new features and systems.
  • Contribute to technical strategy and long‑term infrastructure vision.
  • Drive best practices for infrastructure automation, deployment, and monitoring.
Requirements
  • 5+ years experience as an Infrastructure Engineer or Site Reliability Engineer building and operating large‑scale distributed systems.
  • Skilled in Python and comfortable working with infrastructure‑as‑code tools such as Terraform and Ansible.
  • Familiar with container orchestration systems such as Kubernetes and related tooling like FluxCD, Prometheus, and Grafana.
  • Capable of managing high‑performance GPU environments across cloud and bare metal setups.
  • Highly adaptable, resourceful, and motivated by building things from the ground up.
  • Excited to work in a small, fast‑growing team where autonomy and accountability are key.
  • Comfortable working on‑site in a startup setting where collaboration and speed matter most.
Bonus Points
  • Experience contributing to or maintaining open‑source projects.
  • Background working with AI infrastructure, ML pipelines, or GPU orchestration.
  • Strong computer science fundamentals and ability to work across different programming languages or frameworks.
Skills

fluxcd, ansible, kubernetes, grafana, prometheus, python, terraform, infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Washington

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • San Francisco (CA)

On-site
USD 175,000 - 250,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

The Recruiting Guy • New York (NY)

On-site
USD 175,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior Platform Engineer – AI/ML Infrastructure & Reliability
Senior Platform Engineer – AI/ML Infrastructure & Reliability

StratITech • San Francisco (CA)

On-site
USD 210,000 - 260,000
Equity
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Senior Software Engineer, Infrastructure/Platform
Senior Software Engineer, Infrastructure/Platform

David Joseph & Company • San Francisco (CA)

On-site
USD 250,000 - 350,000
Equity
Site Reliability Engineer
Site Reliability Engineer

BridgeSource Utilities Solutions • United States

Hybrid
USD 140,000 - 190,000
Senior / Lead Infrastructure & Operations Engineer
Senior / Lead Infrastructure & Operations Engineer

Austin Werner • Boston (MA)

On-site
USD 120,000 - 150,000