Senior Site Reliability Engineer

The Recruiting Guy

Washington (District of Columbia)

On-site

USD 175,000 - 250,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A leading technology recruitment firm is looking for a Senior Cloud Infrastructure Engineer to design and maintain infrastructure for AI workloads. This full-time position is strictly on-site in San Francisco, requiring 5+ years of experience and skills in Python, Kubernetes, and GPU management. The ideal candidate will thrive in a fast-paced startup environment.

Qualifications

  • 5+ years of experience as an Infrastructure Engineer or Site Reliability Engineer.
  • Skilled in Python and infrastructure-as-code tools.
  • Experience managing high-performance GPU environments.

Responsibilities

  • Design and maintain infrastructure for AI workloads.
  • Automate GPU compute clusters using Python and Kubernetes.
  • Collaborate to shape infrastructure vision.

Skills

Python
Terraform
Ansible
Kubernetes
FluxCD
Prometheus
Grafana

Job description

Senior Cloud Infrastructure Engineer

Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or be willing to relocate. Relocation Assistance: No. Employment Type: Salaried W2 Full‑Time. Salary Range: $175,000 – $250,000.

About the Company

We represent a pioneering open source technology company in San Francisco that is transforming the way creators interact with generative AI.

They are the team behind a powerful, node‑based visual interface that gives artists, developers, and innovators the ability to design, control, and customize AI workflows with complete flexibility.

Their platform allows users to connect modular components, build complex pipelines, and run everything locally with impressive speed and precision.

Their mission is to make generative AI open, transparent, and accessible to everyone. Built around community collaboration and creative empowerment, their tools help users experiment freely and bring their ideas to life.

Whether it is visual storytelling, image generation, or advanced machine learning, their technology gives creators the freedom to explore without limitations.

About the Role

In this role, you will take the lead on designing, deploying, and maintaining large‑scale distributed systems that power AI workloads. You will work closely with core engineers to shape the company’s long‑term infrastructure vision while ensuring scalability, performance, and reliability across environments.

What You’ll Do
  • Design, build, and maintain the core infrastructure that powers AI workloads at scale
  • Manage and automate GPU compute clusters using tools such as Python, Kubernetes, Terraform, and Ansible
  • Architect and operate systems for orchestration, observability, distributed storage, and networking
  • Ensure reliability, scalability, and performance across production environments
  • Collaborate closely with core engineers to design infrastructure for new features and systems
  • Contribute to technical strategy and long‑term infrastructure vision
  • Drive best practices for infrastructure automation, deployment, and monitoring
Requirements
  • 5+ years of experience as an Infrastructure Engineer or Site Reliability Engineer building and operating large‑scale distributed systems
  • Skilled in Python and comfortable working with infrastructure‑as‑code tools such as Terraform and Ansible
  • Familiar with container orchestration systems such as Kubernetes and related tooling like FluxCD, Prometheus, and Grafana
  • Capable of managing high‑performance GPU environments across cloud and bare metal setups
  • Highly adaptable, resourceful, and motivated by building things from the ground up
  • Excited to work in a small, fast‑growing team where autonomy and accountability are key
  • Comfortable working on‑site in a startup setting where collaboration and speed matter most
Bonus Points
  • Experience contributing to or maintaining open‑source projects
  • Background working with AI infrastructure, ML pipelines, or GPU orchestration
  • Strong computer science fundamentals and ability to work across different programming languages or frameworks
Skills

fluxcd, ansible, kubernetes, grafana, prometheus, python, terraform, infrastructure

Seniority Level

Mid‑Senior level

Employment Type

Full‑time

Job Function

Engineering and Information Technology

Industries

Human Resources Services

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • San Francisco (CA)

On-site
USD 175,000 - 250,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

The Recruiting Guy • New York (NY)

On-site
USD 175,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Motion Recruitment • United States

Remote
USD 140,000 - 170,000
Medical, dental, and vision
Equity / Stock Options
Remote equipment stipend
+3
Software Engineer, Site Reliability
Software Engineer, Site Reliability

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Relocation assistance
Learning and growth opportunities
+1
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 130,000 - 180,000
Medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
+3
Senior Data Center Infrastructure Software Engineer
Senior Data Center Infrastructure Software Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision
401(k) with company match
Paid holidays
Sr. Software Engineer
Sr. Software Engineer

Addison Group • Dallas (TX)

Hybrid
USD 170,000 - 220,000
Competitive annual performance bonus
Employer-paid medical benefits for you
401(k) with generous company match
+7
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Red Oak Technologies Inc. • Mountain View (CA)

On-site
USD 145,000 - 190,000
Bonus opportunities
Comprehensive benefits package