Senior Cloud Infrastructure Engineer

The Recruiting Guy

New York (NY)

On-site

USD 175,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A pioneering technology firm in San Francisco seeks a Senior Cloud Infrastructure Engineer. This role involves designing and maintaining large-scale distributed systems for AI workloads. Ideal candidates have 5+ years of experience, are skilled in Python and have a strong foundation in infrastructure automation tools. You will work closely with engineers to implement best practices. On-site work is required in a collaborative startup environment.

Qualifications

  • 5+ years experience as an Infrastructure Engineer or Site Reliability Engineer.
  • Skilled in Python and infrastructure‑as‑code tools such as Terraform and Ansible.
  • Familiar with container orchestration systems like Kubernetes.

Responsibilities

  • Design and maintain core infrastructure for AI workloads.
  • Manage GPU compute clusters using Python, Kubernetes, Terraform, and Ansible.
  • Collaborate with engineers to design infrastructure for new systems.

Skills

Python
Kubernetes
Terraform
Ansible
Grafana
Prometheus
fluxcd

Job description

Base pay range

$175,000 – $250,000 per year

Job Title: Senior Cloud Infrastructure Engineer

Location: San Francisco, CA (On-site only)

Employment Type: Salaried, W-2 Full-Time

Salary Range: $175,000 – $250,000

About The Company

We represent a pioneering open source technology company in San Francisco that is transforming the way creators interact with generative AI. They are the team behind a powerful, node‑based visual interface that gives artists, developers, and innovators the ability to design, control, and customize AI workflows with complete flexibility. Their platform allows users to connect modular components, build complex pipelines, and run everything locally with impressive speed and precision.

Their mission is to make generative AI open, transparent, and accessible to everyone. Built around community collaboration and creative empowerment, their tools help users experiment freely and bring their ideas to life. Whether it is visual storytelling, image generation, or advanced machine learning, their technology gives creators the freedom to explore without limits.

About The Role

In this role, you will lead the design, deployment, and maintenance of large‑scale distributed systems that power AI workloads. The ideal candidate is deeply technical, self‑sufficient, and motivated by solving complex infrastructure challenges. You will work closely with core engineers to shape the company’s long‑term infrastructure vision while ensuring scalability, performance, and reliability across environments.

What You’ll Do
  • Design, build, and maintain the core infrastructure that powers AI workloads at scale
  • Manage and automate GPU compute clusters using tools such as Python, Kubernetes, Terraform, and Ansible
  • Architect and operate systems for orchestration, observability, distributed storage, and networking
  • Ensure reliability, scalability, and performance across production environments
  • Collaborate closely with core engineers to design infrastructure for new features and systems
  • Contribute to technical strategy and long‑term infrastructure vision
  • Drive best practices for infrastructure automation, deployment, and monitoring
Requirements
  • 5+ years experience as an Infrastructure Engineer or Site Reliability Engineer building and operating large‑scale distributed systems
  • Skilled in Python and comfortable working with infrastructure‑as‑code tools such as Terraform and Ansible
  • Familiar with container orchestration systems such as Kubernetes and related tooling like FluxCD, Prometheus, and Grafana
  • Capable of managing high‑performance GPU environments across cloud and bare‑metal setups
  • Highly adaptable, resourceful, and motivated by building things from the ground up
  • Excited to work in a small, fast‑growing team where autonomy and accountability are key
  • Comfortable working on‑site in a startup setting where collaboration and speed matter most
Bonus Points
  • Experience contributing to or maintaining open‑source projects
  • Background working with AI infrastructure, ML pipelines, or GPU orchestration
  • Strong computer science fundamentals and ability to work across different programming languages or frameworks

Skills: fluxcd, ansible, kubernetes, grafana, prometheus, python, terraform, infrastructure

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • San Francisco (CA)

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Washington

On-site
USD 175,000 - 250,000
Senior Platform Engineer – AI/ML Infrastructure & Reliability
Senior Platform Engineer – AI/ML Infrastructure & Reliability

StratITech • San Francisco (CA)

On-site
USD 210,000 - 260,000
Equity
Director of Infrastructure Engineering
Director of Infrastructure Engineering

Appsierra Group • United States

On-site
USD 350,000 - 500,000
Equity compensation eligibility
Performance-based bonuses
Health insurance reimbursement up to 1
+3
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Stealth Startup • San Francisco (CA)

On-site
USD 250,000 - 400,000
Senior Software Engineer, Infrastructure/Platform
Senior Software Engineer, Infrastructure/Platform

David Joseph & Company • San Francisco (CA)

On-site
USD 250,000 - 350,000
Equity
Senior / Lead Infrastructure & Operations Engineer
Senior / Lead Infrastructure & Operations Engineer

Austin Werner • Boston (MA)

On-site
USD 120,000 - 150,000
Member of Technical Staff, Infrastructure at high-growth AI infrastructure startup
Member of Technical Staff, Infrastructure at high-growth AI infrastructure startup

Jack & Jill • San Francisco (CA)

On-site
USD 180,000 - 240,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000