Software Engineer, AI Automation — Compute Operations

Simplify

San Francisco (CA)

On-site

USD 224,000 - 300,000

Full time

11 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Competitive compensation

Job summary

Simplify is hiring a Software Engineer for its Compute Operations team to help run the physical and software infrastructure behind frontier AI. You will embed with production staff and build systems that keep a global compute fleet healthy and productive.

You'll ship production features in Go, Python, or TypeScript, work with large language model APIs, and participate in on-call rotations to reduce pager noise and improve operator workflows.

Qualifications

  • Shipped production code in Go, Python, or TypeScript.
  • Built features using large language model APIs (OpenAI, Anthropic, etc.).
  • Experience with AI coding tools and autonomous agents.
  • Experience in on-call rotations and reducing pager noise.
  • Ability to design and ship robust infrastructure features.

Responsibilities

  • Develop fleet health systems with real-time telemetry and alerts.
  • Automate repair routing and create operator task lists.
  • Build rack-level workflows for burn-in and validation of hardware.
  • Own and evolve the facility maintenance system and inventory data.
  • Convert SOPs into structured, auditable data and dashboards.
  • Work on site with engineers and facilities staff on rotation.

Skills

Go
Python
TypeScript
LLM APIs
On-call experience

Tools

Kubernetes
Prometheus
Grafana

Job description

Most job posts are designed to collect as many applications as possible.

This is not one of them.

At Simplify, we work directly with companies on high-priority hires.

We're partnering with an AI compute infrastructure company building the data centers behind the AI frontier to hire a Software Engineer for its Compute Operations team. $224–300K compensation + equity.

About the company

One-liner: Building the physical and software infrastructure that brings large-scale AI compute online.

Stage: A rapidly scaling, private AI infrastructure company with in-person teams in Austin, New York City, San Francisco, and Seattle.

The company works across data-center development, construction, and compute operations. This is a forward-deployed engineering seat pointed at the company's physical operations. You embed with warehouse staff, production engineers, technicians, and facility operators and build the software that runs a global compute fleet

You won't be handed a narrowly specified feature. You'll find the bottleneck, understand the decisions behind it, ship a working product, and improve it with the people who use it.

What you'll work on
  • Build a fleet health system with real-time telemetry and tiered health checks across Kubernetes and bare metal, exposed through a shared API, with alarms correlated into incidents and probable causes drafted for on-call responders.
  • Create a tracked repair and return-material-authorization workflow from failure detection through triage, parts, vendor return, and service restoration; automate repair routing, generate production engineers' shift task lists, and report time to return to service.
  • Develop rack-level software workflows for burn-in, performance baselining, and hardware validation so accelerators can be brought online repeatably with acceptance evidence recorded for every machine.
  • Own the facility maintenance system for lockout/tagout and work orders, its rollout to additional sites, asset-register readiness, migration, and retirement of legacy datacenter inventory ahead of new facility activation.
  • Convert standard operating procedures, training records, and technician qualifications into auditable structured data; report site service-level objectives, deployment cycle time, and labor ramp on customer-requested dashboards.
  • Work on site and on the on-call rotation alongside production engineers and facility operators, building systems for their operational use.
What we look for
  • You have shipped production code in Go, Python, or TypeScript, and can learn whichever language the problem requires.
  • You have built production features using large language model APIs, including OpenAI, Anthropic, or open-weight models, MCP servers, and agentic frameworks.
  • You work daily with AI coding tools such as Claude Code and Cursor, and use agents autonomously to complete useful work.
  • You identify problems, design solutions, and ship them without waiting for direction or approval.
  • You have moved quickly under deadlines while building foundations that other engineers can extend.
  • You have participated in an on-call rotation or worked alongside on-call staff, and translated operational pain into systems that reduce pager noise.
  • You have demonstrated product judgment through interfaces and workflows that are clear to engineers on rotation and reflect how the work is performed.
Nice to have
  • Production engineering or site reliability engineering on large GPU fleets.
  • Experience with hardware qualification or burn-in frameworks.
  • Experience with baseboard management controller, Redfish, or IPMI tooling.
  • Experience with computerized maintenance management systems, data center infrastructure management, or asset management systems.Experience with building management systems, electrical power monitoring systems, or SCADA.
  • Experience with Prometheus and Grafana.
  • $224–300K compensation + equity
  • Your software helps run the physical infrastructure behind frontier AI
  • Small team, full ownership from day one: you find the problem, ship the fix, and operate it
  • Engineering seat with frontline users — short feedback loops and visible impact on the next shift
  • Build across a rare boundary: software, GPU hardware, and the physical data center
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, AI Automation — Business Operations
Software Engineer, AI Automation — Business Operations

Simplify • San Francisco (CA)

On-site
USD 180,000 - 300,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 405,000
Equity
Flexible work environment
Health benefits
Industrial Compute
Industrial Compute

OpenAI • United States

On-site
USD 150,000 - 300,000
Staff Software Engineer: Compute
Staff Software Engineer: Compute

Anthropic • Washington

On-site
USD 405,000 - 485,000
Visa sponsorship
Hybrid office policy
Engineering Manager, Production Platform and Orchestration
Engineering Manager, Production Platform and Orchestration

Jobgether SRL • United States

Remote
USD 200,000 - 300,000
Remote-first culture
Equity participation
Medical/dental/vision insurance
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
Full Stack Engineer, Fleet Scheduling
Full Stack Engineer, Fleet Scheduling

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 490,000
Staff Data Center Infrastructure Software Engineer
Staff Data Center Infrastructure Software Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 150,000 - 190,000
Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 490,000
Senior Data Center Infrastructure Software Engineer
Senior Data Center Infrastructure Software Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision
401(k) with company match
Paid holidays