Most job posts are designed to collect as many applications as possible.
This is not one of them.
At Simplify, we work directly with companies on high-priority hires.
We're partnering with an AI compute infrastructure company building the data centers behind the AI frontier to hire a Software Engineer for its Compute Operations team. $224–300K compensation + equity.
About the company
One-liner: Building the physical and software infrastructure that brings large-scale AI compute online.
Stage: A rapidly scaling, private AI infrastructure company with in-person teams in Austin, New York City, San Francisco, and Seattle.
The company works across data-center development, construction, and compute operations. This is a forward-deployed engineering seat pointed at the company's physical operations. You embed with warehouse staff, production engineers, technicians, and facility operators and build the software that runs a global compute fleet
You won't be handed a narrowly specified feature. You'll find the bottleneck, understand the decisions behind it, ship a working product, and improve it with the people who use it.
What you'll work on
- Build a fleet health system with real-time telemetry and tiered health checks across Kubernetes and bare metal, exposed through a shared API, with alarms correlated into incidents and probable causes drafted for on-call responders.
- Create a tracked repair and return-material-authorization workflow from failure detection through triage, parts, vendor return, and service restoration; automate repair routing, generate production engineers' shift task lists, and report time to return to service.
- Develop rack-level software workflows for burn-in, performance baselining, and hardware validation so accelerators can be brought online repeatably with acceptance evidence recorded for every machine.
- Own the facility maintenance system for lockout/tagout and work orders, its rollout to additional sites, asset-register readiness, migration, and retirement of legacy datacenter inventory ahead of new facility activation.
- Convert standard operating procedures, training records, and technician qualifications into auditable structured data; report site service-level objectives, deployment cycle time, and labor ramp on customer-requested dashboards.
- Work on site and on the on-call rotation alongside production engineers and facility operators, building systems for their operational use.
What we look for
- You have shipped production code in Go, Python, or TypeScript, and can learn whichever language the problem requires.
- You have built production features using large language model APIs, including OpenAI, Anthropic, or open-weight models, MCP servers, and agentic frameworks.
- You work daily with AI coding tools such as Claude Code and Cursor, and use agents autonomously to complete useful work.
- You identify problems, design solutions, and ship them without waiting for direction or approval.
- You have moved quickly under deadlines while building foundations that other engineers can extend.
- You have participated in an on-call rotation or worked alongside on-call staff, and translated operational pain into systems that reduce pager noise.
- You have demonstrated product judgment through interfaces and workflows that are clear to engineers on rotation and reflect how the work is performed.
Nice to have
- Production engineering or site reliability engineering on large GPU fleets.
- Experience with hardware qualification or burn-in frameworks.
- Experience with baseboard management controller, Redfish, or IPMI tooling.
- Experience with computerized maintenance management systems, data center infrastructure management, or asset management systems.Experience with building management systems, electrical power monitoring systems, or SCADA.
- Experience with Prometheus and Grafana.
- $224–300K compensation + equity
- Your software helps run the physical infrastructure behind frontier AI
- Small team, full ownership from day one: you find the problem, ship the fix, and operate it
- Engineering seat with frontline users — short feedback loops and visible impact on the next shift
- Build across a rare boundary: software, GPU hardware, and the physical data center