Turn this role into an interview — a resume and cover letter built around what this employer wants.
Crux AI seeks a senior leader to own all facility operations for its AI compute fleet, spanning owned campuses and third-party colocation sites. You will build the organization from zero, establishing staffing templates, 24/7 coverage, and rigorous audit programs to ensure uptime for high-density liquid-cooled infrastructure.
You will command electrical, mechanical, fire/life-safety, and building automation systems, drive fleet-wide standards, and manage O&M with partners under SLAs.
Built to set the gold standard for integrated AI infrastructure
Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.
Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.
You will own critical facility operations for Crux’s entire fleet — every megawatt of electrical, mechanical, and controls infrastructure our TPU clusters depend on, across owned campuses and third-party colocation sites. You are accountable for uptime: the maintenance programs, operating procedures, compliance regime, and 24/7 site teams that keep high-density, liquid-cooled environments continuously available for AI workloads that do not tolerate thermal or power excursions. You will build this organization from zero — campus facility managers, chief engineers, and critical facility technicians — and set the operating standards every current and future site runs on. Where a colocation partner operates the building, you own the standard they are held to and the audit program that proves it.
Build and lead the fleet-wide facility operations organization — campus facility managers, chief engineers, controls specialists, and 24/7 critical facility technician teams — with staffing templates and coverage models that scale ahead of the capacity ramp.
Own fleet uptime and facility performance: availability of electrical, mechanical, fire/life-safety, and building automation systems, plus efficiency metrics (PUE, WUE, energy cost per MW) reported to the executive team monthly.
Author and enforce the operating standards library — SOPs, MOPs, EOPs, change management, and CMMS-driven preventive/corrective maintenance programs — applied consistently across owned and partner-operated sites.
Define the operating model per site archetype: what we self-perform versus contract (O&M providers, OEM service agreements, colo operator scope), and manage colocation partners to contractual SLAs with a real audit and compliance program behind it.
Own the construction-to-operations transition fleet-wide: ops acceptance criteria in every commissioning plan, participation through Level 5/IST, first-year tuning of high-density liquid cooling plants, and warranty enforcement.
Run fleet incident and problem management for facility events — detection, response, root cause, corrective action, and customer communication.
Own the safety program — LOTO, energized work controls, hazardous energy management, working at height — with zero-compromise standards across employees, vendors, and partner staff.
Feed operational lessons back into design: sit in design reviews, hold engineering accountable for maintainability and operability, and keep the fleet design standard honest with field data.
Fleet operating model is drafted (self-perform/contract boundaries, staffing template, coverage model).
Operating standards framework and CMMS platform is selected.
Ops acceptance criteria is embedded in the first site’s commissioning plan.
Hiring is underway for the first site’s facility leadership.
First site is operating with 24/7 coverage under the full standards library.
Preventive maintenance program live in CMMS with zero overdue critical PMs.
Incident and change management program is running.
Colo partner SLA and audit program is in force at partner sites.
15+ years in critical facility or data center operations, including 7+ years leading multi-site facility operations organizations at portfolio scale — you have owned uptime for a fleet.
Deep technical command of electrical (utility through UPS and distribution), mechanical (chilled water, direct-to-chip liquid cooling, heat rejection), fire/life-safety, and building automation/EPMS systems — enough to challenge a chief engineer’s diagnosis on its merits.
Built or scaled a facility operations program through rapid growth: standards libraries, CMMS deployments, commissioning-to-operations handoffs, and vendor/colo O&M management under contractual SLAs.
Outcome driven and hands-on leadership. Close enough to the work to be useful. This is a builder’s seat, not a supervisory one. You will personally walk the plants, sit in the root-cause reviews, and read the switchgear maintenance reports — not because there is no team, but because at this stage the person accountable for uptime has to be close enough to the work to change it. You hire people better than you at their craft, then stay close enough to know when something is off before it shows up in a report.
Uptime through discipline, not heroics. You believe availability is produced by procedures, drills, maintenance rigor, and change control — and that a site saved by a heroic 3am intervention is a site whose program failed. You build the system that makes the next incident boring.
High tolerance for ambiguity. The first sites will energize while the standards library is still being written. You absorb that, translate genuine unknowns into concrete near-term priorities, and never manufacture false certainty about readiness the commissioning data does not support.
AI-agentic first. Fluent with AI agents — or committed to becoming so quickly — and you embed them as first principles in how you and your team work, defaulting to agentic workflows before adding headcount or process.
Hyperscaler facility operations leadership (Meta, Google, AWS, Microsoft) or neocloud fleet operations during a rapid ramp.
Direct-to-chip liquid cooling plant operations at production scale (CDUs, TCS/FWS loops, water chemistry programs).
Experience holding colocation operators to SLA across a leased portfolio, including audit and compliance program design.
Commissioning leadership (Levels 1–5, IST) on mission-critical projects and first-year plant tuning.
Trade certifications (Electrical, HVAC, Controls), professional engineering license, or military/nuclear power background.
The annual salary range for this position has been estimated based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
We offer generous base, bonus and additional incentive based compensation
Health, dental, and vision coverage for you and your dependents
Company-paid life insurance and disability
Full suite of other optional benefits
401(k) Plan with 4% company match (USA employees)