Get more replies from employers
Send a job-specific resume in minutes.
Insight Global in the United States seeks a Compute Engineer to scale GPU and accelerator fleets from first power-on to production across multiple data halls. You will own turn-up, establish firmware baselines, automate hardware workflows, and coordinate with network and data center teams.
This role requires travel and hands-on hardware experience in large facilities. The position emphasizes automation, scalable validation, and cross-functional collaboration to ensure reliable, production-ready
Responsibilities: This Compute Engineer will bring gigawatts of accelerators from first power-on to production. Facility availability to ready-for-service across thousands of racks per site, with a new data hall landing every few weeks.
They will make rack qualification faster than the fleet grows. Firmware baselines, burn-in, and cluster validation proven on every rack before a customer workload touches it, at a pace that never becomes the critical path.
They must be able to scale by tooling, not headcount. Deployed megawatts grow severalfold next year while the team stays near-flat, because anything done twice by hand becomes software.
Must own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.
Ability to qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.
Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.
Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.
Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.
Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.
Ability to travel 20-30% of the time to our Data Centers and Labs, as needed.