Compute Engineer, Deployment

Insight Global

Town of Texas (WI)

On-site

USD 120,000 - 180,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Insight Global in the United States seeks a Compute Engineer to scale GPU and accelerator fleets from first power-on to production across multiple data halls. You will own turn-up, establish firmware baselines, automate hardware workflows, and coordinate with network and data center teams.

This role requires travel and hands-on hardware experience in large facilities. The position emphasizes automation, scalable validation, and cross-functional collaboration to ensure reliable, production-ready

Qualifications

  • Experience bringing up server or GPU fleets at scale, hundreds of nodes to production.
  • Deep Linux experience with BMC, IPMI, and Redfish; automation in Python or Go.
  • Hands-on data hall work: racking, cabling, component swaps, or remote hands.
  • Ability to triage failures across hardware, firmware, and software.

Responsibilities

  • Scale qualification through tooling, not headcount.
  • Own compute turn-up from facility availability to ready-for-service.
  • Coordinate with network, ICT, and hardware teams during turn-up.
  • Support incident responses on freshly-live capacity.
  • Drive qualification queues by automation.

Skills

Server bring-up
Linux & BMC
Automation (Python/Go)
Data hall racking
Remote hands
Hardware triage

Job description

Responsibilities

Responsibilities: This Compute Engineer will bring gigawatts of accelerators from first power-on to production. Facility availability to ready-for-service across thousands of racks per site, with a new data hall landing every few weeks.

They will make rack qualification faster than the fleet grows. Firmware baselines, burn-in, and cluster validation proven on every rack before a customer workload touches it, at a pace that never becomes the critical path.

They must be able to scale by tooling, not headcount. Deployed megawatts grow severalfold next year while the team stays near-flat, because anything done twice by hand becomes software.

Must own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.

Ability to qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.

Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.

Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.

Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.

Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.

Ability to travel 20-30% of the time to our Data Centers and Labs, as needed.

REQUIRED SKILLS AND EXPERIENCE
  • This person has: brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
  • Worked deep within Linux and out-of-band management: BMC, IPMI, and Redfish are daily tools for you, not occasional lookups. Has automated hardware workflows in Python or Go rather than clicking through them, and the second time you do anything by hand you turn it into software.
  • Has worked physically in data halls, racking, cabling, and swapping components, and you're just as effective acting as remote hands or directing them.
  • Be able to triage failures methodically across hardware, firmware, and software, isolating the fault to a component before reaching for a fix.
  • Willing to travel for turn-up windows when a new data hall comes online.
NICE TO HAVE SKILLS AND EXPERIENCE
  • Bonus: Kubernetes-based bare-metal provisioning.
  • Accelerator platform bring up (NVIDIA, AMD, or custom).
  • Burn-in and stress harness design.
  • DCIM and inventory tooling.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
Compute Platform Engineer
Compute Platform Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 120,000 - 180,000
Lead, Hardware Deployment Engineer
Lead, Hardware Deployment Engineer

SpaceXAI • Memphis (TN)

On-site
USD 150,000 - 230,000
Remote-Ready Compute Engineer for Scalable GPU Fleets
Remote-Ready Compute Engineer for Scalable GPU Fleets

Insight Global • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Compute Deployment Engineer
Compute Deployment Engineer

Fluidstack • New York (NY)

On-site
USD 120,000 - 190,000
Kubernetes-based bare-metal Provisiong
Accelerator platform bring-up
Burn-in and stress harness design
Data Center Engineer
Data Center Engineer

TechClub Inc • Kansas City (MO)

On-site
USD 90,000 - 130,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Production Engineer, Compute
Production Engineer, Compute

Jobtailor • San Francisco (CA)

On-site
USD 170,000 - 240,000