Software Engineer, Machine Lifecycle

Career Techniques

New York (NY)

Hybrid

USD 150,000 - 250,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Career Techniques seeks an automation-focused engineer to own the journey of every fleet machine—from rack to production. You will design a GitOps-driven pipeline handling discovery, firmware, BIOS, OS install, and health validation, with out-of-band management and unattended provisioning.

You will replace manual runbooks with code in Python and Ansible, validating changes through CI/CD and automated burn‑in before production handoff.

Qualifications

  • Curious engineer who learns fast and loves automating physical infrastructure end to end.
  • Strong Python for building automation, tooling, and services, not just scripts.
  • Hands-on Ansible experience: writing playbooks and roles you’d code-review, not just run.
  • Solid grasp of CI/CD: pipelines, testing, staged rollouts, version control discipline.

Responsibilities

  • Design end-to-end machine lifecycle pipeline: power-on, network boot, OS install, validation, and production handoff.
  • Automate hardware bring-up via out-of-band management (BMC, Redfish, IPMI): firmware updates, BIOS, boot order, inventory.
  • Automate OS provisioning with network boot (PXE/UEFI) and unattended install.
  • Write and maintain Ansible and Python to configure machines into final roles as code.
  • Apply GitOps and CI/CD: desired state in Git, merge requests, pipelines, drift reconciliation.
  • Build automated validation and burn-in: health checks, stress tests, acceptance criteria before users see it.
  • Model lifecycle as a state machine with clear automated transitions and auditable history.
  • Instrument the pipeline with metrics and logging to track progress and bottlenecks.
  • Collaborate with HPC/datacenter teams to codify operational knowledge into automation.

Skills

Python automation
CI/CD
GitOps mindset
Linux
Communication

Tools

Ansible

Job description

This role owns the journey of every machine in the fleet: from the moment a server is racked, cabled, and powered on, to the moment it is fully configured, validated, and available for users. Your mission is to make that journey zero-touch. You will design and build the automation pipeline that takes a machine through discovery, firmware and BIOS configuration, OS installation, configuration management, health validation and burn-in, and finally handoff into production, treating each stage as code that lives in Git, runs through CI/CD, and can be reviewed, tested, and rolled back like any other software. The guiding principle is GitOps for physical infrastructure: the desired state of the fleet is declared in a repository, and automation continuously reconciles reality against it. A new machine shows up as a commit; a decommission is a deletion; drift is detected and corrected by the pipeline, not by a person with a checklist.

Responsibilities:
  • Design and build the end-to-end machine lifecycle pipeline: from power-on and network boot through OS install, configuration, validation, and production handoff.
  • Automate hardware bring-up via out-of-band management (BMC, Redfish, IPMI): firmware updates, BIOS settings, boot order, and inventory discovery.
  • Automate OS provisioning with network boot (PXE / UEFI HTTP boot) and unattended installation, so no one ever installs a machine by hand.
  • Write and maintain the Ansible and Python that configure machines into their final roles, replacing manual runbooks with reviewed, versioned code.
  • Apply GitOps and CI/CD principles to the fleet: desired state in Git, changes through merge requests, pipelines that test and apply them, and reconciliation that catches drift.
  • Build automated validation and burn-in: health checks, stress tests, and acceptance criteria a machine must pass before users ever see it.
  • Model the lifecycle as a state machine (new, provisioning, validating, in-service, needs-repair, decommissioned) with clear, automated transitions and an auditable history.
  • Instrument the pipeline with metrics and logging so we always know where a machine is in its lifecycle, and where the process is slow or failing.
  • Work with the HPC and datacenter teams to fold their hard-won operational knowledge into the automation, one stage at a time.
Qualifications:
  • A smart, curious engineer who learns fast and is genuinely excited by the challenge of automating physical infrastructure end to end. This matters more to us than any specific line on your resume.
  • Strong Python for building automation, tooling, and services, not just scripts.
  • Hands-on Ansible experience: writing playbooks and roles you would be happy to code-review, not just run.
  • A solid grasp of CI/CD principles: pipelines, testing, staged rollouts, and the discipline of driving change through version control.
  • An automation-first, GitOps mindset: you believe infrastructure state belongs in Git, and that any task done by hand twice should be code.
  • Working knowledge of Linux: comfortable with the boot process, system services, and debugging when a machine does not come up the way it should. Depth here is a real plus, but interest and trajectory count.
  • Sound engineering judgment: you design workflows that fail safely, retry sensibly, and leave an audit trail.
  • Clear communication and the patience to turn tribal operational knowledge into reliable, documented automation.
Nice to Have:
  • Experience with bare-metal provisioning tooling such as MAAS, Tinkerbell, Foreman, Ironic, or a home-grown equivalent.
  • Familiarity with out-of-band management: BMCs, Redfish, IPMI, and vendor variants like iDRAC or iLO.
  • Exposure to hardware validation and burn-in: stress testing, firmware qualification, or failure prediction at fleet scale.
  • Experience with GitOps tooling or declarative infrastructure management in general.
  • Prior work in datacenter, HPC, or large-fleet environments where machines number in the hundreds or thousands.

Comp: 150-250K base + Bonus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Machine Lifecycle
Software Engineer - Machine Lifecycle

Iceberg • New York (NY)

On-site
USD 140,000 - 210,000
Sr Software Engineer- Bare Metal Infrastructure
Sr Software Engineer- Bare Metal Infrastructure

H-E-B • San Antonio (TX)

On-site
USD 120,000 - 160,000
Software Engineer, Machine Lifecycle
Software Engineer, Machine Lifecycle

Tower Research Capital • New York (NY)

Hybrid
USD 150,000 - 250,000
Paid time off
Hybrid work
Meals provided
+5
Senior DevOps / Infrastructure Engineer
Senior DevOps / Infrastructure Engineer

Compunnel, Inc. • San Jose (CA)

On-site
USD 100,000 - 130,000
Staff Engineer, CI/CD & Cloud Infrastructure
Staff Engineer, CI/CD & Cloud Infrastructure

San Diego Stealth Startup • San Diego (CA)

On-site
USD 175,000 - 185,000
Lead DevOps Platform Engineer
Lead DevOps Platform Engineer

Jobtailor • Manassas (VA)

On-site
USD 160,000 - 210,000
Senior HPC DevOps Engineer
Senior HPC DevOps Engineer

Peraton • Maryland

On-site
USD 120,000 - 150,000
Linux Team Lead
Linux Team Lead

Jobtailor • Glen Burnie (MD)

On-site
USD 120,000 - 150,000
HPC Operations Engineer
HPC Operations Engineer

Career Techniques • New York (NY)

Hybrid
USD 175,000 - 225,000
Senior Compute Infrastructure Engineer
Senior Compute Infrastructure Engineer

NextGenEnergyJobs • Portland (OR)

On-site
USD 150,000 - 200,000
Equity in the company
Relocation assistance to Portland
Competitive cash compensation 150k–200