Senior System Software Engineer

Acceler8 Talent

Mountain View (CA)

Hybrid

USD 225,000 - 275,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Acceler8 Talent is sourcing a System Software Engineer for Node & Cluster Management in Mountain View, CA. This hybrid role requires in-person presence Tue–Thu and offers a $250k+ base with RSUs. You will shape how a first-generation AI platform is operated from a single node to a full cluster, interfacing with host software and OpenBMC firmware.

The role emphasizes low-level Linux development, node health, telemetry, PCIe devices, and firmware interactions across hardware boundaries.

Qualifications

  • 8+ years of systems-software experience.
  • Strong Linux systems development and low-level userspace experience.
  • Experience building REST APIs and CLI tools for hardware or infrastructure management.

Responsibilities

  • Build node-level management services for health, telemetry, inventory, and control.
  • Design cluster-management and failover to minimize downtime.
  • Develop REST APIs and CLI tools for diagnostics, firmware updates, and device recovery.
  • Create unified management across host software and OpenBMC firmware.
  • Extend management from nodes to racks and clusters.
  • Debug across APIs, daemons, kernel drivers, firmware and hardware.
  • Build provisioning, test automation, and monitoring tools for lab systems.

Skills

C
Go
Rust
C++
Python
Linux
REST APIs
CLI tools
Device drivers
BMC

Tools

Redfish
OpenBMC
gNMI
IPMI
PCIe
Telemetry

Job description

System Software Engineer, Node & Cluster Management

Hybrid | Mountain View, CA | Tuesday to Thursday on-site

$250k+ base + RSUs

I’m working with a well-funded AI semiconductor startup building a new processor architecture for frontier-scale language models.

The host software team owns the stack that makes its custom AI hardware usable, from low-level Linux interfaces through node and cluster management. This role will help define how a new compute platform is monitored, controlled, updated, and recovered at datacenter scale.

You’ll work on problems such as:

  • Building node-level management services for health, telemetry, inventory, and control
  • Designing cluster-management and failover capabilities that minimize downtime
  • Developing REST APIs and CLI tools for diagnostics, firmware updates, and device recovery
  • Creating unified management across host software and OpenBMC firmware
  • Extending management capabilities from individual nodes to racks and clusters
  • Debugging issues across APIs, daemons, kernel drivers, firmware, and hardware
  • Building provisioning, test automation, and monitoring tools for lab systems.

Looking for engineers who have:

  • 8+ years of systems-software experience
  • Strong Linux systems development and low-level userspace experience
  • Strong C, plus Go, Rust, C++, or Python
  • Experience building REST APIs and CLI tools for hardware or infrastructure management
  • An understanding of device drivers, telemetry paths, PCIe devices, and BMC-managed systems
  • Experience debugging across software, firmware, and hardware boundaries
  • The autonomy to build working software against new and evolving hardware.

Nice to have:

  • Redfish, OpenBMC, gNMI, or IPMI experience
  • Cluster or fleet management for GPUs or AI accelerators
  • Hardware bring-up, lab automation, or manufacturing-test experience
  • Firmware-update, secure-boot, or attestation experience.

This is a systems role rather than a conventional web-services position. You’ll work close to the hardware while shaping how a first-generation AI platform is operated from a single node through a full cluster.

Candidates must be authorized to work in the United States and able to work from the Mountain View office Tuesday through Thursday.

Worth a confidential chat?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Software Engineer: Node & Cluster Management
Senior Systems Software Engineer: Node & Cluster Management

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 225,000 - 275,000
System Software Engineer
System Software Engineer

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 140,000 - 210,000
Host Systems Software Engineer
Host Systems Software Engineer

Slope • San Francisco (CA)

On-site
USD 120,000 - 160,000
Host Systems Software Engineer
Host Systems Software Engineer

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Systems Software Engineer, Silicon Bringup
Systems Software Engineer, Silicon Bringup

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
AI Infra Systems Engineer — Hybrid, GPU & Scale
AI Infra Systems Engineer — Hybrid, GPU & Scale

Delos Data Inc • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits
System Software Engineer - AI
System Software Engineer - AI

Delos Data • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits
System Engineer
System Engineer

Acceler8 Talent • Fremont (CA), Northern (KY)

Hybrid
USD 135,000 - 165,000
Software Engineer, Manufacturing Infrastructure
Software Engineer, Manufacturing Infrastructure

United States Digital Space LLC • United States

Hybrid
USD 180,000 - 260,000
Infrastructure Engineer
Infrastructure Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 390,000
Equity