System Software Engineer — Node & Cluster Management

MatX

Mountain View (CA)

On-site

USD 120,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Life insurance
Health Savings Account
Paid time off
12 holidays

Job summary

MatX is seeking a System Software Engineer to design and build the node-level management plane for its AI systems, exposing health, inventory, telemetry, and control via REST interfaces. You will work across the stack from kernel drivers to on-node daemons and BMC collaboration.

You will design cluster management, develop CLI tools, and ensure cohesive management across host and BMC, enabling scalable, reliable operations in a high-performance environment.

Qualifications

  • BS or higher in Computer Science, Electrical Engineering, or equivalent with 8+ years in systems software.
  • Strong Linux systems development experience, including low-level userspace; comfortable reading and debugging kernel driver/daemon code.
  • Proficient in C and systems languages (Go, Rust, C++, Python).
  • Experience designing and building HTTP/REST APIs and CLI tools for hardware or infrastructure management.
  • Solid understanding of device drivers, telemetry paths, PCIe behavior, and BMC subsystems.

Responsibilities

  • Design and build the node-level management plane exposing health, inventory, telemetry, and control via HTTP/REST endpoints.
  • Design and implement cluster management solutions and failover algorithms.
  • Build management CLI utilities for operators and internal engineers.
  • Collaborate with BMC firmware engineers for unified management across host and BMC.
  • Extend node capabilities to fleet-wide health, inventory, and alerting.

Skills

C/C++
Go
Rust
Python
Linux kernel drivers
HTTP/REST APIs
CLI tooling
Debugging

Education

BS or higher in CS or EE

Tools

OpenBMC
Redfish
PCIe

Job description

System Software Engineer — Node & Cluster Management

Mountain View, CA

What MatX Is Building

MatX's mission is to make the world’s best AI models run as efficiently as allowed by physics, bringing the world years ahead in AI quality and availability. MatX is seeking System Software Engineer to join our team as we create best-in-class silicon for high-performance and sustainable GenAI. Successful candidates for these roles will be responsible for delivering performant and functionally accurate silicon for MatX products across compute, memory management, high-speed connectivity and other key technologies.

The MatX host system software team owns everything that makes our AI silicon and systems usable: from Linux kernel drivers up through node and cluster management. The team also co-owns the BMC/OpenBMC firmware stack, with dedicated firmware engineers, so host software and out-of-band management are designed together rather than bolted together. We're looking for self-driven engineers who can take a hardware spec and a register map and just start building — prototype drivers, low-level utilities that talk directly to the chip, daemons, and tooling — with minimal hand-holding. Each engineer on this team has a primary focus area, but ownership of overlapping components is shared, and you should expect (and want) to venture across the stack.

What You'll Do Here
  • Design and build the node-level management plane for MatX's AI systems: expose node health, inventory, telemetry, and control operations through HTTP/REST endpoints (e.g., Redfish-style or custom APIs)
  • Design and implement cluster management solutions and failover algorithms to minimize downtime
  • Build the management CLI utilities that operators and internal engineers use daily — interacting with the on-node management and telemetry daemons to query state, run diagnostics, update firmware, and recover devices
  • Partner with our BMC firmware engineers to present unified management and observability across in-band and out-of-band paths — so operators see one coherent node, whether data comes from the host daemons or the BMC (e.g., unified Redfish-style views, firmware update orchestration across host and BMC, and recovery flows that work even when the host is down)
  • Extend node-level capabilities to cluster level: fleet-wide health aggregation, device inventory, alerting hooks, and integration points for our customers' own fleet-management systems
  • Get hands-on with the low-level stack: you'll regularly need to drop below the API layer — into the telemetry daemon, driver interfaces, or raw device access utilities — to prototype, debug, or unblock yourself
  • Build tooling and automation for managing lab systems during bring-up: provisioning, test orchestration, regression monitoring
  • Define the software contracts between the on-node daemons, the BMC stack, and the management layer — shared-ownership boundaries you'll co-design
  • Debug production-grade issues spanning management APIs, daemons, kernel drivers, BMC firmware, and hardware
  • Help shape what "manageable at scale" means for a new hardware platform, from single node to full rack to cluster
Who You Are
  • BS or higher in Computer Science, Electrical Engineering, or equivalent practical experience, with 8+ years in systems software — this is not a pure web-services role; deep low-level systems experience is required
  • Strong hands-on Linux systems development experience, including low-level userspace software; comfortable reading and debugging kernel driver and daemon code
  • Strong programming skills in C plus a systems language suited to services and tooling (Go, Rust, C++, and/or Python)
  • Experience designing and building HTTP/REST APIs and CLI tools for hardware or infrastructure management
  • Solid understanding of how the pieces underneath your APIs actually work — device drivers, telemetry paths, PCIe device behavior, BMC-managed subsystems — and the instinct to go look when something misbehaves
  • Experienced debugging across API, daemon, kernel, firmware, and hardware boundaries
  • Comfortable working with firmware engineers to align host-side and BMC-side management capabilities behind common interfaces
  • Self-driven and pragmatic: able to stand up a working management endpoint against brand-new hardware with minimal specification
Bonus Points If You Have
  • Experience with Redfish, OpenBMC, gNMI, IPMI, or other datacenter hardware management standards
  • Cluster/fleet management experience for GPU or accelerator infrastructure
  • Experience with hardware bring-up, lab automation, or manufacturing/qualification test infrastructure
  • Familiarity with firmware update orchestration, secure boot, or attestation flows
Compensation

The US base salary for this full-time position is determined based on a variety of factors including role, experience, location, job related skills, and relevant education and training. Career length is only a guideline for compensation.

  • Early Career - $120,000 - $250,000 + equity
What We Offer
  • A Stake in our successA flexible cash equity compensation mix that fits your needs
  • Heath & WellnessCompany subsidized Health, Dental, Vision, and Life insurance; Pre-tax Health Savings Accounts with generous company contribution (even if you don’t)
  • Time To Recharge4 weeks paid time off (accrued), 12 company holidays, and 3 weeks remote/flexible work per year
  • Support to ParentsUp to 12 weeks of paid parental leave, regardless of your path to parenthood
  • Learning & Development$1,500 yearly towards your professional development e.g. conferences, courses, and other learning opportunities
  • Team ConnectionTeam Lunches, quarterly off-sites, and regular town halls
  • Financial Wellbeing401K and/or Roth IRA, with 5% company contribution, even if you don’t!
  • Flexible Spending AccountsPre-tax spend accounts for medical, dental/vision, dependent care, parking, and transit expenses
  • Commute On UsFor those commuting up to 1 hour, put your rideshare cost on our company card and reclaim the drive-time to get work done!
  • MatX E[x]tras$50 per month to use on the perks you care about most
  • Remote PerksWe work remotely Monday & Friday, supported by home-tech setup, and remote wifi expense reimbursement

As part of our dedication to the diversity of our team and our focus on creating an inviting and inclusive work experience, MatX is committed to a policy of Equal Employment Opportunity and will not discriminate against an applicant or employee on the basis of race, color, religion, creed, national origin or ancestry, sex, gender, gender identity, gender expression, sexual orientation, age, physical or mental disability, medical condition, marital/domestic partner status, military and veteran status, genetic information or any other legally recognized protected basis under federal, state or local laws, regulations or ordinances.

All candidates must be authorized to work in the United States and work from our offices in Mountain View Tuesdays-Thursdays.

This position requires access to information that is subject to U.S. export controls. This offer of employment is contingent upon the applicant's capacity to perform job functions in compliance with U.S. export control laws without obtaining a license from U.S. export control authorities.

MatX does not accept unsolicited resumes from individual recruiters or third-party recruiting agencies in response to job postings. No fee will be paid to third parties who submit unsolicited candidates directly to our hiring managers or People team and any resumes submitted are deemed to be the property of MatX.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

System Software Engineer, Node & Cluster Management
System Software Engineer, Node & Cluster Management

MatX Inc. • Mountain View (CA)

Hybrid
USD 250,000 - 600,000
Time off
Health insurance
Financial wellbeing
+8
System Software Engineer, Linux Kernel and Device Drivers
System Software Engineer, Linux Kernel and Device Drivers

MatX Inc. • Mountain View (CA)

Hybrid
USD 250,000 - 600,000
4 weeks PTO
Holidays & remote time
Medical insurance
+3
Platform Technical Project Manager, Rack-Scale AI Systems
Platform Technical Project Manager, Rack-Scale AI Systems

MatX • Mountain View (WY)

Hybrid
USD 120,000 - 600,000
Health & Wellness
Time To Recharge
Learning & Development
+4
System Software Engineer, Linux Kernel and Device Drivers
System Software Engineer, Linux Kernel and Device Drivers

MatX • Mountain View (CA)

On-site
USD 250,000 - 475,000
Equity
Health insurance
Paid time off
+6
System Software Engineer
System Software Engineer

MatX Inc. • Mountain View (CA)

Hybrid
USD 200,000 - 500,000
4 weeks PTO
12 company holidays
up to 3 weeks remote work
+3
Technical Program Manager, Rack-Scale API Systems
Technical Program Manager, Rack-Scale API Systems

MatX Inc. • Mountain View (CA)

Hybrid
USD 160,000 - 600,000
4 weeks PTO
12 company holidays
Up to 3 weeks remote work
+7
BMC Firmware Engineer
BMC Firmware Engineer

MatX Inc. • Mountain View (CA)

Hybrid
USD 160,000 - 600,000
4 weeks PTO
Remote work up to 3 weeks
Medical insurance
+4
BMC Firmware Engineer
BMC Firmware Engineer

MatX • Mountain View (WY)

Hybrid
USD 250,000 - 475,000
Equity compensation
Health, Dental, Vision, Life insurance
Remote work 3 days per week
+3
Runtime Engineer
Runtime Engineer

MatX Inc. • Mountain View (CA)

Hybrid
USD 160,000 - 475,000
Time off: 4 weeks PTO + 12 holidays +
Health: Company-subsidized Medical + D
Financial Wellbeing: 401K with company
+5
SOC Micro-Architect and RTL Designer
SOC Micro-Architect and RTL Designer

MatX • Mountain View (WY)

On-site
USD 120,000 - 600,000
Equity
Health insurance
Paid time off
+6