Lead Node Systems Engineer — Frontier Inference HPC

Etched

San Jose (CA)

On-site

USD 200,000 - 260,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical coverage
Dental coverage
Vision coverage
Housing subsidy
Relocation assistance
Daily meals in the office
Unlimited compute budget

Job summary

Etched is hiring a Node Systems Lead to join its San Jose-based Supercomputing team. You will own the software layer for each node, optimize host software, and coordinate interfaces between rack components to deliver peak throughput, latency, and reliability for inference workloads.

You will lead a team, shape technical direction for system configuration, and collaborate with hardware, firmware, and inference software teams to deploy scalable, high-performance clusters.

Qualifications

  • Strong experience developing and debugging production systems software in C, C++ or Rust on Linux.
  • Experience leading technical projects and mentoring engineers while staying hands on.
  • Deep understanding of operating systems fundamentals, including scheduling, concurrency, memory management and I/O.
  • Demonstrated ability to investigate performance problems on real hardware and validate improvements through profiling and measurement.
  • Experience taking system software or performance improvements from prototype through production deployment.
  • Understanding of hardware/software interactions and the ability to debug across application, kernel and device boundaries.
  • Ability to work closely with hardware and software teams and translate workload requirements into concrete system changes.

Responsibilities

  • Lead and develop the Node Systems team, setting priorities and giving engineers clear ownership.
  • Set the technical direction for host software, system configuration and rack component interfaces.
  • Guide the team’s performance work across CPU scheduling, memory, networking and host-to-accelerator communication.
  • Ensure system optimizations become tested, maintainable software and configurations ready for deployment.
  • Lead the development of rack simulation and diagnostic capabilities that support platform development and debugging.
  • Partner with inference software, firmware, and hardware teams to architect our system design for current and next-gen products
  • Align with Fleet Software on the configurations and interfaces needed to manage and monitor deployed systems at the cluster-level
  • Work with manufacturing and test engineering to establish the software baseline and test coverage needed to ship reliable systems.
  • Stay close to the implementation through design reviews, code contributions and hands-on debugging.

Skills

C
C++
Rust
Linux
Leadership
OS fundamentals
Performance profiling
System deployment
Hardware-software debug

Tools

PCIe
DMA
RDMA
Device drivers

Job description

Etched is hiring a Node Systems Lead to join its San Jose-based Supercomputing team. You will own the software layer for each node, optimize host software, and coordinate interfaces between rack components to deliver peak throughput, latency, and reliability for inference workloads.

You will lead a team, shape technical direction for system configuration, and collaborate with hardware, firmware, and inference software teams to deploy scalable, high-performance clusters.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Node Systems Architect for Frontier Inference
Lead Node Systems Architect for Frontier Inference

Linuxconfig • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Medical benefits
Housing subsidy
Relocation support
+3
Node Systems Lead
Node Systems Lead

Etched • San Jose (CA)

On-site
USD 200,000 - 260,000
Medical coverage
Dental coverage
Vision coverage
+4
Node Systems Lead
Node Systems Lead

Linuxconfig • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Medical benefits
Housing subsidy
Relocation support
+3
Supercomputing Engineer
Supercomputing Engineer

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical coverage
Dental coverage
Vision coverage
+4
Head of Supercomputing
Head of Supercomputing

The Consensus • San Jose (CA)

On-site
USD 260,000 - 520,000
Medical, dental, and vision packages
Housing subsidy near Santana Row
Relocation support to San Jose
+3
Fleet Platform Lead for Scalable Inference
Fleet Platform Lead for Scalable Inference

etched • San Jose (CA)

On-site
USD 170,000 - 210,000
Medical, dental, and vision packages
Housing subsidy of $2,500 per month
Relocation support to San Jose
+3
Staff Datacenter Networking Engineer: Frontier AI GPU
Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Cluster-Scale AI Systems Engineer
Cluster-Scale AI Systems Engineer

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical coverage
Dental coverage
Vision coverage
+4
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Hardware Systems Engineer, AI Accelerators (100G PCIe)
Hardware Systems Engineer, AI Accelerators (100G PCIe)

The Consensus • San Jose (CA)

On-site
USD 150,000 - 275,000
Medical, dental, and vision coverage
Housing subsidy near office
Relocation support
+3