Member of Technical Staff - Machines

Mixpeek

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Modal is seeking engineers to design, build, and sustain the hardware fleet that powers its serverless platform. You will work on the machines layer, provisioning bare metal and cloud hosts, imaging, monitoring, and repairing them to keep production running smoothly.

You will automate capacity from various hardware providers, benchmark hosts, configure GPUs, RDMA, networking, and storage, and ensure rapid recovery from failures.

Qualifications

  • 5+ years of experience writing high-quality production code.
  • Experience operating fleets of physical hardware or building the related control planes.
  • Strong cloud skills and Linux kernel/driver knowledge.
  • Ability to debug across layers and respond to production incidents.
  • Willingness to participate in on-call rotations.

Responsibilities

  • Design, build, and maintain high-performance systems for Modal's serverless platform.
  • Own the machines layer: provisioning, images, monitoring, and repairs of hardware fleets.
  • Automate integration of new CPU/GPU/storage servers and hardware providers.
  • Audit, benchmark, and maintain machine images, GPUs, RDMA, networking, and storage.
  • Develop automation to keep the fleet healthy with minimal human intervention.
  • Diagnose issues from kernel panics to container runtime on new architectures.

Skills

Cloud computing
Linux kernel & networking
Debugging across layers
On-call readiness
Go programming

Tools

BMC/IPMI
PXE/network boot
Bare metal provisioning
Kernel images & firmware

Job description

About Us:

AI needs a new infrastructure layer. We're building it at Modal.

Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g.,Seaborn,Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role:

We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll automate the integration of new capacity from a growing set of hardware providers; from auditing and benchmarking hosts and clusters, to maintaining our machine images, configuring GPUs, RDMA, networking, and storage, and getting machines into production. You'll build the automation that keeps the fleet healthy without human intervention: detecting bad GPUs, thermals, and disks. You'll dig into whatever is between the hardware and the software that runs on top of it, whether that is a kernel panic, a broadcast storm during boot, or getting our container runtime to run on new architectures and platforms.

Requirements:
  • 5+ years of experience writing high-quality production code

  • Experience operating fleets of physical hardware (bare metal provisioning, BMC/IPMI, PXE or network boot, firmware) or building the control planes that manage them (the more challenges you've worked through, the better)

  • Strong cloud skills

  • Strong knowledge of low-level operating system foundations (Linux kernel, drivers, networking, file systems, containers, etc.)

  • Effective at debugging across layers, from BGP flapping, Linux RPS, and vBIOS bugs to a Python control-plane service

  • Willingness to step into the thick of it with our on-call rotation and respond to production incidents

Nice-to-Haves:
  • Experience with GPUs and the NVIDIA software stack in production (drivers, health monitoring, XIDs, RDMA/NVLink)

  • Prior experience with Go

Key Things the Team Is Working On:
  • Automatic remediation of unhealthy machines (power cycling, reimaging, GPU recovery) to maximize uptime and minimize operator toil.

  • Automatic integration of new CPU, GPU, and storage servers into the fleet while managing hardware and network heterogeneity.

  • Network health monitoring and reliability across many datacenters, and standardization of bare metal network configuration.

  • Automatic hardware acceptance testing and benchmarking (CPU, disk, GPU, interconnect, network).

  • Custom network bootloader, machine image pipeline, and kernel and firmware management across the fleet.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Machines
Member of Technical Staff - Machines

Linuxconfig • Northern (KY), New York (NY)

Hybrid
USD 180,000 - 230,000
Member of Technical Staff - Machines
Member of Technical Staff - Machines

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - Machines
Member of Technical Staff - Machines

Creandum • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff - Lead, Machines
Member of Technical Staff - Lead, Machines

Socket.dev • New York (NY)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Lead, Machines
Member of Technical Staff - Lead, Machines

Linuxconfig • Northern (KY), New York (NY)

Hybrid
USD 190,000 - 230,000
Member of Technical Staff - Lead, Machines
Member of Technical Staff - Lead, Machines

Creandum • New York (NY), Northern (KY)

Hybrid
USD 180,000 - 280,000
Member of Technical Staff - Lead, Machines
Member of Technical Staff - Lead, Machines

Mixpeek • New York (NY)

On-site
USD 210,000 - 270,000
Member of Technical Staff - Platform Engineering
Member of Technical Staff - Platform Engineering

Modal • New York (NY)

On-site
USD 150,000 - 190,000
Member of Technical Staff - Lead, Storage
Member of Technical Staff - Lead, Storage

Socket.dev • San Francisco (CA)

On-site
USD 200,000 - 320,000
Member of Technical Staff - Storage
Member of Technical Staff - Storage

Mixpeek • San Francisco (CA)

On-site
USD 180,000 - 240,000
Remote-friendly options