Senior ML Performance Engineer - GPU & Inference

Modal Labs

New York (NY)

On-site

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job description

About Us:

Modal provides the infrastructure foundation for AI teams. With instant GPU access, sub-second container startups, and native storage, Modal makes it simple to train models, run batch jobs, and serve low-latency inference. We have thousands of customers who rely on us for production AI workloads, including Lovable, Scale AI, Substack, and Suno.

We're a fast-growing team based out of NYC, SF, and Stockholm. We've hit 9-figure ARR and recently raised a Series B at a $1.1B valuation. Our investors include Lux Capital, Redpoint Ventures, Amplify Partners, and Elad Gil.

Working at Modal means joining one of the fastest-growing AI infrastructure organizations at an early stage, with many opportunities to grow within the company. Our team includes creators of popular open-source projects (e.g. Seaborn, Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role

We are looking for strong engineers with experience in making ML systems performant at scale. If you are interested in contributing to open-source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you!

Requirements
  • 5+ years of experience writing high-quality, high-performance code.

  • Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).

  • Familiarity with Nvidia GPU architecture and CUDA.

  • Experience with ML performance engineering (tell us a story about boosting GPU performance — debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead, etc).

  • Nice-to-have: familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Modal Labs • New York (NY)

On-site
USD 150,000 - 190,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Modal • New York (NY)

On-site
USD 120,000 - 160,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Modal • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff ML Performance Engineer - GPU & Inference
Staff ML Performance Engineer - GPU & Inference

Modal • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior ML Performance Engineer
Senior ML Performance Engineer

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - Product (Backend)
Member of Technical Staff - Product (Backend)

Modal Labs • New York (NY)

On-site