Member of Technical Staff - ML Performance

Modal

New York (NY)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

A leading AI infrastructure company based in New York is seeking experienced engineers to enhance the performance of ML systems and contribute to open-source projects. Ideal candidates will have over 5 years of experience in writing high-quality code and familiarity with Nvidia GPU architecture and ML frameworks. This role offers opportunities for significant growth within a fast-growing team and requires in-person collaboration in NYC, San Francisco, or Stockholm.

Qualifications

  • 5+ years of experience writing high-quality, high-performance code.
  • Experience working with high-level ML frameworks.
  • Familiarity with Nvidia GPU architecture and CUDA.

Responsibilities

  • Contribute to open-source projects related to ML systems.
  • Optimize container runtime for better model performance.
  • Debug and optimize GPU performance.

Skills

High-quality code writing
Performance optimization
Experience with torch
Familiarity with Nvidia GPU architecture
Experience with CUDA

Tools

TensorRT
vLLM

Job description

About Us:

Modal provides the infrastructure foundation for AI teams. With instant GPU access, sub‑second container startups, and native storage, Modal makes it simple to train models, run batch jobs, and serve low‑latency inference. Companies like Suno, Lovable, and Substack rely on Modal to move from prototype to production without the burden of managing infrastructure.

We're a fast‑growing team based out of NYC, SF, and Stockholm. We've hit high 8‑figure ARR and recently raised a Series B at a $1.1B valuation. We have thousands of customers who rely on us for production AI workloads, including Lovable, Scale AI, Substack, and Suno.

Working at Modal means joining one of the fastest‑growing AI infrastructure organizations at an early stage, with many opportunities to grow within the company. Our team includes creators of popular open‑source projects (e.g. Seaborn, Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role:

We are looking for strong engineers with experience in making ML systems performant at scale. If you are interested in contributing to open‑source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you!

Requirements:
  • 5+ years of experience writing high‑quality, high‑performance code.
  • Experience working with torch, high‑level ML frameworks, and inference engines (vLLM or TensorRT).
  • Familiarity with Nvidia GPU architecture and CUDA.
  • Experience with ML performance engineering (tell us a story about boosting GPU performance — debugging SM occupancy issues, rewriting an algorithm to be compute‑bound, eliminating host overhead, etc).
  • Nice‑to‑have: familiarity with low‑level operating system foundations (Linux kernel, file systems, containers, etc).
  • Ability to work in‑person, in our NYC, San Francisco or Stockholm office.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Modal Labs • New York (NY)

On-site
USD 150,000 - 190,000
Senior ML Performance Engineer - GPU & Inference
Senior ML Performance Engineer - GPU & Inference

Modal Labs • New York (NY)

On-site
Member of Technical Staff - Product (Backend)
Member of Technical Staff - Product (Backend)

Modal Labs • New York (NY)

On-site
USD 140,000 - 190,000
Member of Technical Staff - Python SDK
Member of Technical Staff - Python SDK

Modal Labs • New York (NY)

On-site
USD 120,000 - 190,000
Member of Technical Staff - Product (Backend)
Member of Technical Staff - Product (Backend)

Modal • New York (NY)

On-site
USD 150,000 - 230,000
Systems Engineering Manager
Systems Engineering Manager

Modal Labs • New York (NY)

On-site
USD 140,000 - 190,000
Member of Technical Staff - Python SDK
Member of Technical Staff - Python SDK

Modal • New York (NY)

On-site
USD 140,000 - 190,000
Customer Engineer
Customer Engineer

Modal Labs • New York (NY)

On-site
Developer Relations Engineer
Developer Relations Engineer

Modal • San Francisco (CA)

On-site
USD 120,000 - 150,000
Staff Python SDK Engineer — Developer Tools
Staff Python SDK Engineer — Developer Tools

Modal Labs • New York (NY)

On-site
USD 120,000 - 190,000