Staff Engineer, Inference Runtime — Performance & Scale

Menlo Ventures

New York (NY)

Hybrid

USD 405,000 - 485,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Menlo Ventures is looking for a Staff Engineer to lead technical efforts in their Inference organization. This senior role involves setting the technical direction for the inference runtime, ensuring it scales efficiently across various platforms.

The ideal candidate will have a strong background in systems engineering and experience with ML infrastructure, optimizing performance in distributed systems. The role is based in New York, NY, with a hybrid work policy.

Qualifications

  • Deep background in systems engineering or ML infrastructure.
  • Experienced with GPU, TPU or Trainium ecosystems.
  • Significant software engineering experience.

Responsibilities

  • Set technical direction for the inference serving stack.
  • Own and evolve the accelerator-agnostic runtime.
  • Drive efficient accelerator usage across computing platforms.

Skills

Systems engineering
Performance profiling
Distributed systems
Technical alignment
Communication skills

Education

Bachelor’s degree or equivalent

Tools

Rust
Python
Kubernetes

Job description

Menlo Ventures is looking for a Staff Engineer to lead technical efforts in their Inference organization. This senior role involves setting the technical direction for the inference runtime, ensuring it scales efficiently across various platforms.

The ideal candidate will have a strong background in systems engineering and experience with ML infrastructure, optimizing performance in distributed systems. The role is based in New York, NY, with a hybrid work policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff Engineer Inference Runtime — Flexible Hours
Senior Staff Engineer Inference Runtime — Flexible Hours

jobr.pro • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Staff Engineer — Inference Runtime Lead
Staff Engineer — Inference Runtime Lead

Anthropic • New York (NY)

Hybrid
USD 405,000 - 485,000
Staff Engineer, Inference Runtime — High-Performance AI Serving
Staff Engineer, Inference Runtime — High-Performance AI Serving

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Staff ML Performance Engineer — Scalable Inference & CUDA
Staff ML Performance Engineer — Scalable Inference & CUDA

Modal • New York (NY)

On-site
USD 120,000 - 160,000
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Staff Inference Runtime Architect (Rust/Python)
Staff Inference Runtime Architect (Rust/Python)

Anthropic • New York (NY)

Hybrid
USD 405,000 - 485,000
Generous vacation
Parental leave
Flexible working hours
+1
Senior Staff Tech Lead — Inference & ML Performance
Senior Staff Tech Lead — Inference & ML Performance

fal • San Francisco (CA)

On-site
USD 150,000 - 200,000
Lead Engineer, Inference Platform & Scale
Lead Engineer, Inference Platform & Scale

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Job stability with startup vitality
Open-source AI research
Simple, non-corporate work culture
Staff Systems Engineer — Inference Runtime Lead
Staff Systems Engineer — Inference Runtime Lead

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 405,000 - 485,000