ML Model Performance Engineer - Inference and Acceleration

Baseten

New York (NY)

On-site

USD 200,000 - 275,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

100% coverage of medical, dental, and vision insurance
Generous PTO policy
Paid parental leave
Company-facilitated 401(k)
Exposure to various ML startups

Job summary

A leading AI platform provider in New York is looking for a Software Engineer focused on ML performance. This role involves optimizing ML model inference using cutting-edge techniques and collaborating with a diverse team. Ideal candidates have a degree in Computer Science or related fields, along with experience in Python or C++. The position offers competitive compensation along with comprehensive benefits including equity, health insurance, and generous PTO.

Qualifications

  • Bachelor’s, Master’s, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.
  • Experience with programming languages, such as Python or C++.
  • Strong familiarity with ML libraries, especially PyTorch and TensorRT.

Responsibilities

  • Implement and productionize cutting-edge techniques for ML model inference.
  • Debug ML performance issues in underlying codebases.
  • Collaborate with a diverse team to design and implement innovative solutions.

Skills

Experience with Python
Experience with C++
Familiarity with LLM optimization techniques
Strong familiarity with PyTorch
Understanding of GPU architecture

Education

Bachelor’s, Master’s, or Ph.D. degree in Computer Science or related field

Tools

Docker
Kubernetes
CUDA

Job description

A leading AI platform provider in New York is looking for a Software Engineer focused on ML performance. This role involves optimizing ML model inference using cutting-edge techniques and collaborating with a diverse team. Ideal candidates have a degree in Computer Science or related fields, along with experience in Python or C++. The position offers competitive compensation along with comprehensive benefits including equity, health insurance, and generous PTO.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Competitive compensation with equity
100% medical, dental, and vision insurance
Generous PTO policy
+2
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
ML Acceleration Engineering Lead
ML Acceleration Engineering Lead

Anthropic • New York (NY)

On-site
USD 425,000 - 560,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
+2
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Model API Engineer - High-Performance Inference
Model API Engineer - High-Performance Inference

Baseten • New York (NY)

On-site
USD 180,000 - 360,000
Competitive compensation including equity
100% medical, dental, and vision insurance
Generous PTO including Winter Break
+2
Engineering Manager, GPU ML Accelerator
Engineering Manager, GPU ML Accelerator

Menlo Ventures • New York (NY)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Staff ML Inference Engineer — Model Efficiency (Remote)
Staff ML Inference Engineer — Model Efficiency (Remote)

Jaide Health • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+2
Performance Engineer, Large-Scale ML Systems
Performance Engineer, Large-Scale ML Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Competitive compensation
Optional equity donation matching
Generous vacation and parental leave
+1
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,200 - 262,200
Health insurance
401(k) plan
Parental leave
+2