ML Model Performance Engineer - Inference and Acceleration

Baseten

New York (NY)

On-site

USD 200,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

100% coverage of medical, dental, and vision insurance
Generous PTO policy
Paid parental leave
Company-facilitated 401(k)
Exposure to various ML startups

Job summary

A leading AI platform provider in New York is looking for a Software Engineer focused on ML performance. This role involves optimizing ML model inference using cutting-edge techniques and collaborating with a diverse team. Ideal candidates have a degree in Computer Science or related fields, along with experience in Python or C++. The position offers competitive compensation along with comprehensive benefits including equity, health insurance, and generous PTO.

Qualifications

  • Bachelor’s, Master’s, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.
  • Experience with programming languages, such as Python or C++.
  • Strong familiarity with ML libraries, especially PyTorch and TensorRT.

Responsibilities

  • Implement and productionize cutting-edge techniques for ML model inference.
  • Debug ML performance issues in underlying codebases.
  • Collaborate with a diverse team to design and implement innovative solutions.

Skills

Experience with Python
Experience with C++
Familiarity with LLM optimization techniques
Strong familiarity with PyTorch
Understanding of GPU architecture

Education

Bachelor’s, Master’s, or Ph.D. degree in Computer Science or related field

Tools

Docker
Kubernetes
CUDA

Job description

A leading AI platform provider in New York is looking for a Software Engineer focused on ML performance. This role involves optimizing ML model inference using cutting-edge techniques and collaborating with a diverse team. Ideal candidates have a degree in Computer Science or related fields, along with experience in Python or C++. The position offers competitive compensation along with comprehensive benefits including equity, health insurance, and generous PTO.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
ML Acceleration Engineering Lead
ML Acceleration Engineering Lead

Anthropic • New York (NY)

On-site
USD 425,000 - 560,000
Senior ML Performance Engineer: LLM Benchmarking & GPU
Senior ML Performance Engineer: LLM Benchmarking & GPU

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary
Equity and bonus opportunities
Medical, dental, and vision coverage
+2
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Staff ML Performance Engineer — Scalable Inference & CUDA
Staff ML Performance Engineer — Scalable Inference & CUDA

Modal • New York (NY)

On-site
USD 120,000 - 160,000
Engineering Manager, ML Inference & Scale
Engineering Manager, ML Inference & Scale

Anthropic • San Francisco (CA)

Hybrid
USD 425,000 - 560,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 150,000 - 200,000
Model API Engineer - High-Performance Inference
Model API Engineer - High-Performance Inference

Baseten • New York (NY)

On-site
USD 180,000 - 360,000