Software Engineer - ML Model Performance

Baseten

San Francisco (CA)

On-site

USD 150,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation with equity
100% medical, dental, and vision insurance
Generous PTO policy
Paid parental leave
Company‑facilitated 401(k)

Job summary

A leading AI technology company based in San Francisco is seeking a motivated Software Engineer to enhance ML performance. This role involves implementing advanced techniques for ML model inference, collaborating in a dynamic team, and tackling performance optimization challenges. Ideal candidates will have a relevant degree, programming experience (Python or C++), and familiarity with LLM optimization. This position offers competitive compensation, comprehensive health coverage, and generous PTO policies.

Qualifications

  • Degree in Computer Science, Engineering, Mathematics, or related field.
  • Experience with general‑purpose programming languages.
  • Familiarity with LLM optimization techniques.

Responsibilities

  • Implement and productionize techniques for ML model inference.
  • Deep dive into codebases to debug ML performance issues.
  • Collaborate with a diverse team to design solutions.

Skills

Experience with Python or C++
Familiarity with LLM optimization techniques
Strong familiarity with PyTorch
Understanding of GPU architecture

Education

Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field

Tools

TensorRT
CUDA
Docker
Kubernetes

Job description

A leading AI technology company based in San Francisco is seeking a motivated Software Engineer to enhance ML performance. This role involves implementing advanced techniques for ML model inference, collaborating in a dynamic team, and tackling performance optimization challenges. Ideal candidates will have a relevant degree, programming experience (Python or C++), and familiarity with LLM optimization. This position offers competitive compensation, comprehensive health coverage, and generous PTO policies.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Model Performance Engineer - Inference and Acceleration
ML Model Performance Engineer - Inference and Acceleration

Baseten • New York (NY)

On-site
USD 200,000 - 275,000
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 150,000 - 200,000
Engineering Manager, ML Inference & Scale
Engineering Manager, ML Inference & Scale

Anthropic • San Francisco (CA)

Hybrid
USD 425,000 - 560,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Staff ML Infrastructure & Performance Engineer
Staff ML Infrastructure & Performance Engineer

Embedding VC • San Mateo (CA)

On-site
USD 120,000 - 150,000
Staff ML Performance Engineer — High-Throughput Inference
Staff ML Performance Engineer — High-Throughput Inference

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Competitive salary and equity
Visa sponsorship
Health, dental, and vision insurance
+1
Senior ML Performance Engineer: LLM Benchmarking & GPU
Senior ML Performance Engineer: LLM Benchmarking & GPU

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary
Equity and bonus opportunities
Medical, dental, and vision coverage
+2
Staff ML Inference Engineer — Model Efficiency (Remote)
Staff ML Inference Engineer — Model Efficiency (Remote)

Jaide Health • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+2