Senior ML Performance Engineer

well-funded deeptech startup

California (MO)

On-site

USD 200,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Join a well-funded deeptech startup in California as a Senior ML Performance Engineer. You will lead the optimization of ML inference pipelines and improve model performance for innovative clients.

This role requires strong skills in Python, TensorFlow, and PyTorch, as well as expertise in performance optimization techniques. The ideal candidate is eager to work on cutting-edge ML projects in an energetic environment with potential for significant impact.

Qualifications

  • Strong proficiency in Python and ML frameworks such as TensorFlow and PyTorch.
  • Deep understanding of ML algorithms and architectures.
  • Expertise in performance optimization techniques including profiling and quantization.

Responsibilities

  • Conduct in-depth performance profiling and analysis of ML models.
  • Design and implement efficient ML inference pipelines.
  • Collaborate with infrastructure teams to optimize configurations.

Skills

Python
TensorFlow
PyTorch
Performance optimization techniques
Kubernetes
AWS
GCP
Azure

Job description

Senior ML Performance Engineer | $200-250K salary + equity | Early-stage ML Infrastructure. 30-50 employees, $50MM+ in funding

We’re seeking a seasoned ML Performance Optimization Specialist to spearhead the development and deployment of high-performance, scalable ML inference pipelines for a key early-stage company. You’ll optimize model performance, reduce latency, and maximize throughput for some of the most innovative companies in the world.

Key Responsibilities
  • Model Optimization: Conduct in-depth performance profiling and analysis of ML models, identifying and eliminating bottlenecks.
  • Pipeline Engineering: Design and implement efficient ML inference pipelines, leveraging technologies like TensorFlow Serving, TorchServe, and NVIDIA Triton Inference Server.
  • Infrastructure Optimization: Collaborate with infrastructure teams to optimize hardware and software configurations for optimal ML performance, including GPU acceleration, distributed training, and model quantization.
  • Performance Benchmarking: Develop and maintain rigorous performance benchmarks to measure and track improvements.
  • Experimentation: Explore cutting‑edge techniques like model quantization, pruning, and knowledge distillation to further enhance performance.
Required Skills and Experience
  • Strong proficiency in Python and ML frameworks (TensorFlow, PyTorch)
  • Deep understanding of ML algorithms and architectures
  • Expertise in performance optimization techniques (profiling, quantization, pruning, etc.)
  • Hands‑on experience with container orchestration platforms (Kubernetes)
  • Proficiency in cloud platforms (AWS, GCP, Azure)
  • Strong problem‑solving and analytical skills
Preferred Qualifications
  • Experience with ML hardware acceleration (GPUs, TPUs)
  • Knowledge of distributed training frameworks (Horovod, DDP)
  • Familiarity with MLIR and compiler optimization techniques

If you’re passionate about pushing the boundaries of ML performance and eager to work on cutting‑edge projects, we encourage you to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Senior ML Performance Engineer — Ultra-Fast Inference + Equity
Senior ML Performance Engineer — Ultra-Fast Inference + Equity

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior ML Performance Engineer - GPU & Inference
Senior ML Performance Engineer - GPU & Inference

Modal Labs • New York (NY)

On-site
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior ML Performance Engineer
Senior ML Performance Engineer

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary
Equity and bonus opportunities
Medical, dental, and vision coverage
+2