Remote Inference Optimization Engineer

Modular Mailing Systems, Inc.

Los Altos (CA)

Hybrid

USD 198,000 - 286,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Premier insurance plans
5% 401k matching
Flexible paid time off
Stock options
Team building events

Job summary

Modular Mailing Systems, Inc. is seeking an experienced Performance Engineer to optimize LLM inference on their cloud platform. This pivotal role involves building optimization infrastructures and collaborating with teams to enhance performance across GPUs and ASICs.

The ideal candidate will have over 5 years of experience in distributed systems, a track record in software tools development, and a collaborative mindset. Flexible hybrid work options are available. Exceptional benefits and competitive compensation packages are also included.

Qualifications

  • 5+ years of experience in distributed systems or performance engineering.
  • A track record of building reusable software tools and libraries.
  • Creativity and curiosity in solving complex problems.

Responsibilities

  • Build the optimization platform for LLMs on Modular Cloud.
  • Shape technical direction for LLM performance.
  • Partner with GTM team to deliver customized LLM inference.
  • Publish blog posts on LLM inference optimization.

Skills

Distributed systems
Performance engineering
Software tools development
Technical leadership
Complex problem solving

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Cloud native technologies

Job description

Modular Mailing Systems, Inc. is seeking an experienced Performance Engineer to optimize LLM inference on their cloud platform. This pivotal role involves building optimization infrastructures and collaborating with teams to enhance performance across GPUs and ASICs.

The ideal candidate will have over 5 years of experience in distributed systems, a track record in software tools development, and a collaborative mindset. Flexible hybrid work options are available. Exceptional benefits and competitive compensation packages are also included.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior Remote LLM Inference Optimization Lead
Senior Remote LLM Inference Optimization Lead

Dragonfly Digital Management, LLC (Dragonfly Capital) • United States

On-site
USD 140,000 - 210,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Inference Runtime Engineer for LLMs & Diffusion
Inference Runtime Engineer for LLMs & Diffusion

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Inference Optimization Engineer United States - Remote · Remote
Inference Optimization Engineer United States - Remote · Remote

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Performance Engineer: AI Inference & Systems (LLM)
Performance Engineer: AI Inference & Systems (LLM)

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
Comprehensive health benefits
Flexible work arrangements