Machine Learning Engineer

IC Resources

San Francisco (CA)

On-site

USD 200,000 - 290,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k)
Unlimited PTO
Modern engineering workspace
Equity participation

Job summary

IC Resources seeks an engineer to accelerate production AI systems, focusing on speed and efficiency of large language model inference. You will optimize GPU-heavy pipelines and scale distributed GPU environments in a fast-moving, innovative AI company.

You will work with research and infrastructure teams to translate cutting-edge models into reliable production systems, improving latency and throughput while leveraging modern hardware.

Qualifications

  • Degree in Computer Science, Electrical Engineering, or related discipline.
  • Strong Python and C++ (CUDA preferred) programming skills.
  • Knowledge of modern LLM serving technologies (vLLM, SGLang, PyTorch, etc.).
  • Understanding of GPU architecture and parallel computing.
  • Experience with model serving, distributed inference, quantization, or batching strategies.
  • Strong profiling, debugging, and systems optimization skills.

Responsibilities

  • Improve speed and efficiency of production AI systems.
  • Evaluate performance bottlenecks and build benchmarking tools.
  • Optimize inference pipelines and scale distributed GPU environments.
  • Collaborate with research and infrastructure teams to deploy new modeling techniques.
  • Continuously improve latency, throughput, and hardware utilization.

Skills

Python
C++
CUDA
vLLM
SGLang
PyTorch
GPU architecture
Distributed inference
Profiling
Debugging

Education

Degree in Computer Science or Electrical Engineering

Tools

CUDA toolkit

Job description

Compensation: $200K–$290K + Equity

Location: San Jose, California

An innovative AI company is expanding its systems engineering team and is searching for an engineer who enjoys making large language models faster, more scalable, and more efficient in production.

If your interests include GPU optimization, distributed computing, and extracting every ounce of performance from modern hardware, this role offers the opportunity to work on some of today's most demanding inference challenges.

Responsibilities

Your primary focus will be improving the speed and efficiency of production AI systems. You'll evaluate performance bottlenecks, build benchmarking tools, optimize inference pipelines, and help scale distributed GPU environments.

Working closely with research and infrastructure teams, you'll transform new modeling techniques into reliable production systems while continuously improving latency, throughput, and hardware utilization.

We're Looking For Someone Who Has

  • A degree in Computer Science, Electrical Engineering, or a related discipline
  • Strong Python and C++ (CUDA preferred) programming skills
  • Knowledge of modern LLM serving technologies such as vLLM, SGLang, PyTorch, or comparable frameworks
  • Understanding of GPU architecture and parallel computing
  • Experience with model serving, distributed inference, quantization, or batching strategies
  • Strong profiling, debugging, and systems optimization skills

Additional Experience That Stands Out

  • Ray or similar distributed computing frameworks
  • Performance tuning at the systems or kernel level
  • High-performance computing environments

What's Offered

  • Competitive salary with meaningful equity participation
  • 401(k)
  • Unlimited PTO
  • Modern engineering workspace with premium employee amenities
  • Opportunity to solve technically complex AI infrastructure challenges alongside a highly experienced engineering team
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Strativ Group • Palo Alto (CA)

On-site
USD 500,000 - 600,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
AI/ML Engineer
AI/ML Engineer

BitWords Inc. • San Francisco (CA)

Hybrid
USD 140,000 - 200,000
Competitive salary
Equity package
Health, dental, vision insurance
+6
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
System Software Engineer - AI
System Software Engineer - AI

Delos Data • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits
System Software Engineer - AI
System Software Engineer - AI

Entrada Ventures • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits package