Embedded ML Inference & Optimization Engineer

Applied Intuition Inc.

Sunnyvale (CA)

Hybrid

USD 159,053 - 199,295

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Applied Intuition Inc. is seeking a software engineer with deep expertise in optimizing ML models for production-grade embedded runtime environments.

You will work across the ML framework stack and target a range of embedded compute platforms used in on- and off-road ADAS/AD stacks. You will collaborate with ML engineers and software developers, profiling model performance, implementing pruning/quantization, and driving efficient architectures for memory-constrained devices.

Qualifications

  • Bachelors in Electrical Engineering or Computer Science, or related field.
  • 3+ years of experience with ML accelerators, GPU/CPU/SoC architecture.
  • Strong software development skills with embedded programming focus.
  • Experience profiling and optimizing model performance on embedded compute platforms.
  • Experience working with deep learning frameworks (PyTorch, JAX, ONNX).

Responsibilities

  • Drive ML performance optimization across technologies for embedded ADAS/AD stacks.
  • Develop compute usage strategies to optimize inference latency and efficiency.
  • Work on model pruning and quantization for memory-constrained platforms.
  • Collaborate with ML engineers and software developers to optimize model architectures.
  • Set up profiling methodologies to identify performance bottlenecks on target hardware.

Skills

ML frameworks experience
Embedded programming
Model profiling
DL frameworks (PyTorch)

Education

Bachelors in Electrical Engineering
Bachelors in Computer Science

Job description

Applied Intuition Inc. is seeking a software engineer with deep expertise in optimizing ML models for production-grade embedded runtime environments.

You will work across the ML framework stack and target a range of embedded compute platforms used in on- and off-road ADAS/AD stacks. You will collaborate with ML engineers and software developers, profiling model performance, implementing pruning/quantization, and driving efficient architectures for memory-constrained devices.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Embedded Systems Performance Engineer (C++, ML)
Senior Embedded Systems Performance Engineer (C++, ML)

Applied Intuition • Mountain View (CA)

On-site
USD 199,295 - 264,500
Equity options
Comprehensive health insurance
401(k) retirement benefits with employer match
On-Device AI Engineer for Android Automotive
On-Device AI Engineer for Android Automotive

Applied Intuition Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 250,000
Equity
Health benefits
401k retirement
ML Runtime Optimization Engineer
ML Runtime Optimization Engineer

Applied Intuition Inc. • Sunnyvale (CA)

On-site
USD 159,053 - 199,295
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Senior On-Device ML Engineer – Mobile Inference Expert
Senior On-Device ML Engineer – Mobile Inference Expert

Unity • California (MO)

On-site
USD 180,000 - 240,000
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,200 - 262,200
Health insurance
401(k) plan
Parental leave
+2
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2