Embedded ML Runtime Optimization Engineer

Applied Intuition

Sunnyvale (CA)

Hybrid

USD 159,000 - 199,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Applied Intuition in Sunnyvale, CA is seeking a software engineer specializing in optimizing ML models for production-grade embedded runtime environments. You will influence the entire ML framework stack across PyTorch, JAX, ONNX, TensorRT, CUDA, XLA and Triton.

You will collaborate with ML engineers and software teams to implement model pruning, quantization, and deployment strategies on memory-constrained embedded compute platforms, delivering efficient, low latency inference.

Qualifications

  • Bachelor's degree in Electrical Engineering or Computer Science
  • 3+ years of experience with ML accelerators, GPU, CPU, SoC architecture and micro-architecture
  • Strong software development skills with the focus on embedded programming
  • Experience profiling and optimizing model performance on embedded compute platforms
  • Experience in working with deep learning frameworks (e.g., PyTorch, JAX, ONNX, etc.)

Responsibilities

  • Drive ML performance optimization on multiple technologies for on-road and off-road ADAS / AD stacks targeting deployment on a variety of embedded compute platforms
  • Develop compute usage strategies to optimize efficiency and latency of model inference for compute boards selected by our customers
  • Work on model pruning and quantization, and support deployment on memory constrained platforms
  • Collaborate closely with ML engineers and software developers on technical efforts to find and optimize efficient model architecture solutions
  • Set up methodologies to profile the model performance on target embedded compute platforms and identify performance bottlenecks as part of stack integration

Skills

Embedded programming
ML optimization
Model performance profiling

Education

Bachelor's degree in Electrical Engineering or Computer Science
B.Sc. in Computer Science, Mathematics, Physics or related field

Tools

PyTorch
JAX
ONNX
CUDA
TensorRT
XLA
Triton

Job description

Applied Intuition in Sunnyvale, CA is seeking a software engineer specializing in optimizing ML models for production-grade embedded runtime environments. You will influence the entire ML framework stack across PyTorch, JAX, ONNX, TensorRT, CUDA, XLA and Triton.

You will collaborate with ML engineers and software teams to implement model pruning, quantization, and deployment strategies on memory-constrained embedded compute platforms, delivering efficient, low latency inference.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Embedded ML Inference & Optimization Engineer
Embedded ML Inference & Optimization Engineer

Applied Intuition Inc. • Sunnyvale (CA)

Hybrid
USD 159,000 - 200,000
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Remote Edge AI Engineer - On-Device ML Optimization
Remote Edge AI Engineer - On-Device ML Optimization

Bright Vision Technologies • South Windsor (CT)

On-site
USD 100,000 - 155,000
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
ML Runtime Optimization Engineer
ML Runtime Optimization Engineer

Applied Intuition Inc. • Sunnyvale (CA)

Hybrid
USD 159,000 - 200,000
ML Runtime Optimization Engineer
ML Runtime Optimization Engineer

Applied Intuition • Sunnyvale (CA)

Hybrid
USD 159,000 - 199,000
On-Device AI Engineer for Android Automotive
On-Device AI Engineer for Android Automotive

Applied Intuition • Sunnyvale (CA)

Hybrid
USD 150,000 - 250,000
Equity in options/RSUs
Comprehensive health insurance
401k with match
+1
AI Performance Engineer: Scale ML Workloads at Datacenter
AI Performance Engineer: Scale ML Workloads at Datacenter

Applied Intuition • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000