Senior On-Device ML Engineer – Mobile Inference Expert

Unity

California (MO)

On-site

USD 180,000 - 240,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Unity seeks a Senior ML Engineer for On-Device & Mobile AI to optimize, deploy, and maintain high-performance inference stacks across mobile, desktop, and embedded hardware. You will push efficient multi-modal models—transformers, diffusion, and VLMs—into shipping features with tight latency and memory budgets.

You will work across WebGPU-based runtimes and native options, tuning kernels, and implementing quantization and pruning to meet power and speed targets on diverse SKUs, while

Qualifications

  • 5+ years in software/ML engineering with on-device or real-time systems.
  • Production deployment of transformer- and/or diffusion-based models on mobile/embedded hardware.
  • Hands-on experience with major inference runtimes (ORT Web, CoreML, TFLite, ExecuTorch).
  • Low-level performance engineering with GPU/compute APIs and profiling tools.

Responsibilities

  • Own the optimization and deployment pipeline for on-device models: export, graph transformation, operator fusion, memory planning, hardware tuning.
  • Implement quantization (INT4/INT8/FP16), weight sharing, pruning, and distillation to meet latency and memory budgets.
  • Write and tune WebGPU compute shaders and native kernels; profile with browser/platform tools and fix bottlenecks.
  • Collaborate with ML researchers to productionize CV and multi-modal architectures for on-device usage.

Skills

On-device inference
Model optimization
WebGPU/WGSL/Metal/Vulkan/CUDA
Profiling & debugging
Python for export pipelines
TypeScript/JavaScript familiarity
Transformer/diffusion deployment
Runtime scheduling

Tools

ONNX Runtime Web
CoreML
TFLite
ExecuTorch
WebGPU
WGSL
Metal
Vulkan
CUDA

Job description

Unity seeks a Senior ML Engineer for On-Device & Mobile AI to optimize, deploy, and maintain high-performance inference stacks across mobile, desktop, and embedded hardware. You will push efficient multi-modal models—transformers, diffusion, and VLMs—into shipping features with tight latency and memory budgets.

You will work across WebGPU-based runtimes and native options, tuning kernels, and implementing quantization and pruning to meet power and speed targets on diverse SKUs, while

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

On-Device ML Engineer — Mobile AI & Optimization
On-Device ML Engineer — Mobile AI & Optimization

Unity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff ML Engineer: On-Device AI for Real-Time Games
Staff ML Engineer: On-Device AI for Real-Time Games

Unity • Mountain View (CA)

On-site
USD 180,000 - 280,000
On-Device ML Engineer — Mobile AI Optimizations
On-Device ML Engineer — Mobile AI Optimizations

Unity Technologies • United States

Remote
USD 218,000 - 284,000
Health insurance
Stock options
Retirement plan
+2
Principal On-Device AI Engineer for Real-Time Game Inference
Principal On-Device AI Engineer for Real-Time Game Inference

Unity • Mountain View (CA)

On-site
USD 190,000 - 270,000
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity • California (MO)

On-site
USD 180,000 - 240,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity • Mountain View (CA)

On-site
USD 180,000 - 280,000
Principal On-Device AI Engineer – Real-Time Game Inference
Principal On-Device AI Engineer – Real-Time Game Inference

Unity Technologies • Mountain View (CA)

On-site
USD 278,000 - 418,000
Comprehensive health, life, and disability insurance
Employee stock ownership
Competitive retirement/pension plans
+1
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3