On-Device AI Engineer: Model Optimization & Edge Deployment

Hyphen Connect

Boston (MA)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hyphen Connect is seeking an AI Specialist Engineer in Boston, Massachusetts, to enhance performance for on-device inference in large language and vision models. You will be responsible for compressing and optimizing these models, developing pipelines specifically for model distillation, and benchmarking performance across various architectures. Ideal candidates will have expertise in distillation and quantization, as well as solid C++ and Python skills, with hands-on experience using TensorRT and ONNX Runtime.

Qualifications

  • Expertise in model distillation, pruning, and quantization techniques required.
  • Hands-on experience with TensorRT, ONNX Runtime, and edge deployment is essential.
  • Strong skills in both C++ and Python are necessary.

Responsibilities

  • Compress and optimize models for on-device inference.
  • Develop pipelines for model distillation and hardware-specific compilation.
  • Benchmark performance across various NPU/GPU architectures.

Skills

Model distillation
Pruning techniques
4-bit/8-bit quantization
C++
Python
TensorRT
ONNX Runtime

Job description

Hyphen Connect is seeking an AI Specialist Engineer in Boston, Massachusetts, to enhance performance for on-device inference in large language and vision models. You will be responsible for compressing and optimizing these models, developing pipelines specifically for model distillation, and benchmarking performance across various architectures. Ideal candidates will have expertise in distillation and quantization, as well as solid C++ and Python skills, with hands-on experience using TensorRT and ONNX Runtime.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI On-Device Engineer: Model Optimization & Edge Deployment
AI On-Device Engineer: Model Optimization & Edge Deployment

Hyphen Connect • Oregon (WI)

On-site
USD 100,000 - 130,000
Edge AI Engineer: On-Device Models, Distillation & Quantization
Edge AI Engineer: On-Device Models, Distillation & Quantization

Hyphen Connect • San Francisco (CA)

On-site
USD 120,000 - 160,000
Edge AI Engineer: On-Device Inference & Optimization
Edge AI Engineer: On-Device Inference & Optimization

Hyphen Connect • Seattle (WA)

On-site
USD 100,000 - 140,000
AI Specialist (AI Engineering)
AI Specialist (AI Engineering)

Hyphen Connect • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Specialist (AI Engineering)
AI Specialist (AI Engineering)

Hyphen Connect • Boston (MA)

On-site
USD 100,000 - 130,000
AI Specialist (AI Engineering)
AI Specialist (AI Engineering)

Hyphen Connect • Seattle (WA)

On-site
USD 100,000 - 140,000
AI Specialist (AI Engineering)
AI Specialist (AI Engineering)

Hyphen Connect • Oregon (WI)

On-site
USD 100,000 - 130,000
Remote AI Inference Engineer — Edge Model Deployment & Optimization
Remote AI Inference Engineer — Edge Model Deployment & Optimization

Quadric Inc. • Burlingame (CA)

On-site
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Life Insurance (Basic, Voluntary & AD&D)
+7
Senior Edge AI Engineer — On-Device ML & Model Compression
Senior Edge AI Engineer — On-Device ML & Model Compression

Axon Enterprise • Seattle (WA)

Hybrid
USD 177,000 - 284,000
Competitive salary and 401k with employer match
Discretionary paid time off
Paid parental leave
+4
Senior AI Inference Engineer — Edge ML, Porting Models
Senior AI Inference Engineer — Edge ML, Porting Models

quadric.io, Inc • Burlingame (CA)

On-site
USD 120,000 - 150,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Life Insurance (Basic, Voluntary & AD&D)
+7