On-Device AI Engineer: Model Optimization & Edge Deployment
Hyphen Connect
Boston (MA)
On-site
USD 100,000 - 130,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
Hyphen Connect is seeking an AI Specialist Engineer in Boston, Massachusetts, to enhance performance for on-device inference in large language and vision models. You will be responsible for compressing and optimizing these models, developing pipelines specifically for model distillation, and benchmarking performance across various architectures. Ideal candidates will have expertise in distillation and quantization, as well as solid C++ and Python skills, with hands-on experience using TensorRT and ONNX Runtime.
Qualifications
Expertise in model distillation, pruning, and quantization techniques required.
Hands-on experience with TensorRT, ONNX Runtime, and edge deployment is essential.
Strong skills in both C++ and Python are necessary.
Responsibilities
Compress and optimize models for on-device inference.
Develop pipelines for model distillation and hardware-specific compilation.
Benchmark performance across various NPU/GPU architectures.
Skills
Model distillation
Pruning techniques
4-bit/8-bit quantization
C++
Python
TensorRT
ONNX Runtime
Job description
Hyphen Connect is seeking an AI Specialist Engineer in Boston, Massachusetts, to enhance performance for on-device inference in large language and vision models. You will be responsible for compressing and optimizing these models, developing pipelines specifically for model distillation, and benchmarking performance across various architectures. Ideal candidates will have expertise in distillation and quantization, as well as solid C++ and Python skills, with hands-on experience using TensorRT and ONNX Runtime.