AI On-Device Engineer: Model Optimization & Edge Deployment
Hyphen Connect
Oregon (WI)
On-site
USD 100,000 - 130,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
Hyphen Connect is seeking an AI Specialist Engineer to enhance the performance of large language and vision models for on-device inference. You will be crucial in developing AI solutions, optimizing model performance across diverse hardware architectures. Key qualifications include expertise in model distillation and quantization techniques, hands-on experience with TensorRT and ONNX Runtime, along with strong C++ and Python skills. This role offers exciting opportunities to work on cutting-edge technology in the United States.
Qualifications
Expertise in model distillation, pruning, and quantization techniques.
Hands-on experience with TensorRT, ONNX Runtime, and edge deployment.
Strong C++ and Python skills.
Responsibilities
Compress and optimize large language and vision models for on-device inference.
Develop pipelines for model distillation and hardware-specific compilation.
Benchmark performance across various NPU/GPU architectures.
Skills
Model distillation
Pruning techniques
Quantization techniques
TensorRT
ONNX Runtime
C++
Python
Job description
Hyphen Connect is seeking an AI Specialist Engineer to enhance the performance of large language and vision models for on-device inference. You will be crucial in developing AI solutions, optimizing model performance across diverse hardware architectures. Key qualifications include expertise in model distillation and quantization techniques, hands-on experience with TensorRT and ONNX Runtime, along with strong C++ and Python skills. This role offers exciting opportunities to work on cutting-edge technology in the United States.