Edge AI Engineer: On-Device Models, Distillation & Quantization
Hyphen Connect
San Francisco (CA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
Hyphen Connect seeks an AI Specialist Engineer in San Francisco to enhance AI performance on diverse hardware architectures. You will efficiently develop and deploy cutting-edge AI solutions. Responsibilities include compressing large models for on-device use, developing model distillation pipelines, and benchmarking across NPU/GPU architectures. Ideal candidates should have expertise in AI model optimization, strong C++ and Python skills, and hands-on experience with TensorRT and ONNX Runtime.
Qualifications
Expertise in model distillation, pruning, and quantization techniques.
Hands-on experience with TensorRT, ONNX Runtime, and edge deployment.
Strong abilities in C++ and Python.
Responsibilities
Compress and optimize large language and vision models for on-device inference.
Develop pipelines for model distillation and hardware-specific compilation.
Benchmark performance across NPU/GPU architectures.
Skills
Model distillation
C++
Python
TensorRT
ONNX Runtime
Job description
Hyphen Connect seeks an AI Specialist Engineer in San Francisco to enhance AI performance on diverse hardware architectures. You will efficiently develop and deploy cutting-edge AI solutions. Responsibilities include compressing large models for on-device use, developing model distillation pipelines, and benchmarking across NPU/GPU architectures. Ideal candidates should have expertise in AI model optimization, strong C++ and Python skills, and hands-on experience with TensorRT and ONNX Runtime.