Principal ML Infra Engineer - GPU Inference & C++ Systems
Franklin Fitch
Dallas (TX)
On-site
USD 100,000 - 140,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A technology solutions provider in Dallas is seeking an AI Infrastructure Engineer with expertise in C++ and CUDA. The role involves designing and optimizing GPU-accelerated systems for deploying machine learning models in production. Ideal candidates will have a Master's or PhD, along with strong experience in GPU optimization and high-performance system design. Responsibilities include building inference pipelines, supporting model deployment, and working closely with ML researchers to ensure models are production-ready.
Qualifications
Strong C++ expertise with experience in production-grade systems.
Hands-on experience with CUDA programming and GPU optimization.
Solid understanding of GPU architectures and memory management.
Responsibilities
Design and maintain GPU-accelerated infrastructure for ML models.
Build and optimize inference pipelines for low-latency performance.
Support model conversion and deployment using inference runtimes.
Partner with ML researchers to transition models to production.
Skills
C++ expertise
CUDA programming
GPU optimization
Linux
Education
Masters or PhD
Tools
TensorRT
Job description
A technology solutions provider in Dallas is seeking an AI Infrastructure Engineer with expertise in C++ and CUDA. The role involves designing and optimizing GPU-accelerated systems for deploying machine learning models in production. Ideal candidates will have a Master's or PhD, along with strong experience in GPU optimization and high-performance system design. Responsibilities include building inference pipelines, supporting model deployment, and working closely with ML researchers to ensure models are production-ready.