Get more replies from employers
Send a job-specific resume in minutes.
EngRadar seeks a senior ML infrastructure engineer to develop and optimize the inference stack for both on-device and cloud AI, integrating new engines and multimodal models. You will work on latency and throughput improvements across CPU/GPU runtimes and contribute to open-source projects.
Required is extensive experience building production ML systems with strong Python and C++ skills, plus deep transformer knowledge and CPU/GPU profiling expertise.
The role involves developing and optimizing the inference stack for on-device and cloud AI, integrating new engines and multimodal models. Responsibilities include improving latency and throughput across various hardware runtimes and contributing to open-source projects.
Requirements: Candidates need significant experience building production ML systems or performance-sensitive infrastructure with strong proficiency in Python and C++. Deep knowledge of transformer architectures and experience profiling CPU/GPU workloads are essential.
Key Skills: Python, C++, Transformer Architectures, Model Inference, CPU Profiling, GPU Profiling, PyTorch, Llama.cpp, MLX, ExecuTorch, vLLM, SGLang, TensorRT-LLM, CUDA, Metal, Vulkan
Benefits: Competitive salary, Equity grants, Medical healthcare plan, Vision healthcare plan, Dental healthcare plan, Catered team lunch, Expensed dinners in the office, Flexible PTO, Flexible WFH