Staff ML Inference Engineer — Model Efficiency (Remote)
Jaide Health
San Francisco (CA)
Presencial
USD 120.000 - 160.000
Jornada completa
14 días+
Generador de candidaturas
Convierte este puesto en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.
Supera los filtros ATS
Ventajas ofrecidas por este puesto de trabajo
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
100% Parental Leave top-up for up to 6 months
6 weeks of vacation (30 working days!)
Descripción de la vacante
Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations and collaborate closely with modeling and systems teams. Ideal candidates will have over 5 years of experience in high-performance coding, plus strong skills in C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and inclusive work culture is celebrated.
Formación
5+ years of experience writing high-performance, production-quality code.
Strong programming skills in C++ or Python (Rust/Go also welcome).
Experience working with large language models and familiarity with the LLM inference ecosystem.
Responsabilidades
Work across the inference stack to improve core performance metrics.
Identify bottlenecks and develop optimizations for model execution.
Collaborate closely with modeling and systems teams to measure and ship improvements.
Conocimientos
High-performance, production-quality code
Programming in C++ or Python
Diagnosing and resolving performance bottlenecks
GPU programming
Experience with large language models
Descripción del empleo
Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations and collaborate closely with modeling and systems teams. Ideal candidates will have over 5 years of experience in high-performance coding, plus strong skills in C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and inclusive work culture is celebrated.
Consigue la evaluación confidencial y gratuita de tu currículum.
Me encontraba atascado, enviaba solicitudes sin respuesta hasta que empecé a usar JobLeads. Hicieron que mi currículum fuera imposible de ignorar por las empresas.
Sophie Reynolds
La evaluación de currículums de JobLeads me ayudó a solucionar algunos errores fundamentales. ¡Empecé a recibir invitaciones a entrevistas casi inmediatamente!
Daniel Fischer
Con la revisión de currículums de JobLeads, ¡mi currículum pasó rápidamente de ignorado a listo para entrevistas!
Puestos de trabajo similares que vale la pena comparar