AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc.

México

A distancia

MXN 900.000 - 1.300.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca la empresa.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Remote-first team
International collaboration
Cutting-edge AI research challenges

Descripción de la vacante

Lever, Inc. is seeking an AI Research Engineer (Kernel & Inference Optimization) based in Mexico.

You will work at the intersection of AI research, systems engineering, and high-performance model inference, focusing on latency, throughput, and memory efficiency across hardware environments. You will develop and optimize model-serving architectures, write custom GPU kernels (MSL), and collaborate with cross-functional teams to translate research into production edge applications.

Formación

  • Degree in Computer Science or related field; PhD in NLP/ML highly relevant.
  • Proven expertise in Metal Shading Language (MSL) and writing custom compute shaders.
  • Experience with low-level kernel optimization and inference optimization on mobile/edge devices.
  • Track record of measurable improvements in latency, throughput, and memory footprint.
  • Knowledge of modern model-serving architectures and optimization techniques for high-performance AI deployment.

Responsabilidades

  • Design and deploy advanced model-serving architectures optimized for high throughput and low latency.
  • Develop inference pipelines across diverse environments including edge devices.
  • Establish performance targets for latency, throughput, memory footprint, and reliability.
  • Build controlled inference benchmarks and track latency, throughput, and errors.
  • Create datasets and simulation scenarios to evaluate model performance.
  • Identify bottlenecks and apply system-level optimizations like batching and memory management.
  • Develop custom GPU kernels and compute shaders for mobile hardware (MSL).
  • Apply pruning, quantization, Flash Attention, KV caching, and speculative decoding.
  • Design distributed inference using tensor/pipeline/expert parallelism for large-scale GPUs.
  • Collaborate with cross-functional teams to integrate optimized inference frameworks.
  • Define evaluation methodologies and document experimental results.
  • Monitor production performance to identify improvement opportunities.

Conocimientos

MSL kernel programming
GPU kernel optimization
Inference optimization
Distributed inference
Benchmarking
English communication

Educación

PhD in NLP or Machine Learning
MS/BS in CS or related

Herramientas

MSL (Metal Shading Language)

Descripción del empleo

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in Mexico.

You will work at the intersection of AI research, systems engineering, and high-performance model inference.
Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments.
You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices.
The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels.
You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers.
Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements.
You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems.

Accountabilities
  • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
  • Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
  • Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
  • Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
  • Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
  • Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
  • Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
  • Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads.
  • Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications.
  • Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies.
  • Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability.
Requirements:
  • Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences.
  • Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch.
  • Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices.
  • Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications.
  • Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment.
  • Strong experience writing GPU kernels for mobile devices such as smartphones.
  • Practical experience developing and deploying end-to-end inference pipelines, from model optimization through production integration on constrained hardware.
  • Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges.
  • Experience designing robust evaluation and benchmarking frameworks for inference systems.
  • Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters.
  • Deep understanding of the mathematical foundations and architecture of diffusion models and Vision Transformers.
  • Familiarity with modern inference optimization techniques including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE.
  • Strong analytical and problem-solving abilities, with an ability to investigate complex system bottlenecks and turn research findings into practical engineering solutions.
  • Excellent English communication skills and the ability to collaborate effectively with distributed, cross-functional technical teams.
Benefits:
  • Opportunity to work on advanced AI systems spanning model serving, inference optimization, mobile computing, edge deployment, and large-scale distributed inference.
  • Remote-first working environment with an international team.
  • Exposure to cutting-edge AI research and practical systems engineering challenges.
  • Opportunity to contribute to performance-critical infrastructure where improvements can have a measurable impact on real-world AI applications.
  • Collaborative environment combining research-driven experimentation with hands-on engineering.
  • Opportunity to work with advanced model architectures including diffusion models, Vision Transformers, and multimodal systems.
  • Access to challenging technical problems involving GPU kernels, inference engines, memory optimization, and distributed computing.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote AI Research Engineer: Kernel & Inference Optimization
Remote AI Research Engineer: Kernel & Inference Optimization

Lever, Inc. • México

A distancia
MXN 900.000 - 1.500.000
Remote-first team
International collaboration
Cutting-edge AI research challenges
AI Engineer (Senior) ID34931
AI Engineer (Senior) ID34931

AgileEngine • Santiago de Querétaro

Presencial
MXN 1.117.734 - 1.862.891
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top-tier clients
+1
AI Architect ID34949– $2,500 Sign-On Bonus
AI Architect ID34949– $2,500 Sign-On Bonus

AgileEngine • Rosarito

Presencial
MXN 1.489.203 - 2.233.804
Professional growth opportunities
Competitive compensation
Flextime
+2
AI DI Engineering Manager - Mexico City
AI DI Engineering Manager - Mexico City

Aily Labs GmbH • Ciudad de México

Presencial
MXN 1.806.358 - 2.348.265
Competitive salary and equity package
Flexible, remote-first culture
Work in top AI environments
Lead AI Engineer
Lead AI Engineer

Blend • Región Centro

Híbrido
MXN 900.000 - 1.500.000
AWS Certifications
Udemy Business
English lessons
AI Software Engineer (Full Stack) ID51390
AI Software Engineer (Full Stack) ID51390

AgileEngine • Región Centro

Presencial
MXN 1.373.155 - 2.059.732
Professional growth assistance
Competitive USD-based compensation
Flexible schedule
+1
Generative AI Engineer
Generative AI Engineer

Altimetrik Mexico • México

Presencial
MXN 558.000 - 781.200
Seguro de salud integral
Bonos de desempeño
Días de vacaciones
+3
Desarrollador Back End (Infraestructura de IA)
Desarrollador Back End (Infraestructura de IA)

Alignerr Corp. • Monterrey

A distancia
MXN 344.000 - 689.000
AI Software Engineer (Full Stack) ID51390
AI Software Engineer (Full Stack) ID51390

AgileEngine • Ciudad de México

Presencial
MXN 1.373.155 - 1.888.088
Mentorship and personalized growth roadmaps
USD-based competitive compensation
Flexible schedule with remote options
+1
AI Software Engineer (Full Stack) ID51390
AI Software Engineer (Full Stack) ID51390

AgileEngine • Santiago de Querétaro

Presencial
MXN 1.544.799 - 2.059.732
Mentorship and personalized growth roadmaps
USD-based compensation with budgets for education and fitness
Flexible schedule with remote work options