Inference Performance Engineer | Flexible, High-Impact ML

adaption

Argentina

Híbrido

ARS 2.000.000 - 4.200.000

Jornada completa

Hace 4 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Flexible work
Travel stipend
Lunch stipend
Well-being

Descripción de la vacante

Adaption is seeking a seasoned ML systems engineer to own the cost and performance of its inference stack. You will shape throughput and tail latency by tuning caching, batching, quantization, and decoding while preserving model quality.

You’ll collaborate with the serving fleet team, optimize routing to providers, and build profiling tools to reveal time, memory, and compute usage. Ideal candidates have 5+ years in ML systems and hands-on experience with vLLM, SGLang, or TensorRT-LLM, plus

Formación

  • 5+ years of ML systems or inference infrastructure experience with measurable improvements.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.

Responsabilidades

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Conocimientos

5+ years in ML systems
Model serving/inference
Python
C++
Rust
CUDA/NCCL/mixed precision
GPU performance

Herramientas

vLLM
SGLang
TensorRT-LLM

Descripción del empleo

Adaption is seeking a seasoned ML systems engineer to own the cost and performance of its inference stack. You will shape throughput and tail latency by tuning caching, batching, quantization, and decoding while preserving model quality.

You’ll collaborate with the serving fleet team, optimize routing to providers, and build profiling tools to reveal time, memory, and compute usage. Ideal candidates have 5+ years in ML systems and hands-on experience with vLLM, SGLang, or TensorRT-LLM, plus

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Inference Performance Engineer
Inference Performance Engineer

adaption • Argentina

Híbrido
ARS 2.000.000 - 4.200.000
Flexible work
Travel stipend
Lunch stipend
+1
ML Inference Platform Engineer — Low-Latency GPU Systems
ML Inference Platform Engineer — Low-Latency GPU Systems

Dialpad • Buenos Aires

Híbrido
ARS 133.982.000 - 193.530.000
Competitive salary
Comprehensive benefits
Opportunities for growth
+1
Software Engineer, ML Inference Platform
Software Engineer, ML Inference Platform

United States Digital Space LLC • Buenos Aires

Presencial
ARS 1.500.000 - 2.100.000
Senior ML Engineer — Flexible Hours, LLM Platform
Senior ML Engineer — Flexible Hours, LLM Platform

Intellectsoft • Argentina

Presencial
ARS 132.725.000 - 221.210.000
Udemy courses
Flexible hours
Career path
+2
Data & Machine Learning Engineer
Data & Machine Learning Engineer

IDT • Buenos Aires

Presencial
ARS 104.200.779 - 133.972.431
Senior Python Backend Engineer - AI/ML & LLMs, Remote
Senior Python Backend Engineer - AI/ML & LLMs, Remote

Intellectsoft • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Awesome projects with an impact
Udemy courses of your choice
Team-building events
+2
Senior AI Engineer - Agentic LLMs & RAG in Production
Senior AI Engineer - Agentic LLMs & RAG in Production

intive • Argentina

Presencial
ARS 1.200.000 - 2.400.000
Remote-friendly work environment
Udemy access & training resources
Mentorship program
+1
Agent Systems Engineer - Build Adaptive, Reliable AI Agents
Agent Systems Engineer - Build Adaptive, Reliable AI Agents

adaption • Argentina

Híbrido
ARS 178.643.000 - 282.852.000
Bay Area collaboration
Adaption Passport travel stipend
Lunch stipend
+2
Production ML/AI Engineer - Scale AI Features & Systems
Production ML/AI Engineer - Scale AI Features & Systems

Silver.dev • Argentina

Híbrido
ARS 800.000 - 1.200.000
AI Engineering Leader: Production ML & RAG Systems
AI Engineering Leader: Production ML & RAG Systems

Blend • Municipio de Esquel

Presencial
ARS 3.000.000 - 6.000.000
AWS certifications
Databricks learning paths
Study plans and courses
+2