ML Inference Platform Engineer — Low-Latency GPU Systems

Dialpad

Buenos Aires

Híbrido

ARS 133.982.000 - 193.530.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Competitive salary
Comprehensive benefits
Opportunities for growth
AI tools training

Descripción de la vacante

Dialpad is hiring a Software Engineer for the ML Inference Platform to build production systems that serve in-house AI models at scale. You will work at the intersection of model development, runtime systems, and cloud infrastructure on NVIDIA GPUs in GCP.

This is an implementation-heavy role focused on model serving, deployment safety, telemetry, and cost-aware optimization. You’ll help define packaging, deployment, benchmarking, and operations across environments.

Formación

  • 6+ years of professional software engineering experience shipping backend services or production platforms.

Responsabilidades

  • Design, build, and improve systems connecting AI capability development to production inference.
  • Inference Serving & Runtime Systems for low-latency, high-throughput workloads.
  • GPU Infrastructure & Utilization on Kubernetes/GCP with NVIDIA GPUs.
  • Model Server Integration with frameworks like vLLM, Triton, or TGI.
  • Traffic, release safety, and canary rollout mechanisms for model-backed services.
  • Benchmarking & evaluation tooling for latency, throughput, cost, and reliability.
  • Artifact lifecycle management across environments and observability capabilities.

Conocimientos

Production engineering
Python/Go coding
Distributed systems
Kubernetes & Linux
Observability & debugging
Collaborative communication

Herramientas

Kubernetes
Linux
CI/CD
Triton
vLLM
GCP

Descripción del empleo

Dialpad is hiring a Software Engineer for the ML Inference Platform to build production systems that serve in-house AI models at scale. You will work at the intersection of model development, runtime systems, and cloud infrastructure on NVIDIA GPUs in GCP.

This is an implementation-heavy role focused on model serving, deployment safety, telemetry, and cost-aware optimization. You’ll help define packaging, deployment, benchmarking, and operations across environments.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Software Engineer, ML Inference Platform
Software Engineer, ML Inference Platform

United States Digital Space LLC • Buenos Aires

Presencial
ARS 1.500.000 - 2.100.000
Software Engineer, ML Inference Platform
Software Engineer, ML Inference Platform

Dialpad • Buenos Aires

Híbrido
ARS 133.982.000 - 193.530.000
Competitive salary
Comprehensive benefits
Opportunities for growth
+1
Senior ML Engineer — Flexible Hours, LLM Platform
Senior ML Engineer — Flexible Hours, LLM Platform

Intellectsoft • Argentina

Presencial
ARS 132.725.000 - 221.210.000
Udemy courses
Flexible hours
Career path
+2
AI Platform Engineer: Scalable Data Infra for AI Mining
AI Platform Engineer: Scalable Data Infra for AI Mining

Dialpad • Buenos Aires

Presencial
ARS 2.000.000 - 3.000.000
Competitive salary
Comprehensive benefits
Growth opportunities
Senior Python Backend Engineer - AI/ML & LLMs, Remote
Senior Python Backend Engineer - AI/ML & LLMs, Remote

Intellectsoft • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Awesome projects with an impact
Udemy courses of your choice
Team-building events
+2
Data & Machine Learning Engineer
Data & Machine Learning Engineer

IDT • Buenos Aires

Presencial
ARS 104.200.779 - 133.972.431
End-to-End Data & ML Engineer (LATAM) — AI-Driven Pipelines
End-to-End Data & ML Engineer (LATAM) — AI-Driven Pipelines

Medium • Municipio de Colalao del Valle

Presencial
ARS 104.200.000 - 148.859.000
Production ML Engineer — MLOps
Production ML Engineer — MLOps

Proofpoint • Córdoba

Presencial
ARS 104.359.000 - 163.994.000
Competitive compensation
Comprehensive benefits
Flexible work environment
+1
AI Transformation Engineer
AI Transformation Engineer

Dialpad • Buenos Aires

Presencial
ARS 1.400.000 - 1.900.000
GenAI & ML Engineer (AWS Cloud) — Remote
GenAI & ML Engineer (AWS Cloud) — Remote

Ideamine Technologies (Acquired by Netrix Global) • Buenos Aires

Presencial
ARS 900.000 - 1.300.000
Remote work allowed
AWS certifications
Microsoft certifications
+3