ML Inference Platform Engineer: Low-Latency AI on GPUs

Dialpad

Buenos Aires

Presencial

ARS 181.094.000 - 271.641.000

Jornada completa

14 días+
Generador de candidaturas

Una candidatura completa en un minuto: currículum y carta de presentación adaptados, listos para enviar.

Supera los filtros ATS

Descripción de la vacante

Dialpad is seeking ML Inference Platform Engineers to build the production systems that serve our in-house AI models at scale. This role lies at the intersection of model development, high-performance runtime systems, and cloud infrastructure.

You will help turn trained models and AI capabilities into reliable, observability-rich, low-latency services running on NVIDIA GPUs in GCP. This is an implementation-heavy role focused on inference machinery, deployment safety, and scalable production

Formación

  • 6+ years of professional software engineering experience shipping backend services, infrastructure systems, or production platforms.
  • Proficiency in writing maintainable production code in Python, Go, or another backend-oriented language.
  • Experience building, operating, or optimizing high-throughput services, distributed systems, or ML infrastructure.
  • Hands-on experience with containers, Kubernetes, Linux environments, CI/CD, deployment automation, and production operations.
  • Strong instinct for reproducibility, observability, rollout safety, failure modes, and resilience.
  • Comfort reasoning about bottlenecks across compute, memory, network, storage, batching, concurrency, and SLOs.
  • Ability to work closely with model developers, product engineers, infrastructure teams, and leadership.

Responsabilidades

  • Design, build, and improve systems connecting AI capability development to production inference.
  • Develop inference serving and runtime systems for low-latency, high-throughput workloads.
  • Operate and optimize GPU-backed containerized workloads on Kubernetes/GCP.
  • Integrate model-serving frameworks and runtimes (e.g., vLLM, Triton, TGI) for internal deployment and observability.
  • Enable shadow serving, canary rollouts, and safe rollback mechanisms for model-backed services.
  • Build tooling to measure latency, throughput, cost, and reliability under production traffic.
  • Improve packaging, versioning, deployment, and rollback of model artifacts across environments.
  • Strengthen telemetry, logging, tracing, dashboards, and alerting for production insight.
  • Contribute to compute efficiency and cost-performance tradeoffs across the platform.

Conocimientos

Production Engineering Experience
Strong Software Fundamentals
Inference or Systems Orientation
Kubernetes & Linux Fluency
Operational Judgment
Performance Awareness
Collaboration

Herramientas

vLLM
Triton
TGI

Descripción del empleo

Dialpad is seeking ML Inference Platform Engineers to build the production systems that serve our in-house AI models at scale. This role lies at the intersection of model development, high-performance runtime systems, and cloud infrastructure.

You will help turn trained models and AI capabilities into reliable, observability-rich, low-latency services running on NVIDIA GPUs in GCP. This is an implementation-heavy role focused on inference machinery, deployment safety, and scalable production

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior AI Platform Engineer: Inference & Production
Senior AI Platform Engineer: Inference & Production

Linuxconfig • Buenos Aires

Híbrido
ARS 136.003.000 - 211.560.000
Senior AI/ML Inference Platform Engineer
Senior AI/ML Inference Platform Engineer

Dialpad • Buenos Aires

Híbrido
ARS 241.615.000 - 392.625.000
Competitive salary
Comprehensive benefits
Growth opportunities
Software Engineer, AI / ML Inference Platform
Software Engineer, AI / ML Inference Platform

Dialpad • Buenos Aires

Presencial
ARS 181.094.000 - 271.641.000
Sr. Software Engineer, AI / ML Inference Platform
Sr. Software Engineer, AI / ML Inference Platform

Dialpad • Buenos Aires

Presencial
ARS 241.615.000 - 392.625.000
Competitive salary
Comprehensive benefits
Growth opportunities
Sr. Software Engineer, AI / ML Inference Platform
Sr. Software Engineer, AI / ML Inference Platform

Linuxconfig • Buenos Aires

Híbrido
ARS 136.003.000 - 211.560.000
Senior Python Backend Engineer - AI/ML & LLMs, Remote
Senior Python Backend Engineer - AI/ML & LLMs, Remote

Intellectsoft • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Awesome projects with an impact
Udemy courses of your choice
Team-building events
+2
Applied AI Engineer: Production-Grade LLM Systems
Applied AI Engineer: Production-Grade LLM Systems

Newfold Digital • Argentina

Presencial
ARS 181.094.000 - 271.641.000
Senior Applied AI Engineer — Production ML & LLMs (Remote)
Senior Applied AI Engineer — Production ML & LLMs (Remote)

phData • Jesús María

Híbrido
ARS 180.802.000 - 271.203.000
Remote-first
Competitive pay
PTO & holidays
+1
AI Engineer — Vertex AI, LLMs & Python
AI Engineer — Vertex AI, LLMs & Python

Qodea • Buenos Aires

Presencial
ARS 97.269.000 - 142.162.000
OSDE 210 for family group
Work from Home Allowance
Birthday leave
+2
Data & Machine Learning Engineer
Data & Machine Learning Engineer

IDT • Buenos Aires

Presencial
ARS 104.200.779 - 133.972.431