Software Engineer, ML Inference Platform

United States Digital Space LLC

Buenos Aires

Presencial

ARS 1.500.000 - 2.100.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

United States Digital Space LLC in Buenos Aires is looking for ML Inference Platform Engineers to build the production systems that serve our in-house AI models at scale.

You will operate at the intersection of model development, high-performance runtime systems, and cloud infrastructure, delivering low-latency, observable services on NVIDIA GPUs in GCP.

Formación

  • Production Engineer

Responsabilidades

  • Design, build and improve systems that connect AI capability development to production inference.
  • Inference Serving & Runtime Systems: Build and improve model-serving pathways for low-latency, high-throughput inference workloads.
  • GPU Infrastructure & Utilization: Operate and optimize containerized workloads on Kubernetes/GCP, with focus on NVIDIA GPUs, memory, storage, and networking.
  • Model Server Integration: Work with model-serving frameworks and runtimes such as vLLM, Triton, TGI.
  • Traffic & Release Safety: Enable shadow serving, canary rollouts, staged deployments, and fast rollback mechanisms.
  • Benchmarking & Evaluation Infrastructure: Build tooling to measure latency, throughput, cost, saturation behavior, and reliability.

Conocimientos

Production Engineer

Descripción del empleo

About the company

the company is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.

Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, the company was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved.

Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust the company. the company is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile.

Being a Dialer

At the company, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more.

We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves.

We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic.

Your role

We are hiring ML Inference Platform Engineers to build the production systems that serve our in-house AI models at scale.

This role sits at the intersection of model development, high-performance runtime systems, and cloud infrastructure. You will help turn trained models and emerging AI capabilities into reliable, observable, low-latency production services running on NVIDIA GPUs in GCP.

This is not a research role, and it is not a generic MLOps or support role. It is an implementation-heavy systems engineering role focused on the machinery of inference: model serving, runtime optimization, GPU utilization, deployment safety, traffic management, benchmarking, and production reliability.

Our mission is to shorten the path from promising model capability to dependable production impact. We build the shared infrastructure, standards, and release pathways that allow models to move from candidate artifacts into scalable, rollback-safe inference services with clear performance, reliability, and cost characteristics.

This is a new team, so the systems and interfaces are still being shaped. You will help define how models are packaged, deployed, benchmarked, monitored, compared, and operated across environments. The work is practical, deeply technical, and closely tied to the company’s broader AI strategy. We are not building one-off demos; we are building the inference platform by which a growing AI organization can repeatedly and safely ship real model-backed products.

What you’ll do
  • You will design, build, and improve the systems that connect AI capability development to production inference.
  • Depending on your strengths, your work may include:
  • Inference Serving & Runtime Systems: Build and improve model-serving pathways for low-latency, high-throughput, high-availability inference workloads.
  • GPU Infrastructure & Utilization: Operate and optimize containerized workloads on Kubernetes/GCP, with a focus on efficient use of NVIDIA GPUs, memory, storage, and networking.
  • Model Server Integration: Work with model-serving frameworks and runtimes such as vLLM, Triton, TGI, or similar systems, adapting them to internal deployment, observability, and release requirements.
  • Traffic & Release Safety: Enable shadow serving, canary rollouts, staged deployments, candidate-versus-incumbent comparisons, and fast rollback mechanisms for model-backed services.
  • Benchmarking & Evaluation Infrastructure: Build tooling to measure latency, throughput, cost, saturation behavior, and reliability under realistic production traffic.
  • Artifact Lifecycle: Improve how model and capability artifacts are packaged, versioned, promoted, deployed, and rolled back across environments.
  • Observability & Debuggability: Strengthen runtime telemetry, structured logging, tracing, dashboards, and alerting so engineers can understand model-serving behavior in production.
  • Efficiency & Scale: Contribute to strategies that improve compute efficiency, GPU utilization, autoscaling behavior, and cost-performance tradeoffs across the inference platform.
Skills you’ll bring
  • Production Engineer
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Software Engineer, ML Inference Platform
Software Engineer, ML Inference Platform

Dialpad • Buenos Aires

Híbrido
ARS 133.982.000 - 193.530.000
Competitive salary
Comprehensive benefits
Opportunities for growth
+1
Engineering Manager, AI Transformation
Engineering Manager, AI Transformation

United States Digital Space LLC • Buenos Aires

Presencial
ARS 178.643.000 - 267.965.000
AI Platform Engineer
AI Platform Engineer

United States Digital Space LLC • Buenos Aires

Presencial
ARS 89.322.000 - 148.869.000
ML Inference Platform Engineer — Low-Latency GPU Systems
ML Inference Platform Engineer — Low-Latency GPU Systems

Dialpad • Buenos Aires

Híbrido
ARS 133.982.000 - 193.530.000
Competitive salary
Comprehensive benefits
Opportunities for growth
+1
Engineering Manager, AI Transformation
Engineering Manager, AI Transformation

Dialpad • Buenos Aires

Híbrido
ARS 1.500.000 - 4.000.000
Competitive salary
Comprehensive benefits
Training program
+2
AI Platform Engineer
AI Platform Engineer

Dialpad • Buenos Aires

Presencial
ARS 2.000.000 - 3.000.000
Competitive salary
Comprehensive benefits
Growth opportunities
Engineering Manager, AI Transformation
Engineering Manager, AI Transformation

Dialpad Japan • Buenos Aires

Presencial
ARS 103.231.000 - 162.221.000
Agent Systems Engineer
Agent Systems Engineer

adaption • Argentina

Híbrido
ARS 178.643.000 - 282.852.000
Bay Area collaboration
Adaption Passport travel stipend
Lunch stipend
+2
Software Engineer - AI
Software Engineer - AI

Qodea • Buenos Aires

Presencial
ARS 97.269.000 - 142.162.000
OSDE 210 for family group
Work from Home Allowance
Birthday leave
+2
Software Engineer - AI
Software Engineer - AI

Bynd • Buenos Aires

Híbrido
ARS 1.000.000 - 2.000.000
OSDE 210 for family group
Work from Home Allowance
Birthday leave
+2