Senior DevOps Engineer, AI Platform

Network Solutions

Argentina

Presencial

ARS 135.892.000 - 226.487.000

Jornada completa

hace 44 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.

Supera los filtros ATS

Descripción de la vacante

Network Solutions is seeking a hands-on Senior DevOps Engineer to build and operate the infrastructure powering our AI platforms, agent runtimes, web applications, and APIs across Azure and OCI. You will collaborate with AI and application engineers to translate designs into reliable, scalable, observable production systems.

Key responsibilities include designing Kubernetes environments, managing ingress and egress networking, CI/CD pipelines, and end-to-end observability.

Formación

  • 7+ years in DevOps/SRE/Platform Engineering or related role.
  • Hands-on production Kubernetes experience with networking, storage, security, and troubleshooting.
  • Strong Azure experience (AKS, networking, identity, storage, monitoring); OCI or cross-cloud familiarity is a plus.
  • Extensive experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and IaC.
  • Proven track record supporting production web apps and backend services (REST APIs, microservices, async architectures).
  • Experience with databases, caching, and messaging systems (PostgreSQL, Redis, RabbitMQ).
  • Observability experience with OpenTelemetry, Grafana, Prometheus, Sentry or similar tools.
  • Strong Linux, systems and production troubleshooting skills.

Responsabilidades

  • Translate designs into production-ready cloud infrastructure with minimal supervision.
  • Design, provision, operate, and troubleshoot Kubernetes environments (AKS and OCI-based where applicable).
  • Support AI workloads and associated runtimes, RAG pipelines, and background processing.
  • Design and manage networking, load balancing, DNS, TLS, and private connectivity.
  • Build and operate infrastructure for web apps and backend services (APIs, databases, caches, queues).
  • Develop CI/CD pipelines with Jenkins/Bitbucket, Docker, Helm, Kubernetes, ArgoCD, and registries.
  • Automate provisioning with Terraform, Helm, Kubernetes manifests, Python, Bash.
  • Implement end-to-end observability with metrics, logs, tracing, dashboards, alerts, and SLOs.
  • Own production readiness, incident response, RCA, scalability, and cost optimization.

Conocimientos

Kubernetes
Azure AKS
Terraform
CI/CD
Docker
Jenkins
Bitbucket
ArgoCD
OpenTelemetry
Grafana
Prometheus
PostgreSQL
Redis
RabbitMQ
Linux
Python
FastAPI
Networking
Security
Observability

Herramientas

Kubernetes
Azure
OCI
Jenkins
Bitbucket
Docker
Terraform
Helm
ArgoCD
OpenTelemetry
Grafana
Prometheus
PostgreSQL
Redis
RabbitMQ

Descripción del empleo

At Network Solutions, we’ve been trusted for decades to help people get online and stay ahead. We’ve been here since the beginning of the internet, and we’re still building for what comes next.

As the original digital identity authority, we help secure domain names, protect brands, and safeguard the infrastructure businesses rely on. We empower our customers to own and manage the assets that define them online, while delivering enterprise-grade security to protect against virtual threats. Our team leverages modern, AI-accelerated tools to streamline how businesses manage their digital presence, making the most of our decades of experience.

The Network Solutions team is here to help online businesses protect what’s theirs and build for tomorrow. That’s why millions trust us to protect their domains, brands, and websites every day.

The impact you’ll make

We are looking for a hands‑on Senior DevOps Engineer to build and operate the infrastructure powering our AI platforms, agent runtimes, web applications, backend services, APIs, and shared platform capabilities. Our environment includes a centralized LLM gateway, Python-based agent runtimes, RAG workers, MCP services, asynchronous processing, databases, caches, queues, and observability services.

You will work with AI engineers, application engineers, and architects who define technical designs, then independently translate those designs into reliable, scalable, secure, and observable production infrastructure across Microsoft Azure and Oracle Cloud Infrastructure.

What you’ll do
  • Translate application and platform technical designs into production ready cloud infrastructure with minimal supervision.
  • Design, provision, operate, and troubleshoot Kubernetes environments, primarily Azure Kubernetes Service and Oracle Kubernetes Engine.
  • Support AI workloads including LiteLLM based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
  • Design and manage ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and service to service communication.
  • Build and operate infrastructure for web applications and backend services, including APIs, databases, caches, queues, scheduled jobs, and event driven workloads.
  • Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
  • Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tooling.
  • Implement end to end observability using metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.
  • Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
  • Create reusable infrastructure patterns that allow engineering teams to launch new services quickly and consistently.
What we’re looking for
  • 7 or more years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or a related role.
  • Strong hands on experience operating production Kubernetes environments and deep knowledge of networking, scheduling, storage, autoscaling, security, and troubleshooting.
  • Strong Microsoft Azure experience, including AKS, networking, identity, storage, and monitoring. OCI experience is preferred, or demonstrated ability to work across cloud providers.
  • Strong cloud networking knowledge across virtual networks, subnets, routing, NAT, load balancers, private networking, DNS, TLS, firewalls, ingress, and egress.
  • Strong experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code.
  • Proven experience supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.
  • Hands on experience with databases, caching, and messaging systems such as PostgreSQL, Redis, RabbitMQ, or equivalent technologies.
  • Experience implementing production observability using OpenTelemetry, Grafana, Prometheus, Sentry, cloud monitoring, or similar tools.
  • Strong Linux, systems, and production troubleshooting skills.

You do not need to be a full time application developer, but you should understand how modern backend systems work and be able to troubleshoot across application and infrastructure boundaries.

  • Working knowledge of Python, especially backend services built with frameworks such as FastAPI.
  • Understanding of HTTP, HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead letter queues, and asynchronous processing.
  • Ability to read application logs and stack traces and diagnose latency, memory, CPU, connection, and dependency issues.
How you’ll work

You will frequently receive a technical design for a new AI workload, application, backend service, or platform capability. From that design, you should be able to independently determine and implement the infrastructure needed to run it in production.

  • Determine the required cloud resources, Kubernetes configuration, namespaces, scaling model, and supporting services.
  • Configure ingress and egress, private connectivity, DNS, TLS, service communication, identities, and secrets.
  • Provision and operate dependencies such as PostgreSQL, Redis, RabbitMQ, storage, and other shared services.
  • Build the Jenkins and Bitbucket CI/CD flow for build, test, container publishing, deployment, validation, and rollback.
  • Define observability, health checks, dashboards, alerts, capacity monitoring, and operational runbooks before production launch.
  • Own infrastructure delivery through UAT and production, partnering with architects and engineers when design tradeoffs require discussion.
What will make you stand out
  • Experience supporting AI or machine learning platforms, LLM gateways, agent runtimes, RAG pipelines, or MCP services.
  • Experience with Cloudflare, Envoy, ArgoCD, GitOps, and OpenTelemetry.
  • Experience building reusable infrastructure platforms for high scale SaaS or customer facing applications.
  • Experience operating distributed backend systems using RabbitMQ, Redis, PostgreSQL, and similar technologies.

AI or machine learning infrastructure experience is helpful, but not required. Strong experience with Kubernetes, web and backend infrastructure, networking, queues, databases, CI/CD, observability, and production cloud operations is the foundation for this role.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior DevOps Engineer for AI Platform & Cloud Infra
Senior DevOps Engineer for AI Platform & Cloud Infra

Network Solutions • Argentina

Presencial
ARS 135.892.000 - 226.487.000
Senior Applied AI Engineer
Senior Applied AI Engineer

Network Solutions • Argentina

Presencial
ARS 1.800.000 - 3.000.000
[8AI] DevOps Engineer
[8AI] DevOps Engineer

Visa Hunt • Argentina

Presencial
ARS 1.400.000 - 2.800.000
Senior DevOps Engineer Azure
Senior DevOps Engineer Azure

Visa Hunt • Argentina

Presencial
ARS 181.190.000 - 271.784.000
Principal DevOps AI IRC303057
Principal DevOps AI IRC303057

GlobalLogic • Argentina

Presencial
ARS 133.982.000 - 193.530.000
Principal DevOps AI IRC303057
Principal DevOps AI IRC303057

GlobalLogic • Municipio de Rincón de los Sauces

Presencial
ARS 3.000.000 - 6.000.000
Competitive salary
Global exposure and travel
Professional development
+1
Senior AI Engineer
Senior AI Engineer

MAS Global Consulting • Argentina

Presencial
ARS 83.351.000 - 111.136.000
Principal DevOps AI IRC303057
Principal DevOps AI IRC303057

GlobalLogic • Buenos Aires

Híbrido
ARS 133.982.000 - 193.530.000
Exciting Projects
Collaborative Environment
Work-Life Balance
+2
Sr DevOps Platform Engineer - Bilingual
Sr DevOps Platform Engineer - Bilingual

Concentrix • Buenos Aires

Presencial
ARS 3.600.000 - 4.800.000
Sr DevOps Engineer
Sr DevOps Engineer

Pyramid Consulting, Inc • Argentina

Presencial
ARS 1.500.000 - 2.100.000