Senior Kubernetes Developer - OPS00016

Dev.Pro

Bogotá

A distancia

COP 310.318.076 - 465.477.114

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Work on cutting-edge GPU infrastructure
Collaborate with a top-tier international team
Continuous learning and conference participation

Descripción de la vacante

An international tech company is seeking a skilled Kubernetes Developer to join their fully remote team. You will design and manage Kubernetes platforms for GPU-intensive workloads, collaborate across time zones, and enhance infrastructure performance. The ideal candidate has over 3 years of experience with Kubernetes, HPC schedulers, and strong programming skills in Go and Python. Opportunities for growth through continuous learning and conference participation are provided.

Formación

  • 3+ years of hands-on Kubernetes experience in production.
  • Experience with HPC schedulers such as Slurm, PBS, or LSF.
  • Strong background in GPU resource management and distributed systems.
  • Programming skills in Go and Python are essential.

Responsabilidades

  • Design and manage Kubernetes platforms for GPU-intensive AI/HPC workloads.
  • Design and build a Slurm-like orchestration layer on Kubernetes.
  • Develop CI/CD pipelines for GPU-intensive workloads.

Conocimientos

Hands-on Kubernetes experience
GPU resource management
Experience with HPC schedulers
Cloud/hybrid cloud architecture
Programming skills in Go
Programming skills in Python
Linux administration skills
Familiarity with PyTorch
Familiarity with TensorFlow
IaC tools (Terraform, Helm)

Descripción del empleo

Overview

Dev.Pro Bogota, D.C., Capital District, Colombia — We invite a skilled Kubernetes Developer to join our fully remote, international team. In this role, you'll build and optimize the Kubernetes orchestration platform and develop custom operators to run HPC/AI workloads efficiently on GPU clusters. You'll enhance infrastructure performance and reliability, create internal tools to improve the developer experience, and ensure multi-tenant HPC workloads remain secure and compliant.

What’s in it for you
  • Work on cutting-edge GPU infrastructure and next-gen HPC/AI workloads
  • Build a Slurm-on-Kubernetes product from scratch and shape its architecture
  • Collaborate with a top-tier international team and grow through continuous learning and conference participation
Key Responsibilities
  • Design, develop, and manage Kubernetes platforms for GPU-intensive AI/HPC workloads
  • Design and build a Slurm-like orchestration layer on Kubernetes for HPC/AI workloads
  • Develop custom operators and controllers for GPU job scheduling and execution
  • Integrate batch schedulers with Kubernetes to provide a hybrid HPC/Cloud product
  • Implement advanced GPU resource management and multi-tenant isolation policies
  • Build internal tools and a self-service platform to simplify AI/HPC job deployment and management
  • Monitor GPU clusters, troubleshoot production issues, and ensure high availability, fault tolerance, and disaster recovery
  • Develop CI/CD pipelines for GPU-intensive workloads
  • Ensure compliance with data sovereignty and international regulations
Qualifications
  • 3+ years of hands-on Kubernetes experience in production
  • Experience with HPC schedulers (Slurm, PBS, LSF, Volcano)
  • Strong background in GPU resource management and distributed systems
  • Experience with cloud/hybrid cloud architectures (AWS, GCP, Azure, on-prem GPU clusters)
  • Knowledge of Kubernetes operators, CRDs, scheduling, networking, and storage
  • Deep knowledge of HPC job scheduling and workload orchestration
  • Expertise in IaC (Terraform, Helm, or GitOps: ArgoCD/Flux) and monitoring & observability (Prometheus, Grafana, Jaeger, ELK)
  • Programming skills in Go, Python, Bash/Shell
  • Familiarity with PyTorch, TensorFlow, distributed training, and model serving
  • Skills in Linux administration, performance tuning, and advanced networking (RDMA, InfiniBand, TCP/IP, DNS, load balancing)
  • Experience in storage management and optimization for large datasets

Note: This role is fully remote and international, with a focus on collaboration across time zones.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Senior Kubernetes Engineer: GPU HPC Orchestration
Remote Senior Kubernetes Engineer: GPU HPC Orchestration

Dev.Pro • Bogotá

A distancia
COP 310.318.000 - 465.478.000
Platform Engineer (DevOps / Python)
Platform Engineer (DevOps / Python)

Nearshore Business Solutions • Bogotá

Presencial
COP 218.762.000 - 328.144.000
Senior DevOps/Platform Engineer - Remote - Colombia
Senior DevOps/Platform Engineer - Remote - Colombia

Kake • Colombia

A distancia
COP 281.171.000 - 437.377.000
Competitive USD pay
Fully Remote
Better Me Fund
+1
DevOps Engineer ID83107
DevOps Engineer ID83107

AgileEngine • Cartagena de Indias

A distancia
COP 60.000.000 - 120.000.000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
DevOps Engineer ID83107
DevOps Engineer ID83107

AgileEngine • Pereira

A distancia
COP 203.068.000 - 281.171.000
Growth opportunities
Competitive pay
Remote work
+3
DevOps Engineer ID83107
DevOps Engineer ID83107

AgileEngine • Centrosur

A distancia
COP 60.000.000 - 90.000.000
Growth without limits
Competitive compensation
Flexibility
+3
Remote Kubernetes Engineer | Scale Clusters & Cloud Apps
Remote Kubernetes Engineer | Scale Clusters & Cloud Apps

BairesDev • Colombia

Presencial
COP 381.667.000 - 572.501.000
Remote work
USD compensation
Hardware provided
+4
Senior AWS DevOps Engineer - Remote - Colombia
Senior AWS DevOps Engineer - Remote - Colombia

FullStack • Perímetro Urbano Medellín

Presencial
COP 185.253.000 - 259.356.000
Competitive Salary
Paid Time Off
100% remote work
+4
Senior AWS DevOps Engineer - Remote - Colombia
Senior AWS DevOps Engineer - Remote - Colombia

FullStack • Centro

Presencial
COP 181.290.000 - 290.066.000
Competitive Salary
Paid Time Off
100% remote work
+5
Senior AWS DevOps Engineer - Remote - Colombia
Senior AWS DevOps Engineer - Remote - Colombia

FullStack • Bogotá

Presencial
COP 222.304.000 - 333.457.000
Competitive Salary
Paid Time Off
100% remote work
+5