Technical Lead - GPU Infrastructure

Jobgether

Turkey

Remote

TRY 500,000 - 900,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Remote work
Leadership role
AI research support
Travel opportunities

Job summary

Jobgether is seeking a Technical Lead for GPU Infrastructure in a fully remote setup. The role focuses on architecting and delivering a large-scale GPU platform, evolving from managed Kubernetes toward bare-metal GPU infra, including Slurm-based research computing and Kubernetes-powered inference.

You will lead a distributed team spanning backend, frontend, DevOps, QA, and documentation while maintaining high technical standards. This is a hands-on leadership role with global collaboration.

Qualifications

  • Extensive hands-on GPU and infra expertise with leadership experience.
  • Experience operating Slurm in production and HPC/GPU clusters.
  • Proven Kubernetes control plane and multi-tenant environments.

Responsibilities

  • Own platform architecture and delivery across GPU infra stack.
  • Lead distributed engineering teams across backend, frontend, DevOps, and QA.
  • Define engineering standards, reviews, and release gates.
  • Design, operate, and scale Slurm service for research workloads.
  • Oversee bare-metal GPU infrastructure and CUDA lifecycle.
  • Manage Kubernetes bootstrap on partner-provided hardware with operators.
  • Drive observability, SLOs, and incident response for the platform.
  • Collaborate with partners, vendors, and research teams to translate requirements.

Skills

GPU infrastructure
Kubernetes
Slurm
NVIDIA GPUs
Prometheus/Grafana
Leadership
English proficiency
Linux systems
KubeVirt/VFIO
Architecture reviews

Education

Bachelor's or higher in CS/Engineering

Tools

Kubernetes
Slurm
CUDA drivers
Fabric Manager
NVIDIA NVSwitch

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Technical Lead - GPU Infrastructure based in Turkey.

This is a hands-on technical leadership role responsible for architecting and delivering a large-scale GPU infrastructure platform.

You will lead the evolution from managed Kubernetes workloads toward bare-metal GPU infrastructure, including Slurm-based research computing and Kubernetes-powered inference.

The role combines deep systems expertise with engineering leadership, team management, and direct ownership of architecture and delivery.

You will oversee a distributed team spanning backend, frontend, DevOps, QA, and documentation while maintaining high technical standards.

Your work will support research, model training, and managed inference workloads requiring reliable, scalable, and observable GPU compute.

You will also serve as the primary technical interface with infrastructure partners, hardware providers, and internal platform consumers.

This is an opportunity to shape the architecture and operational foundations of a sophisticated GPU platform in a fully remote environment.

Accountabilities:

The Technical Lead will own the platform architecture and engineering delivery while remaining deeply involved in technical decisions, infrastructure operations, team leadership, and partner relationships.

  • Own the end-to-end platform architecture, including architecture proposals, high-level and low-level designs, technical reviews, and ongoing architecture documentation.
  • Lead and line-manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation.
  • Establish engineering standards, oversee code and design reviews, manage release gates, conduct one-to-ones, and provide growth and performance feedback.
  • Design, build, and operate a managed Slurm service supporting research and model‑training workloads.
  • Own Slurm controllers, accounting, partitions, login nodes, node onboarding, acceptance testing, driver and CUDA baselines, upgrades, stalled-job detection, node health, draining, autohealing, storage visibility, identity, and workload isolation.
  • Lead GPU infrastructure operations on bare-metal environments, including NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn‑in, and acceptance processes.
  • Own Kubernetes cluster bootstrap and lifecycle on partner‑provided bare metal, including NVIDIA GPU Operator and Network Operator.
  • Oversee GPU isolation using technologies such as KubeVirt and VFIO and manage day‑two infrastructure operations, upgrades, backup, recovery, and node replacement.
  • Define managed inference architecture covering serving, multi‑GPU and multi‑node parallelism, autoscaling, request routing, endpoint reliability, and confidential‑compute capabilities.
  • Establish observability across the control plane, GPU fleet, and application layers through metrics, logging, alerting, and SLOs.
  • Lead incident response, post‑incident reviews, and the development of an on‑call model that is sustainable for a lean engineering organization.
  • Act as the primary technical interface with infrastructure partners and vendors, translating requirements into written specifications and acceptance tests.
  • Manage technical escalations with partners through resolution and contribute to capacity planning and hardware sourcing decisions.
  • Work directly with research, model‑training, and product teams to translate workloads into platform requirements and manage capacity constraints.
  • Hire and develop members of the platform team while maintaining a high technical bar.
  • Contribute to architecture decisions involving distributed systems, high‑performance computing, networking, storage, virtualization, and GPU workloads.
Requirements:

The ideal candidate combines deep hands‑on GPU and infrastructure expertise with proven technical leadership experience. They should be comfortable operating complex production systems, making architecture decisions, leading distributed teams, and remaining close to the code and infrastructure.

  • 8+ years of hands‑on engineering experience, including at least 3 years leading teams responsible for infrastructure platforms used by other teams.
  • Bachelor's or Master's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • Extensive hands‑on experience operating Slurm in production, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health, and upgrades.
  • Experience operating HPC or GPU training clusters for research or model‑development users.
  • Strong experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycles, Fabric Manager, NVSwitch, DCGM, MIG, node burn‑in, and acceptance.
  • Deep knowledge of InfiniBand, subnet configuration, RDMA, SR‑IOV, and diagnosing multi‑node NCCL performance issues.
  • Strong Linux systems expertise, including kernel modules, drivers, PCIe passthrough, vfio‑pci, cgroups, namespaces, and performance tuning.
  • Proven production Kubernetes experience covering control planes, upgrades, CNI, CSI, operators, custom controllers, and multi‑tenancy.
  • Experience with HPC storage and large‑scale data movement, including shared filesystems such as VAST, Lustre, or NFS and node‑local NVMe caching.
  • Experience distributing large model weights and datasets across multiple nodes.
  • Strong observability and operations experience with Prometheus, Grafana, Loki, or comparable platforms, including SLOs, incident response, and post‑incident reviews.
  • Working proficiency in JavaScript and Node.js sufficient to review control‑plane, CLI, and worker services and make architecture decisions.
  • Experience delivering a multi‑tenant IaaS, PaaS, research computing service, or comparable platform with resource isolation, quotas, usage metering, APIs, and CLI interfaces.
  • Demonstrated people leadership across time zones and the ability to lead cross‑functional technical reviews.
  • Strong written architecture and decision‑making skills, including documenting alternatives and trade‑offs.
  • Confidence communicating technical decisions and respectfully challenging partners or executives when necessary.
  • Excellent written and spoken English.
  • Fully remote availability with a working location between UTC and UTC+5:30 to provide overlap with teams and partners across Europe and India.
  • Willingness to travel occasionally to partner sites and team events.
Desirable experience includes:
  • Slurm operators on Kubernetes, such as Soperator or Slinky, or Kubernetes-native schedulers such as Kueue, Volcano, KAI, or Kubeflow Trainer.
  • Modern model‑serving technologies such as vLLM, SGLang, or TensorRT-LLM.
  • GPU parallelism strategies, quantization trade‑offs, and GPU memory planning.
  • Multi‑tenant GPU isolation using KubeVirt, Kata Containers, QEMU/KVM, Firecracker, or similar technologies.
  • Confidential computing technologies such as Intel TDX, AMD SEV‑SNP, or NVIDIA confidential‑compute capabilities.
  • Cluster API, kubeadm, Cilium, GPU autohealing, infrastructure as code, and GitOps.
  • Experience working on the operator side of a GPU cloud, university or national HPC center, or AI research platform.
  • Peer‑to‑peer or distributed‑systems experience.
  • Experience working with hardware providers responsible for provisioning but not operating infrastructure, including establishing contracts and acceptance tests.
Benefits:
  • 100% remote position.
  • Opportunity to lead the architecture and delivery of a sophisticated GPU infrastructure platform.
  • High‑impact technical leadership role spanning bare‑metal GPU infrastructure, Slurm, Kubernetes, inference, and observability.
  • Leadership responsibility for a distributed engineering organization across multiple technical disciplines.
  • Significant ownership over architecture, engineering standards, delivery planning, and team development.
  • Direct involvement with infrastructure partners and hardware providers.
  • Opportunity to support advanced AI research, model training, and managed inference workloads.
  • International and distributed working environment with colleagues and partners across Europe and India.
  • Occasional opportunities for travel to partner locations and team events.
  • Opportunity to work at the intersection of high-performance computing, AI infrastructure, distributed systems, and cloud-native technologies.
How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote: Technical Lead, GPU Infrastructure & HPC Platform
Remote: Technical Lead, GPU Infrastructure & HPC Platform

Jobgether • Turkey

Remote
TRY 500,000 - 900,000
Remote work
Leadership role
AI research support
+1
Software Engineer, Infrastructure
Software Engineer, Infrastructure

The Consensus • Turkey

On-site
TRY 400,000 - 800,000
Interesting and challenging work
Learning and growth opportunities
Regular team events and offsites
Software Engineer, Platform
Software Engineer, Platform

Fal.ai Inc. • Turkey

On-site
TRY 600,000 - 900,000
Interesting and challenging work
Learning and growth opportunities
Regular team events and offsites
Tech Lead/Sr Backend (Go) Developer
Tech Lead/Sr Backend (Go) Developer

Jobgether • Turkey

Remote
TRY 350,000 - 520,000
Senior leadership role
Go & Kubernetes work
Cloud platform exposure
+1
Software Engineer, Site Reliability
Software Engineer, Site Reliability

The Consensus • Turkey

On-site
TRY 600,000 - 1,000,000
Interesting and challenging work
Learning and growth opportunities
Regular team events and offsites
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Jobgether • Turkey

Remote
TRY 300,000 - 600,000
Fully remote
Leadership opportunities
Collaborative culture
AI Systems Architect
AI Systems Architect

UltaHost • Turkey

On-site
TRY 1,500,000 - 2,500,000
Head of AI Technologies
Head of AI Technologies

UltaHost • Turkey

On-site
TRY 720,000 - 1,800,000
Lead GCP DevOps Engineer
Lead GCP DevOps Engineer

EPAM Systems • Turkey

On-site
TRY 5,842,000 - 8,763,000
Lead Cloud Infrastructure Engineer (Next-Gen Virtualization & Automation)
Lead Cloud Infrastructure Engineer (Next-Gen Virtualization & Automation)

Turknet • Fatih

On-site
TRY 400,000 - 700,000
Continuous learning opportunities
Empowerment and significant responsibility
Dynamic and fun work environment