An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Kyndryl is seeking a senior engineer to own production-grade NVIDIA AI platforms in Jakarta. You will tackle cross-layer issues across Linux, CUDA, Docker/Kubernetes, and GPU health, shaping standards for reliable training and inference environments.
You will lead RCA, build runbooks, mentor engineers, and drive platform upgrades and performance investigations across multi-GPU systems.
In Person
Jakarta (KID51709) GOWork
Full time
R-69816
At Kyndryl, we run and reimagine the mission-critical technology systems that drive advantage for the world’s leading businesses. We are at the heart of progress; with proven expertise and a continuous flow of AI-powered insight, enabling smarter decisions, faster innovation, and a lasting competitive edge. For our people—Kyndryls—that means doing purposeful work that powers human progress. Join us and experience a flexible, supportive environment where your well-being is prioritized and your potential can thrive.
At Kyndryl, we run and reimagine the mission-critical technology systems that drive advantage for the world’s leading businesses. We are at the heart of progress; with proven expertise and a continuous flow of Artificial Intelligence (AI)-powered insight, enabling smarter decisions, faster innovation, and a lasting competitive edge. For our people - Kyndryls - that means doing purposeful work that powers human progress. Join us and experience a flexible, supportive environment where your well-being is prioritized and your potential can thrive.
Ready to become the technical authority behind production-grade NVIDIA AI platforms? In this senior engineering role, you will solve the toughest cross-layer issues across Linux, NVIDIA drivers, Compute Unified Device Architecture (CUDA), container platforms, Graphics Processing Unit (GPU) health and distributed AI workloads. You will shape the standards that keep training and inference environments reliable, observable and ready to scale.
Lead Root Cause Analysis (RCA), automate recurring operations, develop Level 1 (L1) and Level 2 (L2) runbooks, mentor engineers and act as technical authority during Priority 0 / Priority 1 (P0/P1) incidents.
Define platform compatibility, upgrade, canary and rollback standards, and design GPU observability and remediation thresholds using DCGM, Prometheus, Grafana or equivalent tooling.
Architect and optimize local or cloud-based AI training and inference platforms using PyTorch and NVIDIA technologies such as TensorRT, Triton, NVIDIA Inference Microservices (NIM), NVIDIA GPU Cloud (NGC) and NVIDIA NeMo.
Lead NVIDIA Collective Communications Library (NCCL) and distributed multi-GPU or multi-node performance investigations across compute, network/fabric and storage.
Lead complex diagnosis across Linux, NVIDIA drivers, Compute Unified Device Architecture (CUDA), container runtime, Kubernetes or Red Hat OpenShift, NVIDIA GPU Operator, Data Center GPU Manager (DCGM) and Artificial Intelligence (AI) workloads, including recurring NVIDIA driver error identifiers (XIDs) and Multi-Instance GPU (MIG) issues.
This role gives you the opportunity to deepen your expertise in accelerated computing, production AI infrastructure and large-scale platform engineering. You can grow as a technical authority, broaden into AI infrastructure architecture or technical leadership, and develop experience across complex customer environments and emerging NVIDIA technologies.
You’re good at what you do and possess the required experience to prove it. However, equally as important - you have a growth mindset; keen to drive your own personal and professional development. You are customer-focused - someone who prioritizes customer success in your work. And finally, you’re open and borderless - naturally inclusive in how you work with others.
Strong Python and Bash automation, observability engineering and deep platform troubleshooting. General AI application or Retrieval-Augmented Generation (RAG) development without NVIDIA platform depth is insufficient.
Strong NCCL, multi-GPU or multi-node performance troubleshooting, platform lifecycle upgrades, canary and rollback experience.
Strong PyTorch expertise plus one or more of vLLM, llama.cpp, TensorRT, Triton, NIM, NGC or NeMo.
Expert Linux, NVIDIA driver, CUDA, container runtime and Kubernetes knowledge, with deep hands-on experience in NVIDIA GPU Operator, DCGM and multi-GPU scheduling and health.
5+ years building and optimizing production AI training or inference platforms, with substantial operations responsibility.
Experience with Red Hat OpenShift and production Kubernetes environments.
Experience troubleshooting network/fabric and storage dependencies affecting distributed GPU workloads.
Experience with NVIDIA DGX systems or large enterprise multi-GPU platforms.
Experience mentoring engineers, developing operational standards and leading critical incidents.
The “Kyn” in Kyndryl means kinship, which represents the strong bonds we have with each other, our customers and our communities. We focus on ensuring all Kyndryls feel included and we welcome people of all cultures, backgrounds, and experiences. Even if you don’t meet every requirement, we encourage you to apply. We believe in growth, and we’re excited to see what you can bring. At Kyndryl, employee feedback has told us that our number one driver of employee engagement is belonging. That sense of belonging - being a valued, respected, trusted member of the team - is fundamental to our culture and fueling great experiences for our customers. This dedication to welcoming everyone into our company means that Kyndryl gives you the ability to thrive and contribute to our culture of empathy and shared success. That’s The Kyndryl Way.
Your career with us isn’t just a job - it’s an adventure with purpose. We offer a dynamic, hybrid-friendly culture that supports your well-being and empowers you to grow. Our Be Well programs support your financial, mental, physical, and social health. From your first day, you’ll contribute to impactful work that powers systems our customers rely on. You’ll have tools to chart your career path, development goals aligned with your ambitions, continuous feedback, and access to learning opportunities and hands‑on experiences that help you build in-demand skills.
The “Kyn” in Kyndryl means kinship, which represents the strong bonds we have with each other, our customers and our communities. We focus on ensuring all Kyndryls feel included and we welcome people of all cultures, backgrounds, and experiences. Even if you don’t meet every requirement, we encourage you to apply. We believe in growth, and we’re excited to see what you can bring. At Kyndryl, employee feedback has told us that our number one driver of employee engagement is belonging. That sense of belonging — being a valued, respected, trusted member of the team — is fundamental to our culture and fueling great experiences for our customers. This dedication to welcoming everyone into our company means that Kyndryl gives you the ability to thrive and contribute to our culture of empathy and shared success. That’s The Kyndryl Way.
Your career with us isn’t just a job—it’s an adventure with purpose. We offer a dynamic, hybrid-friendly culture that supports your well-being and empowers you to grow. Our Be Well programs are thoughtfully designed to support your financial, mental, physical, and social health—because we know that when you feel your best, you do your best.
From your very first day, you’ll dive into impactful work that powers the systems our customers rely on every day. You won’t just contribute—you’ll make a difference, tackling meaningful projects that sharpen your skills and fuel your growth.
We’re here to champion your journey. With powerful tools to chart your career path, personalized development goals aligned with your ambitions, and continuous feedback to keep you inspired and on track, you’ll have everything you need to thrive and evolve. You’ll develop in-demand skills to grow your career and achieve your ambitions with access to cutting‑edge learning opportunities—from certifications with Microsoft, Google, and Amazon to coaching and hands‑on experiences. And through it all, you’ll be part of a culture that values empathy, restless learning, and a devotion to shared success.
We want you to thrive here—and we’re committed to helping you do just that. Ready to make an impact? Join us and help shape what’s next.
Can’t find a suitable job opening?
At Kyndryl, we achieve progress the world depends on, with purpose. Beginning with the purpose that matters to you . Because here, you will be part of a culture designed with purpose. One that is restless, empathetic and devoted. Where we are committed to sustainable progress for our customers and supporting the communities where we work and live. All of you is what we want. And what we need. Join us, and together, we can advance the vital systems that power human progress.
Kyndryl is committed to creating a diverse environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, pregnancy, disability, age, veteran status, or other characteristics. Kyndryl is also committed to compliance with all fair employment practices regarding citizenship and immigration status.