About the Role
LanceSoft is expanding its AI infrastructure practice and is looking for GPU Platform Engineers to design, build, operate and optimize GPU-accelerated platforms for enterprise clients. You will work on NVIDIA DGX and HGX based clusters, high-speed fabrics and AI workload platforms, and you will be the technical owner of platform health, performance and scalability from deployment through steady-state operations.
Required Certification
- NVIDIA-Certified Professional: AI Operations (NCP-AIO). Must be current and valid at the time of joining.
Preferred certifications
- NVIDIA-Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
- Cloud certifications (AWS, Azure or GCP)
- CKA or RHCSA
- Terraform Associate
Key Responsibilities-
Platform design and deployment
- Design and deploy GPU clusters on NVIDIA DGX, HGX and OEM platforms (Dell, HPE, Supermicro and others), including compute, fabric and storage layers.
- Provision and manage clusters using NVIDIA Base Command Manager or equivalent tooling.
- Implement multi-tenant GPU platforms with workload isolation, quotas and fair-share scheduling.
- Support both on-prem and air-gapped deployments, including offline software repositories and container registries.
Orchestration and workload management
- Configure and operate Slurm and Kubernetes (including the NVIDIA GPU Operator, Network Operator and MIG configurations).
- Set up and maintain container workflows using Docker, containerd and NVIDIA NGC containers.
- Support AI and ML teams in onboarding training and inference workloads, and in right-sizing GPU allocation.
Networking and storage
- Configure and validate InfiniBand and RoCE fabrics, NVLink and NVSwitch topologies, and manage fabrics through UFM.
- Validate collective communication performance using NCCL tests and resolve bottlenecks.
- Integrate high-throughput storage for AI (parallel file systems, NFS, object storage) and tune data pipelines.
Performance, reliability and observability
- Benchmark and tune GPU, network and storage performance for training and inference workloads.
- Build observability with DCGM, Prometheus and Grafana, and define alerts, dashboards and SLOs.
- Diagnose GPU faults, XID errors, ECC and memory issues, thermal and power problems, and drive root cause analysis.
- Act as the escalation point for the NOC and manage vendor cases with NVIDIA and OEM support.
Automation and lifecycle management
- Automate provisioning, configuration and operations using Ansible, Terraform, Python and Bash.
- Own firmware, driver, CUDA and software stack compatibility and upgrade planning.
- Maintain reference architectures, runbooks, standards and platform documentation.
- Plan capacity and forecast GPU demand with clients and internal teams.
Client and team engagement
- Contribute technical input to client solution reviews, proposals and pre-sales discussions.
- Mentor L1 to L3 NOC engineers and share operational best practice.
Required Skills
- Deep understanding of the NVIDIA AI stack: DGX, Base Command, DCGM, NCCL, GPU Operator, NGC and NVIDIA AI Enterprise.
- Strong Linux administration, including kernel, driver and package management.
- Hands-on experience with Slurm and Kubernetes in GPU environments.
- Working knowledge of InfiniBand, RoCE and high-performance networking.
- Proficiency in scripting and infrastructure as code (Python, Bash, Ansible, Terraform).
- Solid grasp of monitoring and observability tooling and incident troubleshooting.
- Strong written and verbal communication, with the ability to explain platform decisions to technical and business stakeholders.
Qualifications
- Bachelor's degree in Computer Science, IT, Electronics or a related field (or equivalent experience).
- NCP-AIO certification, current and valid.
- Willingness to travel to client sites and to join on-call rotations for critical escalations where required.
Nice to Have
- Experience with NVIDIA Triton Inference Server, NeMo or other LLM serving and training stacks.
- Experience with MLOps platforms and GPU cost or utilization management.
- Experience with air-gapped or public sector AI environments.
- Familiarity with liquid cooling, power and data center design considerations for high-density GPU racks.
- ITIL 4 Foundation