AI Operations & Infrastructure Engineer

Invictus International Consulting, LLC

Fort Meade (MD)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Invictus International Consulting, LLC in Fort Meade, MD is seeking an AI Operations & Infrastructure Engineer to manage and maintain advanced AI computing platforms. The role demands hands-on experience with NVIDIA technologies and requires a current TS/SCI clearance with a CI Polygraph.

The ideal candidate possesses strong skills in containerization, networking, and infrastructure management, alongside an active NVIDIA Professional Certification. This position is essential for ensuring optimal operation of AI systems.

Qualifications

  • Hold an active NVIDIA Professional Certification in AI Networking, Infrastructure, or Operations.
  • Experience administering NVIDIA GPU and DPU technologies for AI workloads.
  • Proficient in using Docker, Kubernetes, and Slurm for workload orchestration.

Responsibilities

  • Manage and maintain AI computing platforms, including GPUs.
  • Oversee the AI software stack and implement containerization technologies.
  • Diagnose and resolve networking issues impacting AI workloads.

Skills

NVIDIA Professional Certification
Hands-on experience with NVIDIA GPU
Containerization skills (Docker, Kubernetes)
Expertise in AI networking
Knowledge of Slurm and NVIDIA Base Command Manager

Tools

NVIDIA Base Command Manager
BlueField DPUs

Job description

Title: AI Operations & Infrastructure Engineer

Location: Fort Meade, MD

Clearance: TS/SCI with a CI Polygraph

Job Details
  • Manage and maintain AI computing platforms, including GPUs and other specialized hardware
  • Install and configure GPU drivers and software
  • Oversee the AI software stack and tools
  • Implement and manage containerization technologies like Docker and Kubernetes
  • Configure and optimize networking infrastructure for AI workloads, including InfiniBand and Ethernet
  • Manage storage solutions for AI data, considering performance and capacity requirements
  • Deploy and manage data processing units (DPUs) to accelerate data center workloads
  • Monitor and manage AI cluster health and resource utilization
  • Implement workload management and scheduling tools like Slurm and Kubernetes
  • Ensure efficient power and cooling for AI infrastructure to maintain optimal operating conditions
  • Configure high-performance networking solutions for AI and machine learning workloads
  • Optimize network performance to ensure maximum throughput and minimal latency for AI computations
  • Implement and fine-tune network protocols to enhance data transfer speeds and efficiency
  • Integrate NVIDIA networking products with existing AI infrastructure, including servers, GPUs, and storage systems
  • Deploy networking solutions in data centers to ensure seamless connectivity between AI components
  • Diagnose and resolve networking issues impacting AI workloads to maintain optimal system performance
  • Provide technical support and guidance to teams managing AI infrastructure
  • Collaborate with data scientists, researchers, and IT professionals to understand networking requirements and challenges
  • Lead deployment and validation of servers and systems for AI enabled platforms
  • Configure and manage network topologies, BMC, OOB, TPM, power, and cooling
  • Install, upgrade, and validate GPU-based servers, BlueField DPUs, cables, and transceivers
  • Perform firmware upgrades, hardware validation, and storage setup
  • Configure and administer physical and logical resources, including MIG partitioning and BlueField platforms
  • Install and configure operating systems, cluster software, drivers, containers (Docker), and NGC CLI
  • Manage and orchestrate clusters using NVIDIA Base Command Manager, Slurm, Pyxis, Enroot, and Run: Ai
  • Perform stress, benchmarking, and burn-in tests using HPL, NCCL, NVIDIA Nemo, and ClusterKit
  • Verify cabling, firmware/software versions, and network signal quality
  • Troubleshoot and resolve hardware, software, storage, and performance faults
  • Replace faulty components and optimize systems for AMD/Intel platforms
  • Monitor, document, and report on cluster health, resource usage, and job performance
  • Ensure secure, efficient, and scalable operation of NVIDIA AI infrastructure, including user access and workload management
Requirements
  • Qualified candidates must hold an active NVIDIA Professional Certification in either AI Networking, AI Infrastructure, or AI Operations
  • Prior direct, hands-on professional experience administering NVIDIA GPU and data processing unit (DPU) technologies, AI software stacks, and data center environments for high-performance AI workloads
  • Comprehensive expertise in deploying and maintaining AI compute platforms, requiring proficiency in containerization and workload orchestration using Docker, Kubernetes, Slurm, NVIDIA Base Command Manager, and Run:Ai
  • Must be capable of configuring physical and logical resources, including Multi-Instance GPU (MIG) partitioning and BlueField platforms, while overseeing critical facility elements such as power, cooling, and storage solutions
  • The ability to demonstrate advanced skills in AI networking, specifically configuring and optimizing high-performance InfiniBand and Ethernet fabrics to ensure maximum throughput and minimal latency
  • Current active TS/SCI clearance with a CI Polygraph

Equal Opportunity Employer/Veterans/Disabled

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Operations & Infrastructure Engineer
AI Operations & Infrastructure Engineer

Invictus International • Geraghty Village (MD)

On-site
USD 100,000 - 130,000
AI Infra & HPC Engineer — GPU/DPU, TS/SCI
AI Infra & HPC Engineer — GPU/DPU, TS/SCI

Invictus International Consulting, LLC • Fort Meade (MD)

On-site
USD 100,000 - 130,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior AI Infra & Ops Engineer – GPU, Networking, HPC
Senior AI Infra & Ops Engineer – GPU, Networking, HPC

Invictus International • Geraghty Village (MD)

On-site
USD 100,000 - 130,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Datacenter Operations Manager
Datacenter Operations Manager

Larsen & Toubro • Concord (CA)

On-site
USD 120,000 - 150,000
Medical insurance
Vision insurance
401(k)
Senior AI Compute Engineer - HPC Infra & Customer Delivery
Senior AI Compute Engineer - HPC Infra & Customer Delivery

NVIDIA • Santa Clara (CA)

On-site
Senior AI Compute Engineer - NVIS
Senior AI Compute Engineer - NVIS

NVIDIA • Santa Clara (CA)

On-site
USD 148,000 - 288,000
Equity
Benefits
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer

NVIDIA • Colorado

On-site
USD 124,000 - 196,000