Senior Site Reliability Engineer

CSS

Atlanta (GA)

Hybrid

USD 120,000 - 180,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

CSS seeks a Senior Site Reliability Engineer to deploy, validate, and operationalize AI, HPC, Kubernetes, and enterprise infrastructure environments. The role transforms installed hardware into production platforms through automated provisioning, testing, and validation, ensuring readiness for handoff and long-term success.

The candidate will deploy AI/HPC stacks, manage Kubernetes platforms, validate NVIDIA GPU technologies, perform readiness assessments, and develop automation with PowerShell,

Qualifications

  • 5+ years of experience in Infrastructure/Platform/SRE roles.
  • Experience deploying and validating AI, GPU, HPC, or large-scale enterprise infrastructure environments.
  • Strong Kubernetes and Linux administration skills.

Responsibilities

  • Provide technical readiness for deployment across customer environments.
  • Deploy, configure, and validate AI, GPU, and HPC infrastructure solutions.
  • Prepare and administer Kubernetes platforms, container runtimes, storage, networking, and cluster infrastructure.
  • Install and validate NVIDIA GPU technologies and drivers.
  • Perform infrastructure readiness assessments and validation activities.
  • Configure and support server infrastructure (iDRAC, iLO, BMC, firmware).
  • Deploy and administer Windows, Linux, VMware ESXi, Hyper‑V, and KVM environments.
  • Apply security hardening standards and best practices.
  • Develop automation workflows using PowerShell, Python, Bash, and IaC.
  • Create deployment documentation and technical reports.
  • Troubleshoot hardware, OS, virtualization, containerization, networking, and AI platform issues.
  • Engage in ongoing training to stay current with cloud, infrastructure, and AI platforms.
  • Collaborate with engineering, architecture, and service delivery teams.

Skills

Kubernetes
Linux Administration
Automation
Python
PowerShell
Bash
IaC
GPU/ NVIDIA CUDA
Networking
VMware ESXi
Hyper-V
KVM
iDRAC/iLO

Education

Associate degree
Bachelor's degree in CS or related field

Tools

NVIDIA GPUs
Docker

Job description

As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible for deploying, validating, and operationalizing AI, HPC, Kubernetes, and enterprise infrastructure environments. This role transforms newly installed hardware into production-ready platforms through standardized provisioning, automation, testing, and infrastructure validation activities. Working as part of a holistic team strategy, you will support large, complex customer deployments and ensure infrastructure environments are ready for operational handoff and long-term success.

Responsibilities:

  • Provide technical expertise and engagement to support infrastructure readiness, platform engineering, and deployment activities across customer environments.
  • Deploy, configure, and validate AI, GPU, and High Performance Computing (HPC) infrastructure solutions.
  • Prepare and administer Kubernetes platforms, container runtimes, storage integrations, networking components, and cluster infrastructure.
  • Install, configure, and validate NVIDIA technologies including GPU drivers, CUDA, GPU Operators, AI Enterprise prerequisites, and telemetry solutions.
  • Validate accelerated networking technologies including InfiniBand, RoCE, RDMA, and GPU‑to‑GPU communications.
  • Perform infrastructure readiness assessments, burn‑in testing, operational acceptance testing, and performance validation activities.
  • Configure and support server infrastructure including iDRAC, iLO, BMC, firmware, storage, and networking components.
  • Deploy and administer Windows, Linux, VMware ESXi, Hyper‑V, and KVM‑based environments.
  • Apply security hardening standards, compliance requirements, and operational best practices throughout deployment and validation activities.
  • Develop and maintain automation workflows utilizing PowerShell, Python, Bash, and Infrastructure‑as‑Code methodologies.
  • Create customer‑facing deployment documentation, technical reports, readiness assessments, and operational validation deliverables.
  • Troubleshoot complex hardware, operating system, virtualization, containerization, networking, and AI platform issues.
  • Participate in advanced technical training and continued education to maintain expertise in cloud, infrastructure, AI, and platform technologies.
  • Support technical engagements across customer environments and collaborate with internal engineering, architecture, and service delivery teams.

Qualifications:

  • Associate degree (U.S.)/College Diploma (Canada) or equivalent combination of education and technical experience required.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related technical discipline preferred.
  • 5+ years of experience in Infrastructure Engineering, Platform Engineering, Site Reliability Engineering (SRE), Systems Administration, or related technical roles.
  • Experience deploying, supporting, or validating AI, GPU, HPC, or large‑scale enterprise infrastructure environments.
  • Experience with Kubernetes, container platforms, and enterprise Linux administration.
  • Strong knowledge of server provisioning, virtualization, storage, networking, and infrastructure operations.
  • Experience with VMware ESXi, Hyper‑V, KVM, or related virtualization technologies.
  • Experience developing automation and scripting solutions using PowerShell, Python, Bash, or similar tools.
  • Knowledge of Infrastructure‑as‑Code and automated deployment methodologies.
  • Experience with NVIDIA GPU technologies, CUDA, AI Enterprise, or related AI infrastructure platforms preferred.
  • Knowledge of InfiniBand, RDMA, RoCE, or high‑performance networking technologies preferred.
  • Demonstrated troubleshooting, root‑cause analysis, and problem‑solving skills.
  • Possess a customer‑centric mindset and strong written and verbal communication skills.
  • Possess intermediate computer skills, including proficiency with Microsoft Office applications.
  • Ability to travel up to 25%.

Preferred Certifications

  • Certified Kubernetes Administrator (CKA)
  • Red Hat Certified System Administrator (RHCSA) or equivalent Linux certification
  • NVIDIA certifications related to AI, GPU, or DGX platforms
  • VMware Certified Professional (VCP) or equivalent

#LI-VR1 #Hybrid

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

CSS • Dallas (TX)

Hybrid
USD 120,000 - 190,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Senior Staff Site Reliability Operations
Senior Staff Site Reliability Operations

NVIDIA Corporation • Seattle (WA)

On-site
USD 184,000 - 265,000
Equity
Benefits
Senior Staff Site Reliability Operations
Senior Staff Site Reliability Operations

NVIDIA Corporation • Washington

Hybrid
USD 184,000 - 265,000
Equity
Benefits
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 248,000 - 397,000
Equity
Benefits
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA Corporation • Durham (CA), Northern (KY)

On-site
USD 152,000 - 288,000
Senior Staff Site Reliability Operations
Senior Staff Site Reliability Operations

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 265,000
Senior Staff Site Reliability Operations
Senior Staff Site Reliability Operations

NVIDIA Gruppe • Seattle (WA)

On-site
USD 184,000 - 265,000
Equity
Benefits