Senior Site Reliability Engineer

CSS

Dallas (TX)

Hybrid

USD 120,000 - 190,000

Full time

13 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

CSS is seeking a Senior Site Reliability Engineer to deploy, validate, and operationalize AI, HPC, Kubernetes, and enterprise infrastructure. This role transforms newly installed hardware into production-ready platforms through standardized provisioning, automation, testing, and infrastructure validation activities.

Working as part of a holistic team strategy, you will support large, complex customer deployments and ensure infrastructure environments are ready for operational handoff and

Qualifications

  • 5+ years of infrastructure engineering or SRE experience.
  • Experience deploying, supporting, or validating AI, GPU, HPC, or large-scale enterprise infrastructure environments.
  • Experience with Kubernetes, container platforms, and enterprise Linux administration.
  • Strong knowledge of server provisioning, virtualization, storage, networking, and infrastructure operations.
  • Experience with VMware ESXi, Hyper-V, KVM, or related virtualization technologies.
  • Experience developing automation and scripting solutions using PowerShell, Python, Bash, or similar tools.

Responsibilities

  • Provide technical expertise to support infrastructure readiness, platform engineering, and deployment activities across customer environments.
  • Deploy, configure, and validate AI, GPU, and HPC infrastructure solutions.
  • Prepare and administer Kubernetes platforms, container runtimes, storage integrations, networking components, and cluster infrastructure.
  • Install, configure, and validate NVIDIA technologies including GPU drivers, CUDA, GPU Operators, and telemetry solutions.
  • Validate high-performance networking technologies and GPU-to-GPU communications.
  • Develop automation workflows using PowerShell, Python, Bash and Infrastructure-as-Code methodologies.
  • Create customer-facing deployment documentation and readiness assessments.
  • Troubleshoot hardware, OS, virtualization, containerization, networking, and AI platform issues.

Skills

Customer-centric mindset
Written and verbal communication
Troubleshooting
Problem-solving
Communication skills

Education

Associate degree / College Diploma
Bachelor's degree in CS/IT/Engineering

Tools

Kubernetes
VMware ESXi
Hyper-V
KVM
PowerShell
Python
Bash

Job description

As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible for deploying, validating, and operationalizing AI, HPC, Kubernetes, and enterprise infrastructure environments. This role transforms newly installed hardware into production-ready platforms through standardized provisioning, automation, testing, and infrastructure validation activities. Working as part of a holistic team strategy, you will support large, complex customer deployments and ensure infrastructure environments are ready for operational handoff and long-term success.

Responsibilities:

  • Provide technical expertise and engagement to support infrastructure readiness, platform engineering, and deployment activities across customer environments.
  • Deploy, configure, and validate AI, GPU, and High Performance Computing (HPC) infrastructure solutions.
  • Prepare and administer Kubernetes platforms, container runtimes, storage integrations, networking components, and cluster infrastructure.
  • Install, configure, and validate NVIDIA technologies including GPU drivers, CUDA, GPU Operators, AI Enterprise prerequisites, and telemetry solutions.
  • Validate accelerated networking technologies including InfiniBand, RoCE, RDMA, and GPU‑to‑GPU communications.
  • Perform infrastructure readiness assessments, burn‑in testing, operational acceptance testing, and performance validation activities.
  • Configure and support server infrastructure including iDRAC, iLO, BMC, firmware, storage, and networking components.
  • Deploy and administer Windows, Linux, VMware ESXi, Hyper‑V, and KVM‑based environments.
  • Apply security hardening standards, compliance requirements, and operational best practices throughout deployment and validation activities.
  • Develop and maintain automation workflows utilizing PowerShell, Python, Bash, and Infrastructure‑as‑Code methodologies.
  • Create customer‑facing deployment documentation, technical reports, readiness assessments, and operational validation deliverables.
  • Troubleshoot complex hardware, operating system, virtualization, containerization, networking, and AI platform issues.
  • Participate in advanced technical training and continued education to maintain expertise in cloud, infrastructure, AI, and platform technologies.
  • Support technical engagements across customer environments and collaborate with internal engineering, architecture, and service delivery teams.

Qualifications:

  • Associate degree (U.S.)/College Diploma (Canada) or equivalent combination of education and technical experience required.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related technical discipline preferred.
  • 5+ years of experience in Infrastructure Engineering, Platform Engineering, Site Reliability Engineering (SRE), Systems Administration, or related technical roles.
  • Experience deploying, supporting, or validating AI, GPU, HPC, or large‑scale enterprise infrastructure environments.
  • Experience with Kubernetes, container platforms, and enterprise Linux administration.
  • Strong knowledge of server provisioning, virtualization, storage, networking, and infrastructure operations.
  • Experience with VMware ESXi, Hyper‑V, KVM, or related virtualization technologies.
  • Experience developing automation and scripting solutions using PowerShell, Python, Bash, or similar tools.
  • Knowledge of Infrastructure‑as‑Code and automated deployment methodologies.
  • Experience with NVIDIA GPU technologies, CUDA, AI Enterprise, or related AI infrastructure platforms preferred.
  • Knowledge of InfiniBand, RDMA, RoCE, or high‑performance networking technologies preferred.
  • Demonstrated troubleshooting, root‑cause analysis, and problem‑solving skills.
  • Possess a customer‑centric mindset and strong written and verbal communication skills.
  • Possess intermediate computer skills, including proficiency with Microsoft Office applications.
  • Ability to travel up to 25%.

Preferred Certifications

  • Certified Kubernetes Administrator (CKA)
  • Red Hat Certified System Administrator (RHCSA) or equivalent Linux certification
  • NVIDIA certifications related to AI, GPU, or DGX platforms
  • VMware Certified Professional (VCP) or equivalent

#LI-VR1 #Hybrid

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

CSS • Atlanta (GA)

Hybrid
USD 120,000 - 180,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Senior Staff Site Reliability Operations
Senior Staff Site Reliability Operations

NVIDIA Corporation • Washington

Hybrid
USD 184,000 - 265,000
Equity
Benefits
Senior Staff Site Reliability Operations
Senior Staff Site Reliability Operations

NVIDIA Corporation • Seattle (WA)

On-site
USD 184,000 - 265,000
Equity
Benefits
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 248,000 - 397,000
Equity
Benefits
Senior Staff Site Reliability Operations
Senior Staff Site Reliability Operations

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 265,000
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA Corporation • Durham (CA), Northern (KY)

On-site
USD 152,000 - 288,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits