Linux Admin

TechDigital Group

United States

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A cloud infrastructure company in the United States is seeking a Cloud Infrastructure Engineer to manage and optimize scalable and secure cloud environments, particularly focusing on GPU workloads. The ideal candidate has over three years of experience in DevOps or site reliability engineering, with a strong background in Linux administration and knowledge of high-speed networking technologies. This role presents an opportunity to work in a dynamic environment supporting cutting-edge AI applications.

Qualifications

  • 3+ years of experience in DevOps, SRE, or cloud infrastructure management.
  • Strong knowledge of Linux system administration.
  • Experience with high-performance networking technologies.

Responsibilities

  • Provision and maintain scalable cloud infrastructure for AI workloads.
  • Administer and optimize Linux-based servers for GPU compute.
  • Implement network security measures in GPU environments.

Skills

DevOps
Linux Administration
High-Speed Networking
GPU Comprehension
Cloud Platforms (AWS/Azure/GCP)
Networking & Security Knowledge

Job description

Key Responsibilities
  • Infrastructure Management: Provision, deploy, and maintain scalable, secure, and high‑availability cloud infrastructure on platforms such as Cloud to support AI workloads.
  • System Management: Administer and maintain Linux‑based servers and clusters optimized for GPU compute workloads, ensuring high availability and performance.
  • GPU Infrastructure: Configure, monitor, and troubleshoot GPU hardware (e.g., NVIDIA GPUs) and related software stacks (e.g., CUDA, cuDNN) for optimal performance in AI/ML and HPC applications.
  • Troubleshooting: Diagnose and resolve hardware and software issues related to GPU compute nodes and performance issues in GPU clusters.
  • High‑Speed Interconnects: Implement and manage high‑speed networking technologies like RDMA over Converged Ethernet (RoCE) to support low‑latency, high‑bandwidth communication for GPU workloads.
  • CI/CD Pipelines: Build and optimize continuous integration and deployment (CI/CD) pipelines for testing GPU‑based servers and managing deployments using tools like GitHub Actions.
  • Monitoring & Performance: Set up and maintain monitoring, logging, and alerting systems (e.g., Prometheus, Victoria Metrics, Grafana) to track system performance, GPU utilization, resource bottlenecks, and uptime of GPU resources.
  • Security and Compliance: Implement network security measures, including firewalls, VLANs, VPNs, and intrusion detection systems, to protect the GPU compute environment and comply with standards like SOC 2 or ISO 27001.
Required Qualifications
  • Experience: 3+ years of experience in DevOps, Site Reliability Engineering (SRE), or cloud infrastructure management, with at least 1 year working on GPU‑based computer environments in the cloud.
  • Linux Administration: Strong knowledge of Linux system administration for managing network services and tools in a GPU compute environment.
  • High‑Speed Interconnects: Experience with high‑performance networking technologies like RoCE or 100 GbE Ethernet in compute‑intensive environments.
  • GPU‑Specific Networking: Proficiency with NVIDIA GPU networking technologies, such as Mellanox ConnectX adapters, and configuring Netplan to support their drivers and firmware.
  • Cloud Platforms: Hands‑on experience with at least one major cloud provider (AWS, Azure, GCP).
  • Networking & Security: Knowledge of networking concepts (VPC, subnets) and security best practices (IAM, encryption, firewall configurations).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Infrastructure Engineer, GPU & Compute
Infrastructure Engineer, GPU & Compute

Jobtailor • California (MO)

On-site
USD 110,000 - 160,000
Member of Technical Staff – AI Cloud Infrastructure
Member of Technical Staff – AI Cloud Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Network Operations Center Technician II
Network Operations Center Technician II

Cirrascale Cloud Services • Austin (TX)

On-site
USD 80,000 - 110,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
GPUaaS Kubernetes Platform Engineer
GPUaaS Kubernetes Platform Engineer

Veriipro • Irving (TX)

On-site
USD 140,000 - 180,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Compute Platform Engineer
Compute Platform Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 120,000 - 180,000