Staff Engineer, Senior Manager

Jobtailor

Connecticut

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced Infrastructure Engineer to design and own HPC/ML cloud platforms in AWS or Google Cloud. You will lead containerization and deployment of HPC services, translating stakeholder needs into scalable, cost-efficient infrastructures.

The role focuses on IaC with Terraform, Kubernetes, and monitoring stacks, ensuring production reliability across compute and storage layers. Responsibilities include building reusable modules, managing lifecycle from provisioning to

Qualifications

  • BS in computer science, life science, data science or similar fields.
  • 6+ years of cloud infra engineering experience with robust IaC deployments.
  • Experience running scientific computing workloads in enterprise environments.
  • Advanced experience with AWS or GCP including core compute/storage services for HPC.
  • Knowledge of cloud networking, identity, and security controls.
  • Familiarity with HPC tools (Slurm) and GPU computing.

Responsibilities

  • Design, implement, operate, and own robust HPC/ML infra in AWS/GCP.
  • Lead containerization and deployment of HPC platforms (Slurm, Open On Demand, Prometheus/Grafana).
  • Translate stakeholder input into scalable, cost-efficient computing platforms.
  • Develop IaC workflows and reusable Terraform modules; enforce standards.
  • Operationalize Docker/Kubernetes-based solutions; manage lifecycle from provisioning to teardown.
  • Monitor, log, and alert infrastructure (CloudWatch, Prometheus, Grafana).

Skills

Cloud Infra
IaC (Infrastructure as Code)
Terraform
CloudFormation
Kubernetes
Docker
Slurm
HPC
AWS
GCP
EKS/GKE
GPU Computing

Education

B.S. in Computer Science

Tools

Terraform
CloudFormation
Slurm
Docker
Kubernetes
Prometheus
Grafana
AWS ParallelCluster
GCP Cluster Toolkit

Job description

  • Design, implement, operate, and own robust and dependable infrastructure for HPC and ML/AI workloads in a cloud environment (AWS/GCP)
  • Lead containerization, deployment, and operation of user- and admin-facing HPC platforms (Slurm, Open On Demand, Prometheus/Grafana)
  • Translate stakeholder input into robust, high-performance, scalable, cost-effective computing platforms
  • Partner with HPC specialists to capture institutional knowledge and manual processes in IaC workflows
  • Develop and maintain infrastructure automation using IaC tools like Terraform and CloudFormation
  • Create reusable Terraform modules and enforce standards
  • Operationalize containerized solutions using Docker and Kubernetes
  • Own the full lifecycle of infrastructure management, from provisioning to operations, support, updating, and teardown of production computing platforms
  • Develop and maintain monitoring, logging, and alerting for the infrastructure (e.g., CloudWatch, Prometheus/Grafana)
Requirements
  • B.S. in computer science, life science, data science or similar fields
  • 6+ years of experience in cloud infrastructure engineering with a proven track record of developing and supporting robust IaC deployments
  • Experience managing scientific computing workloads in an enterprise environment
  • Advanced experience with at least one of AWS and GCP, including knowledge of core compute and storage services relevant to HPC
  • Solid understanding of cloud networking, identity, and security controls
  • Prior experience with HPC deployment utilities including AWS ParallelCluster, AWS Parallel Computing Services, and Google Cloud Cluster Toolkit (preferred)
  • Proficiency with distributed computing environments, especially EKS/GKE/Kubernetes (preferred)
  • Familiarity with HPC environments, job schedulers (Slurm), HPC application containers (Docker, Singularity, Apptainer) and NVIDIA GPU computing (preferred)
Core Competencies

Demonstrates expertise in designing and managing cloud infrastructure for HPC and ML/AI workloads, utilizing IaC tools like Terraform and CloudFormation. Proficient in containerization and orchestration technologies such as Docker and Kubernetes, with a strong focus on operational efficiency and performance optimization.

Highest-signal resume keywords
  • Cloud Infrastructure Engineering
  • Infrastructure As Code (IaC)
  • HPC Deployment Utilities
  • Containerization (Docker, Kubernetes)
  • Monitoring and Logging (CloudWatch, Prometheus)
ATS Optimization Keywords
Hard Skills
  • Infrastructure Management
  • Cloud Networking
  • AWS
  • GCP
  • Terraform
  • CloudFormation
  • Slurm
  • Docker
  • Kubernetes
  • NVIDIA GPU Computing
Industry Keywords
  • High-Performance Computing (HPC)
  • Machine Learning (ML)
  • Artificial Intelligence (AI)
  • Distributed Computing
  • Enterprise Environment
Tools & Technologies
  • AWS ParallelCluster
  • AWS Parallel Computing Services
  • Google Cloud Cluster Toolkit
  • Prometheus
  • Grafana
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff – AI Cloud Infrastructure
Member of Technical Staff – AI Cloud Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Virtualization & Orchestration Engineer
Virtualization & Orchestration Engineer

Jobtailor • Bellevue (WA)

On-site
USD 150,000 - 210,000
AI and ML Infra Software Engineer, GPU Clusters
AI and ML Infra Software Engineer, GPU Clusters

Jobtailor • California (MO)

On-site
USD 120,000 - 190,000
Systems Engineer
Systems Engineer

Jobtailor • Nashville (TN)

Hybrid
USD 130,000 - 185,000
Staff Engineer, CI/CD & Cloud Infrastructure
Staff Engineer, CI/CD & Cloud Infrastructure

San Diego Stealth Startup • San Diego (CA)

On-site
USD 175,000 - 185,000
Software Engineer – Deployment
Software Engineer – Deployment

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 210,000
Principal Cloud Engineer – AI
Principal Cloud Engineer – AI

Jobtailor • West Chester

On-site
USD 150,000 - 210,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Software Engineer – AI & Cloud Engineering
Software Engineer – AI & Cloud Engineering

Jobtailor • Massachusetts

On-site
USD 110,000 - 170,000
Cloud Engineer
Cloud Engineer

Jobtailor • Idaho Falls (ID)

On-site
USD 120,000 - 150,000