Senior Cloud HPC Architect for AWS/EDA Workloads

Ayar Labs

San Jose (CA)

On-site

USD 120,000 - 150,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ayar Labs is seeking a Senior Cloud HPC Platform Engineer to design, build, and operate a cloud HPC platform in AWS, migrating from a Red Hat environment. You will own workloads, licenses, and data movement, chair capacity planning, and implement scalable compute with Slurm as the central scheduler.

You will collaborate with IT, design teams, and EDA partners, delivering self‑service access, secure connectivity, and reliable, cost‑aware operations.

Qualifications

  • 7+ years building and operating Linux infrastructure, incl. 3+ years in AWS or similar cloud.
  • Deep hands-on experience with Enterprise Linux in production as a System Administrator.
  • Experience designing or operating HPC, batch compute, large-scale simulation.
  • Deep production experience administering Slurm as the primary HPC scheduler, incl. clusters and policies.
  • Strong AWS experience across EC2, IAM, VPC, S3, CloudWatch, SSM, KMS, and high-performance storage.
  • Strong IaC skills using Terraform or OpenTofu; reusable modules, state, testing.
  • Experience automating Linux images with Packer, Ansible, Python, and Bash.
  • Knowledge of cloud networking, DNS, routing, firewalls, VPN, and hybrid connectivity.
  • Experience with FlexNet/FlexLM network licensing.
  • Monitoring, alerting, incident response, capacity management, and cost controls.

Responsibilities

  • Lead the AWS migration: inventory workloads, data, licenses, requirements; plan waves, cutovers, and rollbacks.
  • Build and operate scalable AWS compute with EC2, accelerated networking, autoscaling, and isolation.
  • Own scheduling and job execution using Slurm as primary control plane; deploy via Slurm/ParallelCluster.
  • Translate run manifests into CPU/core, memory, GPU, wall-time, storage I/O, and budget needs.
  • Design high‑performance storage and tiering across FSx, EFS, S3, and on-premise systems.
  • Create reproducible RHEL-compatible environments for Cadence, Synopsys, Ansys, and tools.
  • Manage licenses ensure access across hybrid and cloud environments; monitor usage.
  • Automate infrastructure with Terraform/OpenTofu, Packer, Ansible; maintain patch lifecycle.
  • Prove performance and correctness; validate runtime, queue time, storage, cost, and results.
  • Deliver self-service VDI/DVI access; publish apps, images, SSO, MFA, and policy governance.
  • Ensure security, least-privilege access, logging, and secure connectivity (VPN/Direct Connect).
  • Operate for reliability; runbooks, incident response, disaster recovery testing, capacity plans.
  • Control cloud cost with tagging, budgets, and cost/performance optimization; use Slurm accounting.
  • Eliminate fragile manual steps; implement versioned automation and clear documentation.

Skills

Linux infrastructure
AWS
Slurm
Terraform/OpenTofu
Packer/Ansible/Python/Bash
HPC/batch compute
VDI/DVI administration
Cloud cost optimization
Monitoring/incident response
Documentation & collaboration

Education

Bachelor's degree in Computer Science, Engineering, Information Systems, or related field

Tools

Terraform/OpenTofu
Packer
Ansible
AWS ParallelCluster
Citrix Virtual Apps and Desktops

Job description

Ayar Labs is seeking a Senior Cloud HPC Platform Engineer to design, build, and operate a cloud HPC platform in AWS, migrating from a Red Hat environment. You will own workloads, licenses, and data movement, chair capacity planning, and implement scalable compute with Slurm as the central scheduler.

You will collaborate with IT, design teams, and EDA partners, delivering self‑service access, secure connectivity, and reliable, cost‑aware operations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Engineer, Cloud HPC Platform
Sr. Engineer, Cloud HPC Platform

Ayar Labs • San Jose (CA)

On-site
USD 120,000 - 150,000
Cloud HPC-EDA GTM Strategist for Chip Design
Cloud HPC-EDA GTM Strategist for Chip Design

Amazon Web Services (AWS) • Chicago (IL)

On-site
USD 148,000 - 200,000
Health benefits
RSUs
HPC EDA GTM Lead for Cloud Chip Design
HPC EDA GTM Lead for Cloud Chip Design

Amazon Web Services (AWS) • Arlington (VA)

On-site
USD 148,000 - 200,000
Cloud-Scale HPC-EDA GTM Specialist
Cloud-Scale HPC-EDA GTM Specialist

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 148,000 - 200,000
Senior HPC Cloud Solutions Architect (CAE/ML)
Senior HPC Cloud Solutions Architect (CAE/ML)

Amazon Web Services (AWS) • Herndon (VA)

On-site
USD 154,000 - 208,000
Cloud GTM Lead for HPC EDA & Semiconductor Workloads
Cloud GTM Lead for HPC EDA & Semiconductor Workloads

Amazon Web Services (AWS) • East Palo Alto (CA)

On-site
USD 163,000 - 220,000
HPC-EDA GTM Lead for Cloud-Based Chip Design
HPC-EDA GTM Lead for Cloud-Based Chip Design

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 148,000 - 200,000
Senior Cloud HPC Solutions Architect
Senior Cloud HPC Solutions Architect

Rescale • United States

Remote
USD 120,000 - 160,000
HPC EDA Cloud GTM Leader for Semiconductors
HPC EDA Cloud GTM Leader for Semiconductors

Amazon Web Services (AWS) • Boston (MA)

On-site
USD 148,000 - 200,000
AWS HPC Cloud Engineer — Scalable Secure Infra
AWS HPC Cloud Engineer — Scalable Secure Infra

Jobtailor • Illinois

On-site
USD 140,000 - 190,000