Sr. Engineer, Cloud HPC Platform

Ayar Labs

San Jose (CA)

On-site

USD 120,000 - 150,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ayar Labs is seeking a Senior Cloud HPC Platform Engineer to design, build, and operate a cloud HPC platform in AWS, migrating from a Red Hat environment. You will own workloads, licenses, and data movement, chair capacity planning, and implement scalable compute with Slurm as the central scheduler.

You will collaborate with IT, design teams, and EDA partners, delivering self‑service access, secure connectivity, and reliable, cost‑aware operations.

Qualifications

  • 7+ years building and operating Linux infrastructure, incl. 3+ years in AWS or similar cloud.
  • Deep hands-on experience with Enterprise Linux in production as a System Administrator.
  • Experience designing or operating HPC, batch compute, large-scale simulation.
  • Deep production experience administering Slurm as the primary HPC scheduler, incl. clusters and policies.
  • Strong AWS experience across EC2, IAM, VPC, S3, CloudWatch, SSM, KMS, and high-performance storage.
  • Strong IaC skills using Terraform or OpenTofu; reusable modules, state, testing.
  • Experience automating Linux images with Packer, Ansible, Python, and Bash.
  • Knowledge of cloud networking, DNS, routing, firewalls, VPN, and hybrid connectivity.
  • Experience with FlexNet/FlexLM network licensing.
  • Monitoring, alerting, incident response, capacity management, and cost controls.

Responsibilities

  • Lead the AWS migration: inventory workloads, data, licenses, requirements; plan waves, cutovers, and rollbacks.
  • Build and operate scalable AWS compute with EC2, accelerated networking, autoscaling, and isolation.
  • Own scheduling and job execution using Slurm as primary control plane; deploy via Slurm/ParallelCluster.
  • Translate run manifests into CPU/core, memory, GPU, wall-time, storage I/O, and budget needs.
  • Design high‑performance storage and tiering across FSx, EFS, S3, and on-premise systems.
  • Create reproducible RHEL-compatible environments for Cadence, Synopsys, Ansys, and tools.
  • Manage licenses ensure access across hybrid and cloud environments; monitor usage.
  • Automate infrastructure with Terraform/OpenTofu, Packer, Ansible; maintain patch lifecycle.
  • Prove performance and correctness; validate runtime, queue time, storage, cost, and results.
  • Deliver self-service VDI/DVI access; publish apps, images, SSO, MFA, and policy governance.
  • Ensure security, least-privilege access, logging, and secure connectivity (VPN/Direct Connect).
  • Operate for reliability; runbooks, incident response, disaster recovery testing, capacity plans.
  • Control cloud cost with tagging, budgets, and cost/performance optimization; use Slurm accounting.
  • Eliminate fragile manual steps; implement versioned automation and clear documentation.

Skills

Linux infrastructure
AWS
Slurm
Terraform/OpenTofu
Packer/Ansible/Python/Bash
HPC/batch compute
VDI/DVI administration
Cloud cost optimization
Monitoring/incident response
Documentation & collaboration

Education

Bachelor's degree in Computer Science, Engineering, Information Systems, or related field

Tools

Terraform/OpenTofu
Packer
Ansible
AWS ParallelCluster
Citrix Virtual Apps and Desktops

Job description

Location: San Jose, CA

Job Id:686

# of Openings:0

Ayar Labs is shattering AI data bottlenecks by moving data at the speed of light. As pioneers of co-packaged optics (CPO), we are using light instead of electricity to move data faster, further, and with a fraction of the energy needed to fuel the explosive growth of AI models.

Backed by industry giants like NVIDIA, AMD and Intel and manufactured in partnership with the world’s leading semiconductor ecosystem, Ayar Labs’ co-packaged optics solution is key to unleashing next-generation AI scale-up architectures.

Ayar Labs is moving silicon engineering compute workloads into AWS. The Senior Cloud HPC Platform Engineer will design, build, and operate the cloud HPC platform that supports EDA, simulation, verification, physical design, AMS, and related engineering workflows.

You will lead the migration from the current RHEL-based compute environment to AWS while protecting engineering productivity, design data, license availability, and output correctness. You will work closely with IT, TFM, ASIC, AMS, verification, physical design, security, finance, and EDA vendors.

This is a hands-on infrastructure role. You will own the platform from workload discovery and architecture through migration, production operations, cost management, and on-call support.

Key Responsibilities
  • Lead the AWS migration: Inventory engineering workloads, dependencies, data, licenses, and performance requirements; define migration waves, cutover plans, rollback procedures, and acceptance criteria.
  • Build the cloud HPC platform: Design and operate scalable AWS compute using appropriate EC2 instance families, accelerated networking, autoscaling, placement strategies, and workload isolation.
  • Own scheduling and job execution: Make Slurm the primary scheduler and operational control plane for interactive, batch, regression, and multi-day simulation workloads. Deploy and operate Slurm directly and/or through AWS ParallelCluster where appropriate; configure partitions, QoS, priorities, fair-share, reservations, preemption, accounting, dependencies, job arrays, and policy-based autoscaling.
  • Plan engineering run capacity: Partner with design and verification teams before major regressions, simulations, and tapeout milestones to translate run manifests and workload forecasts into CPU/core, memory, GPU, wall-time, scratch and capacity I/O, network, license-token, Slurm partition/reservation, and budget requirements; publish capacity scenarios, reservations, and readiness risks.
  • Engineer storage and data movement: Design high-performance storage and tiering across services such as Amazon FSx, EFS, S3, and on-prem systems. Establish backup, lifecycle, replication, and recovery controls.
  • Enable EDA workloads: Build reproducible RHEL-compatible environments for Cadence, Synopsys, Ansys, and other engineering tools. Support PDKs, third‑party IP, shared flows, and controlled releases.
  • Manage licenses: Design reliable FlexNet/FlexLM access across hybrid and cloud environments, monitor utilization, and prevent licensing from becoming a scaling bottleneck.
  • Automate the environment: Define infrastructure through Terraform or OpenTofu and automate images, configuration, patching, and application deployment with tools such as Packer and Ansible.
  • Prove performance and correctness: Benchmark representative workloads before and after migration. Validate runtime, queue time, storage performance, reliability, cost, and quality-of-results with engineering owners.
  • Deliver self-service access: Architect and operate secure virtual desktop infrastructure (VDI/DVI) for engineering workflows, including Citrix Virtual Apps and Desktops and/or NICE DCV/VNC. Own application publishing, golden images, patching, SSO/MFA, session brokering and policies, profile and storage integration, GPU/graphics support, clipboard and file-transfer controls, monitoring, capacity, high availability, and performance troubleshooting; provide documented self-service paths for launching jobs and remote sessions.
  • Own security and connectivity: Implement least‑privilege IAM, network segmentation, encryption, secrets management, logging, vendor access controls, and secure connectivity through VPN and/or Direct Connect.
  • Operate for reliability: Establish observability, service objectives, incident response, runbooks, change controls, and disaster‑recovery testing. Partner with engineering on run‑demand forecasting and capacity planning for major regressions, simulations, and tapeout milestones.
  • Control cloud cost: Implement tagging, budgets, chargeback/showback, scheduling policies, idle‑resource controls, and workload‑specific cost/performance optimization. Use Slurm accounting and workload forecasts to provide engineering with resource and cost estimates before large runs.
  • Reduce operational fragility: Replace undocumented manual steps and one‑off scripts with versioned, tested, supportable automation and clear documentation.
Basic Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience.
  • 7+ years building and operating Linux infrastructure, including 3+ years in AWS or a comparable cloud environment.
  • Deep hands‑on experience with Enterprise Linux in production as a System Administrator
  • Experience designing or operating HPC, batch compute, large‑scale simulation, or similarly compute‑intensive platforms.
  • Deep production experience administering Slurm as the primary HPC scheduler, including partitions, QoS, priorities, fair‑share, reservations, preemption, accounting, job arrays and dependencies, failure recovery, upgrades, and integration with AWS ParallelCluster or equivalent cloud capacity.
  • Strong AWS experience across EC2, IAM, VPC, S3, CloudWatch, Systems Manager, KMS, and high‑performance storage services.
  • Strong infrastructure‑as‑code skills using Terraform or OpenTofu, including reusable modules, state management, review, and testing.
  • Experience automating Linux images and configuration with Packer, Ansible, Python, and/or Bash.
  • Strong knowledge of high‑performance and shared storage, Linux file systems, data transfer, backup, and recovery.
  • Production experience operating secure virtual desktop infrastructure (VDI/DVI) for engineering workloads, preferably Citrix Virtual Apps and Desktops, including application publishing, image and patch lifecycle, SSO/MFA, session brokering and policy, profile and storage integration, GPU/graphics, clipboard and file‑transfer controls, monitoring, capacity and high availability, and performance troubleshooting.
  • Strong knowledge of cloud networking, DNS, routing, firewalls, VPN, and hybrid connectivity.
  • Experience supporting FlexNet/FlexLM or another network‑license system.
  • Experience establishing monitoring, alerting, incident response, capacity management, and cost controls for production infrastructure.
  • Ability to partner directly with engineers, translate run manifests and workload forecasts into CPU/core, memory, GPU, wall‑time, storage I/O and capacity, network, license‑token, Slurm partition/reservation, and budget requirements, and communicate capacity and migration risks clearly.
  • Clear written documentation, design proposals, operating procedures, and post‑incident reviews.
Preferred Qualifications
  • Experience supporting semiconductor EDA environments, including Cadence, Synopsys, Ansys, Siemens EDA, PDKs, and IP libraries.
  • Experience migrating EDA, HPC, simulation, or verification workloads from on‑prem infrastructure to AWS.
  • Experience with Amazon FSx for Lustre, FSx for OpenZFS, EFA, AWS Batch, ParallelCluster, or equivalent HPC services.
  • Experience designing and operating Citrix Virtual Apps and Desktops or comparable virtual desktop infrastructure (VDI/DVI) for engineering workloads, including application publishing, image and patch lifecycle, SSO/MFA, session brokering and policies, GPU/graphics, profile and storage integration, high availability, monitoring, and performance troubleshooting.
  • Experience benchmarking workload runtime, queue time, storage I/O, scaling efficiency, quality‑of‑results, and cost.
  • Familiarity with GitLab CI, artifact repositories, observability platforms, and controlled release processes.
  • AWS Professional or Specialty certification.

Salary range: $120,000 - $150,000

Ayar Labs is an Equal Opportunity Employer and is strongly committed to all policies which will afford equal opportunity employment to all qualified persons without regard to age, sex, national origin, race, color, ethnicity, creed, religion, gender identity, sexual orientation, disability, veteran status, or any other characteristic protected by law. It is the policy of Ayar Labs to provide reasonable accommodation when requested by a qualified applicant or employee with a disability, unless such accommodation would cause an undue hardship. Veterans are more than welcome and encouraged to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Staff Engineer, Modeling Infrastructure and Software Development
Sr. Staff Engineer, Modeling Infrastructure and Software Development

Ayar Labs • San Jose (CA)

On-site
USD 170,000 - 223,000
Engineer, Hardware Test Software
Engineer, Hardware Test Software

Ayar Labs • San Jose (CA)

On-site
USD 130,000 - 160,000
HPC-EDA Technical GTM Specialist, WWSO HPC Team
HPC-EDA Technical GTM Specialist, WWSO HPC Team

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 148,000 - 200,000
Principal Engineer, Network Protocol Architect
Principal Engineer, Network Protocol Architect

Ayar Labs • San Jose (CA)

On-site
USD 180,000 - 250,000
HPC-EDA Technical GTM Specialist, WWSO HPC Team
HPC-EDA Technical GTM Specialist, WWSO HPC Team

Amazon Web Services (AWS) • Chicago (IL)

On-site
USD 148,000 - 200,000
Health benefits
RSUs
HPC-EDA Technical GTM Specialist, WWSO HPC Team
HPC-EDA Technical GTM Specialist, WWSO HPC Team

Amazon Web Services (AWS) • Arlington (VA)

On-site
USD 148,000 - 200,000
HPC-EDA Technical GTM Specialist, WWSO HPC Team
HPC-EDA Technical GTM Specialist, WWSO HPC Team

Amazon Web Services (AWS) • Boston (MA)

On-site
USD 148,000 - 200,000
HPC-EDA Technical GTM Specialist, WWSO HPC Team
HPC-EDA Technical GTM Specialist, WWSO HPC Team

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 148,000 - 200,000
Sr. Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering
Sr. Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
Health insurance
RSUs
Career growth opportunities
Senior Cloud HPC Architect for AWS/EDA Workloads
Senior Cloud HPC Architect for AWS/EDA Workloads

Ayar Labs • San Jose (CA)

On-site
USD 120,000 - 150,000