Heavy AWS + HPC

Zeal Solutions Inc

Charlotte (NC)

Remote

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Zeal Solutions Inc. is seeking a DevOps Engineer with strong HPC expertise to design, automate, and maintain compute-focused infrastructure for pharma/ life sciences workloads. The role blends DevOps with HPC cluster management and cloud-based AI/ML support.

You will design scalable cloud infrastructure, work with SLURM/PBS schedulers, and build CI/CD for ML model training and deployment, ensuring secure, cost-efficient operations in a remote US setting.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
  • Proven experience in cloud computing (AWS, Azure, GCP) and cloud architecture.
  • Strong background in AI/ML technologies, with experience in deploying ML models.
  • Proficiency in scripting languages (Python, Bash) and containerization technologies (Docker, Kubernetes).
  • Proficiency with virtual compute environments (EC2).
  • Hands-on experience with High Performance Computing (HPC) and server node Cluster Management
  • Strong Knowledge of Linux/Unix operating systems (RHEL/Ubuntu)
  • Experience with job schedulers (like SLURM, PBS), resource management, and system monitoring tools (DynaTrace).
  • Understanding of storage solutions and file systems used in HPC (such as Lustre, GPFS).
  • Experience with infrastructure as code (IaC) tools like Terraform or CloudFormation.
  • Knowledge of networking, security, and database technologies in a cloud environment.
  • Excellent problem-solving, communication, and team collaboration skills.

Responsibilities

  • Design, implement, and manage cloud-based infrastructure that supports AI/ML workloads.
  • Collaborate with data scientists and ML engineers to deploy scalable machine learning models into production.
  • Ensure the security, scalability, and reliability of AI/ML systems in the cloud.
  • Optimize cloud resources for cost-effective and efficient use.
  • Stay current with the latest in cloud services, AI/ML tools, and industry best practices.
  • Provide technical leadership and guidance in cloud and AI/ML architecture.
  • Develop and maintain CI/CD pipelines for AI/ML model training and deployment.
  • Monitor and troubleshoot AI/ML applications and cloud environments.
  • Document system design and operational procedures.
  • Collaborate with AI/ML and HPC teams to understand their computing and storage needs.

Skills

Python
Bash
Docker
Kubernetes
EC2
SLURM
PBS
Terraform
CloudFormation
RHEL
HPC
Git

Education

Bachelor’s degree in Computer Science

Tools

Terraform
CloudFormation
DynaTrace

Job description

Heavy AWS + HPC ==== High Performance Computing, Parallel Cluster * Slrum is a must

Client: Pharma Client

Location: US, Remote

We're seeking a DevOps Engineer with High Performance Computing (HPC) expertise to design, automate, and maintain the infrastructure supporting the compute-intensive scientific and research workloads (e.g., genomics, molecular modeling, drug discovery simulations). This role bridges traditional DevOps practices with specialized HPC cluster management in a life sciences/pharma environment.

Key Responsibilities:
  • Design, implement, and manage cloud-based infrastructure that supports AI/ML workflows. for
  • Collaborate with data scientists and ML engineers to deploy scalable machine learning models into production.
  • Ensure the security, scalability, and reliability of AI/ML systems in the cloud.
  • Optimize cloud resources for cost-effective and efficient use.
  • Stay current with the latest in cloud services, AI/ML tools, and industry best practices.
  • Provide technical leadership and guidance in cloud and AI/ML architecture.
  • Develop and maintain CI/CD pipelines for AI/ML model training and deployment.
  • Monitor and troubleshoot AI/ML applications and cloud environments.
  • Document system design and operational procedures.
  • Collaborate with AI/ML and HPC teams to understand their computing and storage needs.
Qualifications:
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
  • Proven experience in cloud computing (AWS, Azure, GCP) and cloud architecture.
  • Strong background in AI/ML technologies, with experience in deploying ML models.
  • Proficiency in scripting languages (Python, Bash) and containerization technologies (Docker, Kubernetes).
  • Proficiency with virtual compute environments (EC2).
  • Hands-on experience with High Performance Computing (HPC) and server node Cluster Management
  • Strong Knowledge of Linux/Unix operating systems (RHEL/Ubuntu)
  • Experience with job schedulers (like SLURM, PBS), resource management, and system monitoring tools (DynaTrace).
  • Understanding of storage solutions and file systems used in HPC (such as Lustre, GPFS).
  • Experience with infrastructure as code (IaC) tools like Terraform or CloudFormation.
  • Knowledge of networking, security, and database technologies in a cloud environment.
  • Excellent problem-solving, communication, and team collaboration skills.
Preferred Skills:
  • Familiarity with machine learning frameworks (TensorFlow, PyTorch) and data pipelines.
  • Certifications in cloud architecture (AWS Certified Solutions Architect, Google Cloud Professional Cloud Architect, etc.).
  • Experience in an Agile development environment.
  • Prior work with distributed computing and big data technologies (Hadoop, Spark).
  • Operational experience running large scale platforms, including AI/ML platforms
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote HPC & AWS DevOps Engineer for Pharma AI/ML
Remote HPC & AWS DevOps Engineer for Pharma AI/ML

Zeal Solutions Inc • Charlotte (NC)

Remote
USD 120,000 - 180,000
HPC (High-Performance Computing) Consultant @ Remote
HPC (High-Performance Computing) Consultant @ Remote

BURGEON IT SERVICES LLC • United States

Remote
USD 120,000 - 180,000
HPC on AWS Lead /Specialist/ SME- REMOTE
HPC on AWS Lead /Specialist/ SME- REMOTE

Simple Solutions • Jacksonville (FL)

Remote
USD 140,000 - 210,000
Lead HPC and Systems Engineer
Lead HPC and Systems Engineer

EPAM Systems • United States

On-site
USD 150,000 - 190,000
ZR_2817_JOB
ZR_2817_JOB

Zohorecruit • Jacksonville (FL)

Remote
USD 150,000 - 210,000
HPC Systems Administrator
HPC Systems Administrator

Eli Lilly and Company • California (MO)

On-site
USD 150,000 - 190,000
Hybrid work model
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
HPC Customer Solutions Engineer
HPC Customer Solutions Engineer

GTN Technical Staffing • United States

On-site
USD 120,000 - 180,000
Senior HPC Engineer
Senior HPC Engineer

RCH Solutions • United States

Remote
USD 120,000 - 180,000
Competitive salary + bonus
Health and wellness benefits
401(k) plan with match
+2
Platform Architect - HPC, Kubernetes
Platform Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 140,000 - 230,000