HPC/AI Cloud Engineer

TEKsystems

Farmington (CT)

Remote

USD 103,000 - 117,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical
Dental & Vision
401(k)

Job summary

TEKsystems is seeking a Platform Engineer to design, deploy, and support scalable cloud infrastructure for AI, ML, and HPC workloads. You will build GPU-enabled platforms, manage Kubernetes-based container platforms, and develop IaC using Terraform across OCI, Azure, and AWS.

You will contribute to performance optimization, security, and observability while coordinating with data science and engineering teams to enable AI workloads.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or Information Technology.
  • 5+ years of experience supporting cloud infrastructure and platform engineering.
  • Experience with OCI, Kubernetes, Linux, and Terraform.
  • Experience supporting AI/ML, GPU, or HPC environments.
  • Strong scripting and automation skills using Python, Bash, or similar languages.
  • Knowledge of cloud networking, security, and platform operations.
  • Preferred: NVIDIA GPUs, CUDA, and accelerated computing platforms.
  • Experience with OCI AI Infrastructure, OCI Supercluster, or similar cloud AI services.
  • Familiarity with PyTorch, TensorFlow, Hugging Face, or MLOps platforms.
  • Experience with HPC schedulers such as Slurm.
  • Cloud certifications (Oracle, AWS, Azure, or GCP).
  • Technical Skills Required: OCI, Kubernetes, Linux, Terraform, Python, Cloud Networking, Infrastructure Automation, CI/CD

Responsibilities

  • Design, deploy, and support AI and HPC infrastructure in Cloud environments (OCI, Azure, AWS).
  • Build and optimize GPU-enabled platforms for model training, inference, and large-scale compute workloads.
  • Engineer and administer Kubernetes-based container platforms.
  • Develop Infrastructure as Code (IaC) solutions using Terraform and automation tools.
  • Optimize performance across compute, storage, networking, and orchestration layers.
  • Implement monitoring, observability, security, and operational best practices.
  • Partner with data science, engineering, and architecture teams to enable AI and analytics workloads.
  • Support platform upgrades, incident response, capacity planning, and disaster recovery.

Skills

OCI
Kubernetes
Linux
Terraform
Python
Cloud Networking
Infrastructure Automation
CI/CD

Education

Bachelor's degree in Computer Science/Engineering/IT

Tools

Docker
ArgoCD
MLflow
Prometheus
Grafana
NVIDIA GPUs
CUDA
Slurm
PyTorch
TensorFlow
OCI AI Infrastructure

Job description

Description

The Enterprise Cloud Platform team is seeking a Platform Engineer focused on building, automating, and optimizing cloud infrastructure for AI, Machine Learning, Generative AI, and High-Performance Computing (HPC) workloads. This role will design and operate scalable cloud platforms that provide GPU compute, high-performance storage, advanced networking, and cloud-native services to support enterprise AI initiatives.

Key Responsibilities
  • Design, deploy, and support AI and HPC infrastructure in Cloud environments (OCI, Azure, AWS).
  • Build and optimize GPU-enabled platforms for model training, inference, and large-scale compute workloads.
  • Engineer and administer Kubernetes-based container platforms.
  • Develop Infrastructure as Code (IaC) solutions using Terraform and automation tools.
  • Optimize performance across compute, storage, networking, and orchestration layers.
  • Implement monitoring, observability, security, and operational best practices.
  • Partner with data science, engineering, and architecture teams to enable AI and analytics workloads.
  • Support platform upgrades, incident response, capacity planning, and disaster recovery.
Skills
  • HPC, Rescale
Top Skills Details
  • HPC,Rescale
Additional Skills & Qualifications
  • Required Qualifications Bachelor's degree or equivalent experience in Computer Science, Engineering, or Information Technology.
  • 5+ years of experience supporting cloud infrastructure and platform engineering.
  • Experience with OCI, Kubernetes, Linux, and Terraform.
  • Experience supporting AI/ML, GPU, or HPC environments.
  • Strong scripting and automation skills using Python, Bash, or similar languages.
  • Knowledge of cloud networking, security, and platform operations.
  • Preferred Qualifications Experience with NVIDIA GPUs, CUDA, and accelerated computing platforms.
  • Experience with OCI AI Infrastructure, OCI Supercluster, or similar cloud AI services.
  • Familiarity with PyTorch, TensorFlow, Hugging Face, or MLOps platforms.
  • Experience with HPC schedulers such as Slurm.
  • Cloud certifications (Oracle, AWS, Azure, or GCP).
  • Technical Skills Required: OCI, Kubernetes, Linux, Terraform, Python, Cloud Networking, Infrastructure Automation, CI/CD
  • Preferred: Rescale HPC platform, Domino AI orchestration, NVIDIA GPUs, CUDA, Slurm, OCI AI Infrastructure, PyTorch, TensorFlow, Docker, ArgoCD, MLflow, Prometheus, Grafana
Experience Level

Intermediate Level

Job Type & Location

This is a Contract position based out of Farmington, CT.

Pay and Benefits

The pay range for this position is $75.00 - $85.00/hr.

Individual compensation offered for this position within this range will depend on many factors, including qualifications, skills, relevant experience, job knowledge, geographic location, internal equity, and other pertinent job-related factors.

Eligibility requirements apply to some benefits and may depend on your job classification and length of employment. Benefits are subject to change and may be subject to specific elections, plan, or program terms. If eligible, the benefits available for this temporary role may include the following:

  • Medical
  • dental & vision
  • Critical Illness, Accident, and Hospital
  • 401(k) Retirement Plan Pre-tax and Roth post-tax contributions available
  • Life Insurance (Voluntary Life & AD&D for the employee and dependents)
  • Short and long-term disability
  • Health Spending Account (HSA)
  • Transportation benefits
  • Employee Assistance Program
  • Time Off/Leave (PTO, Vacation or Sick Leave)
Workplace Type

This is a fully remote position.

Application Deadline

This position is anticipated to close on Oct 10, 2026.

About TEKsystems

We're partners in transformation. We help clients activate ideas and solutions to take advantage of a new world of opportunity. We are a team of 80,000 strong, working with over 6,000 clients, including 80% of the Fortune 500, across North America, Europe and Asia. As an industry leader in Full-Stack Technology Services, Talent Services, and real-world application, we work with progressive leaders to drive change. That's the power of true partnership. TEKsystems is an Allegis Group company.

The company is an equal opportunity employer and will consider all applications without regards to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.

About TEKsystems and TEKsystems Global Services

Were a leading provider of business and technology services. We accelerate business transformation for our customers. Our expertise in strategy, design, execution and operations unlocks business value through a range of solutions. Were a team of 80,000 strong, working with over 6,000 customers, including 80% of the Fortune 500 across North America, Europe and Asia, who partner with us for our scale, full-stack capabilities and speed. Were strategic thinkers, hands-on collaborators, helping customers capitalize on change and master the momentum of technology. Were building tomorrow by delivering business outcomes and making positive impacts in our global communities. TEKsystems and TEKsystems Global Services are Allegis Group companies. Learn more at TEKsystems.com.

The company is an equal opportunity employer and will consider all applications without regard to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.

San Francisco Fair Chance Ordinance: Pursuant to the San Francisco Fair Chance Ordinance, for all positions located in the city and county of San Francisco, we will consider for employment qualified applicants with arrest and conviction records.

Massachusetts Lie Detector: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Use of Artificial Intelligence (AI): We may use Artificial Intelligence (AI) to support parts of our hiring process, including sourcing, screening, and evaluating candidates. AI helps assess applications and qualifications, but final decisions are made by our hiring team. By applying, you acknowledge and agree that your application may be reviewed using AI tools.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC/AI Cloud Engineer
HPC/AI Cloud Engineer

TEKsystems • United States

Remote
USD 103,000 - 117,000
Medical benefits
401(k) retirement plan
Life Insurance
+3
Support Engineer
Support Engineer

TEKsystems • Atlanta (GA)

On-site
USD 44,000 - 63,000
Medical, dental & vision
401(k) Retirement Plan
Life Insurance
+3
Data Center Operations Technician
Data Center Operations Technician

TEKsystems • City of Rochester (NY)

On-site
USD 33,000 - 44,000
Medical, dental & vision
401(k) Retirement Plan
Life Insurance
+6
Platforms Engineer
Platforms Engineer

TEK Systems • Denver (CO)

On-site
USD 79,000 - 117,000
Medical benefits
401k plan
Life insurance
+5
Forward Deployed Engineer
Forward Deployed Engineer

TEKsystems • United States

Remote
USD 150,000 - 280,000
Medical coverage
HSA contributions
401k match
+6
Data Center Technician
Data Center Technician

TEKsystems • San Jose (CA)

On-site
USD 48,000 - 51,000
Medical, dental & vision
401(k) Retirement Plan
Life Insurance
+2
Software Engineer (Chris Tsukichi)
Software Engineer (Chris Tsukichi)

TEKsystems • Milwaukee (WI)

Hybrid
USD 69,000 - 83,000
Medical, dental & vision
401(k) Retirement Plan
Life Insurance
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

TEKsystems • United States

Hybrid
USD 83,000 - 90,000
Medical, dental & vision
401(k) Retirement Plan
Life Insurance
+2
SDE1
SDE1

TEKsystems • Seattle (WA)

On-site
USD 58,000 - 80,000
Medical, dental & vision
401(k) Retirement Plan
Life Insurance
+5
Senior FinOps Analyst - AI
Senior FinOps Analyst - AI

TEK Systems • Cary (AR)

Remote
USD 96,000 - 110,000
Medical, dental & vision
401(k) plan and other benefits
Paid time off