ML Infra Engineer: Scale GPU Clusters | Equity

Awake Solutions

Mexico

Hybrid

MXN 23,000 - 37,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Equity
GPU credits
Research budget
Conference support
Home office stipend

Job summary

Awake Solutions is seeking an ML Infrastructure Engineer to design and operate scalable infrastructure for training and serving large ML workloads with a focus on cost-efficiency, reliability and reproducibility.

You will architect and maintain GPU/TPU clusters, implement storage and feature stores, and automate capacity planning and scheduling for training workloads. Collaboration on security and multi-tenancy will be essential to support research infra.

Qualifications

  • 4+ years in infra engineering with experience in cloud GPU/accelerator environments.
  • Experience with cluster orchestration, provisioning and performance tuning.
  • Familiarity with storage systems, feature stores and large-data workflows.
  • Strong scripting and automation skills (Python, Bash, Terraform).

Responsibilities

  • Architect and maintain scalable GPU/TPU clusters and provisioning workflows.
  • Implement data storage, versioning and feature store infrastructure.
  • Automate capacity planning, cost controls and job scheduling for training workloads.
  • Collaborate on security, multi-tenancy and access patterns for research infra.

Skills

Python
Bash
Terraform
GPU/TPU infrastructure
Cluster orchestration

Tools

Kubernetes

Job description

Awake Solutions is seeking an ML Infrastructure Engineer to design and operate scalable infrastructure for training and serving large ML workloads with a focus on cost-efficiency, reliability and reproducibility.

You will architect and maintain GPU/TPU clusters, implement storage and feature stores, and automate capacity planning and scheduling for training workloads. Collaboration on security and multi-tenancy will be essential to support research infra.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer
ML Infrastructure Engineer

Awake Solutions • Mexico

Hybrid
MXN 23,000 - 37,000
Competitive salary
Equity
GPU credits
+3
Remote Distributed Training Engineer — Multi-GPU Scaling
Remote Distributed Training Engineer — Multi-GPU Scaling

BairesDev • Mexico

On-site
PHP 7,299,000 - 11,557,000
100% remote work
Excellent USD compensation
Hardware and software setup
+4
Senior ML Engineer: Build Scalable AI & MLOps
Senior ML Engineer: Build Scalable AI & MLOps

Indra Group • Metro Manila

On-site
PHP 900,000 - 1,600,000
On-site ML Engineer: Production AI Solutions
On-site ML Engineer: Production AI Solutions

Ubiquity Global Services, Inc. • Taguig

On-site
PHP 800,000 - 1,200,000
Career development programs
People-first culture
ML Platform Engineer
ML Platform Engineer

Alloy Enterprises • Mexico

Hybrid
PHP 7,299,000 - 9,124,000
Machine Learning Specialist {LLM}
Machine Learning Specialist {LLM}

NCS Group • Taguig

On-site
ML Engineer: Lifecycle, Deployment & Responsible AI
ML Engineer: Lifecycle, Deployment & Responsible AI

Smart Communications, Inc. • Philippines

On-site
PHP 1,200,000 - 1,600,000
Machine Learning Specialist
Machine Learning Specialist

Indra Group • Metro Manila

On-site
PHP 900,000 - 1,600,000
Remote Deep Learning Engineer – Distributed Training
Remote Deep Learning Engineer – Distributed Training

BairesDev • Mexico

On-site
PHP 7,299,000 - 10,949,000
Remote work 100%
Competitive compensation (USD or local
Home office setup provided
+3
Staff Engineer, Scalable Inference Platform
Staff Engineer, Scalable Inference Platform

Cerebras Systems • Binangonan

On-site
PHP 11,077,000 - 14,769,000