Pioneering AI Infrastructure Engineer

Goaly

Palo Alto (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Goaly is building a world where every company can become its own AI lab. We are seeking an exceptional systems or performance engineer to optimize large-scale AI workloads, spanning post-training, model training, rollout inference, and GPU cluster platforms.

You will work with researchers and systems engineers to identify bottlenecks, develop scalable infrastructure, and translate investigations into durable improvements across distributed workloads.

Qualifications

  • Strong programming skills in Python and at least one systems language (C++, Rust, or Go).
  • Experience solving large-scale performance or systems problems.

Responsibilities

  • Profile AI workloads and identify bottlenecks across GPUs, memory and networking.
  • Build low-latency inference systems and optimize GPU execution.
  • Improve distributed training performance across heterogeneous hardware.
  • Design quantitative performance models and scheduling mechanisms.
  • Develop benchmarks and tooling to monitor system behavior.

Skills

Python
C++
Distributed systems
Performance engineering
Go
Rust

Tools

CUDA
Triton
Kubernetes

Job description

Goaly is building a world where every company can become its own AI lab. We are seeking an exceptional systems or performance engineer to optimize large-scale AI workloads, spanning post-training, model training, rollout inference, and GPU cluster platforms.

You will work with researchers and systems engineers to identify bottlenecks, develop scalable infrastructure, and translate investigations into durable improvements across distributed workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 180,000
Founding Backend Engineer, AI Agent Platform
Founding Backend Engineer, AI Agent Platform

RiseMe • Palo Alto (CA)

Hybrid
USD 180,000 - 240,000
Hybrid in Palo Alto
Visa sponsorship available
Meals & perks: lunch, dinner, snacks,
Staff Engineer, Cluster Infrastructure & AI Compute
Staff Engineer, Cluster Infrastructure & AI Compute

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Meals and office benefits
Visa sponsorship
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
GPU Infrastructure Engineer — Scalable AI Training
GPU Infrastructure Engineer — Scalable AI Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Senior AI Infrastructure Lead - GPU Clusters & Model Serving
Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Founding Senior AI Infrastructure Engineer
Founding Senior AI Infrastructure Engineer

Goaly • Palo Alto (CA)

On-site
USD 140,000 - 210,000
AI Engineer: Model Training, Inference & GPU Infra
AI Engineer: Model Training, Inference & GPU Infra

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 170,000 - 210,000