Senior AI Systems Performance Engineer

Goaly

Menlo Park, Northern (CA, KY)

Hybrid

USD 150,000 - 180,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Goaly seeks a systems/performance engineer to optimize AI workloads at scale. You will identify bottlenecks across model code, GPU kernels, memory, and orchestration, building high-throughput, fault-tolerant systems for training and inference.

Collaborate with researchers and other engineers to translate insights into reusable platform capabilities and durable improvements across a production-grade ML stack.

Qualifications

  • Significant software-engineering and ML-infrastructure experience at scale.

Responsibilities

  • Profile end-to-end AI workloads and identify bottlenecks across model code, GPUs, memory, and networking.

Skills

Distributed systems
Python
Systems programming
Performance engineering
Debugging
Collaboration with researchers
Owner mindset
Communication

Tools

CUDA
Triton
Megatron
Kubernetes

Job description

Goaly seeks a systems/performance engineer to optimize AI workloads at scale. You will identify bottlenecks across model code, GPU kernels, memory, and orchestration, building high-throughput, fault-tolerant systems for training and inference.

Collaborate with researchers and other engineers to translate insights into reusable platform capabilities and durable improvements across a production-grade ML stack.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Staff Engineer, RL Systems & ML Infrastructure
Staff Engineer, RL Systems & ML Infrastructure

Goaly • Menlo Park (CA)

Hybrid
USD 180,000 - 240,000
Meals and office benefits
Visa sponsorship
Location-based hybrid policy
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
AI Training Performance Engineer — Large-Scale GPU Training
AI Training Performance Engineer — Large-Scale GPU Training

Figure • San Jose (CA)

On-site
USD 200,000 - 400,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Performance Optimization Engineer for AI Infrastructure
Performance Optimization Engineer for AI Infrastructure

Fireworks AI • San Mateo (CA)

On-site
USD 180,000 - 260,000
AI Inference Performance & Scale Engineer
AI Inference Performance & Scale Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 210,000
Benefits at a glance
Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

NVIDIA • Austin (TX)

On-site
USD 272,000 - 432,000
Equity
Benefits
AI Performance Engineer: Scale ML Workloads at Datacenter
AI Performance Engineer: Scale ML Workloads at Datacenter

Applied Intuition • Sunnyvale (CA)

On-site
USD 215,000 - 285,000