Senior AI Platform Architect — GPU & Inference

Accellor

Mountain View (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI systems powering ChatGPT and OpenAI API workloads. The role focuses on inference runtime, model serving, GPU infrastructure, and distributed systems for production reliability.

The ideal candidate has 10–12 years of software/ML infra experience, deep knowledge of PyTorch/JAX, and hands-on GPU optimization.

Qualifications

  • 10–12 years of software/ML infra experience with strong architecture skills.
  • Proven ability to design and operate large AI systems and inference stacks.
  • Hands-on expertise across GPUs, model serving, and distributed runtimes.

Responsibilities

  • Design and evolve large‑scale AI systems supportingChatGPT/OpenAI API workloads.
  • Own latency, throughput, safety, cost, and reliability across platforms.
  • Lead inference runtime, model serving, and GPU scheduling at scale.
  • Collaborate with research to enable frontier model workloads.

Skills

Python
Distributed systems
GPU CUDA
Model serving
Kubernetes
LLM inference
Context engineering
System architecture

Tools

PyTorch
JAX
TensorFlow
Triton

Job description

Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI systems powering ChatGPT and OpenAI API workloads. The role focuses on inference runtime, model serving, GPU infrastructure, and distributed systems for production reliability.

The ideal candidate has 10–12 years of software/ML infra experience, deep knowledge of PyTorch/JAX, and hands-on GPU optimization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Architect — Senior Platform Leader
AI Systems Architect — Senior Platform Leader

Accellor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Systems Architect - Inference & Platform
Senior AI Systems Architect - Inference & Platform

Worky • Mountain View (CA)

On-site
USD 260,000 - 320,000
Senior AI Systems Architect - GPU & Platform Internals
Senior AI Systems Architect - GPU & Platform Internals

Accellor • San Francisco (CA)

On-site
USD 210,000 - 320,000
AI Systems & Platform Internals - Technical Architect
AI Systems & Platform Internals - Technical Architect

Worky • Mountain View (CA)

On-site
USD 260,000 - 320,000
GPU Infra Engineer for Scalable AI Compute
GPU Infra Engineer for Scalable AI Compute

OpenAI • United States

On-site
USD 180,000 - 240,000
AI Principal Engineer
AI Principal Engineer

Accellor • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Principal Engineer
AI Principal Engineer

Accellor • San Francisco (CA)

On-site
USD 210,000 - 320,000
AI Principal Engineer
AI Principal Engineer

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Inference Engineer — GPU & Edge Optimization
Senior AI Inference Engineer — GPU & Edge Optimization

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000