Senior AI Platform Architect: Inference, GPUs & Reliability

Accellor

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Accellor is an AI-native services firm delivering measurable outcomes through AI, data, and engineering capabilities. We seek a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize the systems powering advanced AI workloads.

The role focuses on inference runtime, model serving, GPU infrastructure, context engineering, and release safety, requiring a senior hands-on architect able to drive architecture across latency, throughput, and cost.

Qualifications

  • 10–12 years in software engineering, systems architecture, or ML infrastructure.
  • Strong Python and at least one systems/backend language (C++, Go, Rust, Java, TS).
  • Deep understanding of distributed systems, reliability, and scalability.
  • Experience with APIs, microservices, orchestration, and monitoring.
  • Knowledge of AI/ML systems, model serving, context engineering, evaluation.
  • GPU-focused experience: CUDA/Triton, distributed inference, memory optimization.

Responsibilities

  • AI Systems Architecture: design large-scale AI systems for ChatGPT/OpenAI APIs and research workloads.
  • Inference Runtime & Model Serving: build high-throughput low-latency inference across GPU clusters.
  • GPU, Kernel & Distributed Performance: optimize kernels, memory, and interconnects.
  • Context Engineering: define prompt structure, retrieval, and context management.
  • Cost Optimization: reduce token use, caching, and efficient inference paths.
  • Training & Research Infra: support distributed training and experiment velocity.
  • Release Safety & Validation: implement gates for safe, reliable platform releases.
  • Reliability & Observability: telemetry, dashboards, runbooks, SLOs, post-incident learning.

Skills

Python
C++
Distributed systems
GPU
PyTorch
JAX

Tools

CUDA
Triton
NCCL/RCCL
PyTorch
JAX
Kubernetes

Job description

Accellor is an AI-native services firm delivering measurable outcomes through AI, data, and engineering capabilities. We seek a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize the systems powering advanced AI workloads.

The role focuses on inference runtime, model serving, GPU infrastructure, context engineering, and release safety, requiring a senior hands-on architect able to drive architecture across latency, throughput, and cost.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Architect - Inference & Platform
Senior AI Systems Architect - Inference & Platform

Worky • Mountain View (CA)

On-site
USD 260,000 - 320,000
AI Systems Architect — Senior Platform Leader
AI Systems Architect — Senior Platform Leader

Accellor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Platform Architect — GPU & Inference
Senior AI Platform Architect — GPU & Inference

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Cloudflare • Austin (TX)

Hybrid
USD 180,000 - 240,000
AI Systems & Platform Internals - Technical Architect
AI Systems & Platform Internals - Technical Architect

Accellor • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
AI Principal Engineer
AI Principal Engineer

Accellor • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Systems & Platform Internals - Technical Architect
AI Systems & Platform Internals - Technical Architect

Worky • Mountain View (CA)

On-site
USD 260,000 - 320,000
AI Principal Engineer
AI Principal Engineer

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
Senior Architect, Scaled AI Inference & Systems — Equity
Senior Architect, Scaled AI Inference & Systems — Equity

NVIDIA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Senior AI Inference Architect — Disaggregated Serving
Senior AI Inference Architect — Disaggregated Serving

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package