Senior AI Systems Architect - GPU & Platform Internals

Accellor

San Francisco (CA)

On-site

USD 210,000 - 320,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Accellor seeks a Senior Technical Architect to lead AI systems, inference runtimes, and platform internals. You will shape architecture across GPUs, distribution, and safety while balancing performance and cost.

The role requires hands-on leadership, deep expertise in AI infrastructure, and collaboration with research, product, and deployment teams to deliver scalable, reliable AI platforms.

Qualifications

  • 10-12 years of experience in software engineering, systems architecture, ML infrastructure, distributed systems, platform engineering, inference systems, cloud infrastructure, or large-scale backend engineering.
  • Strong hands-on engineering experience with Python and at least one systems/backend language such as C++, Go, Rust, Java, or TypeScript.
  • Deep understanding of distributed systems, production infrastructure, reliability engineering, scalability, observability, and fault-tolerant architecture.
  • Experience designing or operating large-scale systems involving APIs, microservices, distributed compute, orchestration, job scheduling, caching, high-availability infrastructure, and production monitoring.
  • Strong understanding of AI/ML systems, especially model serving, inference workflows, context engineering, retrieval systems, evaluation pipelines, and production model deployment.

Responsibilities

  • Design and evolve large-scale AI systems that support ChatGPT, OpenAI API, Codex, agentic workflows, multimodal models, and research workloads.
  • Define architecture across inference runtime, model serving, request routing, batching, KV-cache handling, GPU scheduling, distributed execution, observability, release gates, and production rollout.
  • Own technical trade-offs across latency, throughput, reliability, correctness, safety, scalability, cost, and infrastructure efficiency.
  • Architect high-throughput, low-latency inference systems across large-scale GPU clusters.
  • Work across inference engines, serving layers, scheduling systems, caching, streaming, deployment pipelines, and runtime optimization.
  • Partner with engineering teams to improve model-serving efficiency, tail latency, GPU utilization, memory efficiency, correctness under load, and cost per request.

Skills

Python
Distributed systems
Architecture design
Communication
Team leadership

Tools

PyTorch
JAX
TensorFlow
Triton
vLLM-style serving
Kubernetes
CUDA
NCCL/RCCL

Job description

Accellor seeks a Senior Technical Architect to lead AI systems, inference runtimes, and platform internals. You will shape architecture across GPUs, distribution, and safety while balancing performance and cost.

The role requires hands-on leadership, deep expertise in AI infrastructure, and collaboration with research, product, and deployment teams to deliver scalable, reliable AI platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Architect — Senior Platform Leader
AI Systems Architect — Senior Platform Leader

Accellor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Platform Architect — GPU & Inference
Senior AI Platform Architect — GPU & Inference

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
Senior AI Systems Architect - Inference & Platform
Senior AI Systems Architect - Inference & Platform

Worky • Mountain View (CA)

On-site
USD 260,000 - 320,000
Lead GPU Architect for AI Accelerators & Clusters
Lead GPU Architect for AI Accelerators & Clusters

EngineersOfAI • Milpitas (CA)

On-site
USD 140,000 - 190,000
Senior Infrastructure Architect - AI-Driven Platforms
Senior Infrastructure Architect - AI-Driven Platforms

Socket.dev • Santa Clara (CA)

On-site
USD 224,000 - 357,000
AI Principal Engineer
AI Principal Engineer

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
AI Principal Engineer
AI Principal Engineer

Accellor • San Francisco (CA)

On-site
USD 210,000 - 320,000
Senior AI Developer Tech Engineer - GPU & HPC Optimizer
Senior AI Developer Tech Engineer - GPU & HPC Optimizer

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer - ML Accelerators
Senior AI Infrastructure Engineer - ML Accelerators

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
High-value equity
Senior AI Infrastructure Architect — Enterprise GPU Clusters
Senior AI Infrastructure Architect — Enterprise GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits