Senior AI Systems Architect - Inference & Platform

Worky

Mountain View (CA)

On-site

USD 260,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Accelloris is an AI-native services firm focused on operationalizing AI at scale, delivering measurable business outcomes through advanced AI, data, and engineering capabilities. A Technical Architect — AI Systems, Inference & Platform Internals is sought to design, scale, and optimize systems powering ChatGPT, OpenAI API, Codex, and agentic workflows.

This senior role emphasizes GPU-level performance, distributed inference, reliability, and cost-efficient operations.

Qualifications

  • 10-12 years of experience in software engineering, systems architecture, ML infrastructure, distributed systems, platform engineering, inference systems, cloud infrastructure, or large-scale backend engineering.
  • Strong hands-on engineering experience with Python and at least one systems/backend language such as C++, Go, Rust, Java, or TypeScript.
  • Deep understanding of distributed systems, production infrastructure, reliability engineering, scalability, observability, and fault-tolerant architecture.
  • Experience designing or operating large-scale systems involving APIs, microservices, distributed compute, orchestration, job scheduling, caching, high-availability infrastructure, and production monitoring.
  • Strong understanding of AI/ML systems, especially model serving, inference workflows, context engineering, retrieval systems, evaluation pipelines, and production model deployment.
  • Practical understanding of GPU systems, accelerator-based workloads, CUDA/Triton-style programming, distributed inference, GPU profiling, memory optimization, and communication libraries such as NCCL or RCCL.
  • Experience with ML frameworks and serving stacks such as PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Apache Ray, Kubernetes-based serving, or internal model-serving systems.
  • Ability to debug complex problems across model behavior, runtime systems, distributed infrastructure, networking, GPU execution, context quality, retrieval quality, evaluation harnesses, and production services.
  • Strong communication skills with the ability to write clear architecture documents, evaluate trade-offs, review implementation quality, and align teams around technically sound decisions.

Responsibilities

  • 1. AI Systems Architecture: Design and evolve large-scale AI systems that support ChatGPT, OpenAI API, Codex, agentic workflows, multimodal models, and research workloads.
  • 2. Inference Runtime & Model Serving: Architect high-throughput, low-latency inference systems across large-scale GPU clusters.
  • 3. GPU, Kernel & Distributed Performance: Analyze and improve performance across GPU kernels, memory movement, collective communication, orchestration, and runtime scheduling.
  • 4. Context Engineering: Design and guide context engineering frameworks that determine what information should be passed to the model, how it should be structured, how much context should be used, and how context quality should be measured.
  • 5. Cost Optimization Frameworks: Design and build cost optimization frameworks for large-scale LLM and GenAI workloads.
  • 6. Training & Research Infrastructure: Collaborate with research and training infrastructure teams to support large-scale model training and post-training workflows.
  • 7. Release Safety, Validation & Evaluation Gates: Architect validation and release systems that ensure model updates, inference engine changes, runtime images, prompt changes, context changes, and platform releases are correct, safe, performant, and regression-free.
  • 8. Reliability, Observability & Production Operations: Design systems that make AI infrastructure observable, debuggable, reliable, and operationally safe.
  • 9. Agentic & Multimodal Platform Internals: Support architecture for AI agents, tool use, memory, function calling, multimodal interaction, long-running workflows, and internal or external agent deployment.
  • 10. Technical Leadership: Work closely with multiple teams; cut across layers, resolve ambiguity, drive architecture decisions, and mentor engineers.

Skills

Python
Distributed systems
GPU infra
Model serving
CI/CD
Kubernetes

Tools

PyTorch
JAX
TensorFlow
Triton
vLLM-style serving
CUDA

Job description

Accelloris is an AI-native services firm focused on operationalizing AI at scale, delivering measurable business outcomes through advanced AI, data, and engineering capabilities. A Technical Architect — AI Systems, Inference & Platform Internals is sought to design, scale, and optimize systems powering ChatGPT, OpenAI API, Codex, and agentic workflows.

This senior role emphasizes GPU-level performance, distributed inference, reliability, and cost-efficient operations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Architect — Senior Platform Leader
AI Systems Architect — Senior Platform Leader

Accellor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Platform Architect — GPU & Inference
Senior AI Platform Architect — GPU & Inference

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
Senior AI Systems Architect - GPU & Platform Internals
Senior AI Systems Architect - GPU & Platform Internals

Accellor • San Francisco (CA)

On-site
USD 210,000 - 320,000
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Systems & Platform Internals - Technical Architect
AI Systems & Platform Internals - Technical Architect

Worky • Mountain View (CA)

On-site
USD 260,000 - 320,000
AI Principal Engineer
AI Principal Engineer

Accellor • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Principal Engineer
AI Principal Engineer

Accellor • San Francisco (CA)

On-site
USD 210,000 - 320,000
AI Principal Engineer
AI Principal Engineer

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
Senior AI Engineer: Real-Time Inference & Agent Systems
Senior AI Engineer: Real-Time Inference & Agent Systems

Arcana Analytics • United States

On-site
USD 120,000 - 160,000
Senior AI Inference Platform Architect
Senior AI Inference Platform Architect

Google • Sunnyvale (CA)

On-site
USD 262,000 - 365,000
Bonus target 25%
Equity
Benefits