Staff AI Product Engineer

Nscale

Seattle (WA)

On-site

USD 220,000 - 293,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Nscale seeks a Staff AI Engineer (Specialized) to drive the architecture of its inference and reinforcement learning systems at the core of the AI services platform. You will own how models are served on Nscale’s GPU cloud, and how RL workloads run, exposing clean, high-performance interfaces to customers and internal teams.

You will define standards for model efficiency, API boundaries, and platform tooling, collaborating with multiple teams to raise engineering quality and ensure scalable,

Qualifications

  • 8–12 years of engineering experience in production AI systems at scale.
  • Ability to set technical direction for an AI systems domain.
  • Deep expertise in production LLM inference and RL for LLMs.
  • APIs/SDKs designed for engineers and external customers.

Responsibilities

  • Set technical direction for inference serving architecture: routing, scheduling and batching.
  • Lead RL and post-training systems: RLHF, reward modelling, and agentic RL.
  • Define shared infrastructure for RL loops and data curation workflows.
  • Design developer-facing APIs and SDKs that are clean and versioned.
  • Coach engineers and align strategy with business direction.

Skills

Production AI systems
LLM inference
RL for LLMs
APIs & SDKs
Python & PyTorch
Transformer architectures
Cross-team influence
GPU/accelerator workloads

Tools

CUDA
Kubernetes
Triton
OpenAPI 3.x

Job description

About Nscale

Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As a Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.

About Nscale

Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As a Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.

About the Role

Nscale is looking for a Staff AI Engineer (Specialized) to set technical direction for the inference and reinforcement learning systems at the core of our AI services platform - and for the APIs through which other engineers consume them. This role owns the architecture of how models are served on Nscale’s GPU cloud, how RL and post-training workloads run on it, and how both are exposed to customers and internal teams as clean, reliable, high-performance interfaces. You’ll work across 2-4 teams spanning serving, post-training, and platform, defining how these systems are built and establishing the standards that create engineering leverage across the organization. As a Staff engineer, you are the technical authority for this area of the AI stack. Your decisions determine the latency, throughput, and cost profile of every token Nscale serves, and the correctness and efficiency of every RL run on our platform. You resolve ambiguous architectural questions where the answer space is genuinely open - disaggregated versus co-located serving, on-policy versus off-policy RL infrastructure, where the API boundary should sit - and your solutions become the standards others build on.

Responsibilities
Inference
  • Set technical direction for Nscale’s inference serving architecture: request routing, scheduling, continuous batching, KV cache management, prefix caching, and speculative decoding
  • Drive the strategy for model efficiency in production - quantization (FP8, INT8/4), sparsity, pruning, distillation, and MoE serving - and the trade-offs between cost, latency, throughput, and model quality
  • Lead the resolution of systemic performance and reliability challenges across the serving stack, from kernel-level bottlenecks to fleet-level capacity and multi-tenant isolation
Reinforcement learning & post-training
  • Own the architecture of Nscale’s RL and post-training systems: RLHF, DPO/GRPO-style methods, reward modelling, and agentic RL with tool calling, off-policy training, and decoupled sampling and policy updates
  • Define how inference and training share infrastructure in RL loops - rollout generation, sample buffering, weight synchronization, and the serving engine’s role inside the training system
  • Establish standards for fine-tuning services (LoRA, QLoRA, adapters, full fine-tuning) and the data curation and processing workflows that feed them
Developer APIs & platform
  • Set the design standards for Nscale’s developer-facing APIs, SDKs, and tooling - OpenAI-compatible and native interfaces, OpenAPI 3.x specifications, versioning, and rate limiting - so that inference and RL capabilities are consumable by engineers who never see the underlying systems
  • Create reusable frameworks and tooling that multiply the effectiveness of other AI engineers across Nscale
Leadership & direction
  • Coach and grow more junior engineers across teams; raise technical capability and engineering quality broadly
  • Collaborate with research, product, and infrastructure leadership to align inference and RL platform strategy with customer demand and business direction
  • Evaluate emerging serving engines, RL frameworks, and accelerator technologies; make decisive build/adopt/contribute recommendations
  • Represent the inference and RL domain in cross-organizational architecture reviews and strategic planning
Requirements
  • 8–12 years of engineering experience, with significant depth in production AI systems at scale (AI labs, hyperscalers, or leading ML infrastructure companies)
  • Demonstrated ability to set technical direction for an AI systems domain at scale
  • Deep expertise in production LLM inference: serving architectures, KV cache and memory management, batching and scheduling strategies, speculative decoding, and low-precision inference
  • Strong hands-on expertise in RL for LLMs - RLHF, DPO and related preference-optimization methods, reward modelling, or agentic/multi-turn RL - including the systems that make them run efficiently on GPU clusters
  • Proven ability to design developer-facing APIs and SDKs that are clean, versioned, and adopted by other engineers and external customers
  • Strong cross-team influence: track record of creating standards and practices adopted by multiple teams
  • Experience designing AI infrastructure with clear control plane / data plane separation and cell-based architecture patterns for scale-out and blast-radius isolation
  • Experience with large-scale GPU/accelerator workloads: CUDA or ROCm, memory bandwidth optimization, and distributed compute paradigms (data/model/tensor parallelism, sharding, scheduling)
  • Strong proficiency in Python and PyTorch, with a track record of building maintainable, well-tested, production-grade ML systems
  • Deep understanding of transformer and LLM architectures and their behavior under production load
  • Ability to design architecture that is both technically excellent and practically adoptable across a diverse engineering organization
Preferred
  • Contributions to widely-used open-source inference or RL frameworks (vLLM, SGLang, TensorRT-LLM, verl, OpenRLHF, TRL, DeepSpeed, etc.)
  • Experience at an AI lab (OpenAI, DeepMind, Anthropic, Meta AI, etc.), hyperscaler AI team, or leading ML infrastructure company
  • Published research or technical writing on inference systems or RL infrastructure (MLSys, NeurIPS, ICLR systems tracks, blog posts)
  • Experience with hardware-software co-design for AI accelerators (custom CUDA kernels, Triton, etc.)
  • Experience operating inference platforms in containerized, distributed environments (Kubernetes, large-scale clusters)
  • Experience building evaluation and benchmarking systems for model quality, safety, and system performance

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range: $220,000 USD - $293,333 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Product Engineer
Staff AI Product Engineer

Greenhouse Software, Inc. • New York (NY), Northern (KY)

Hybrid
USD 220,000 - 330,000
Staff AI Product Engineer
Staff AI Product Engineer

Nscale • San Francisco (CA)

On-site
USD 220,000 - 293,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Staff AI Product Engineer
Staff AI Product Engineer

Nscale • New York (NY)

On-site
USD 220,000 - 293,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Senior Network Engineer
Senior Network Engineer

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Staff Cloud Native Software Engineer New Houston; San Francisco; Seattle
Staff Cloud Native Software Engineer New Houston; San Francisco; Seattle

Nscale • Houston (TX), Northern (KY)

On-site
USD 220,000 - 265,000
Medical benefits
Dental benefits
Flexible PTO
Senior Engineer, Storage Services
Senior Engineer, Storage Services

Nscale • United States

On-site
USD 150,000 - 300,000
Flexible paid time off
Parental leave
Retirement plan participation
Staff Cloud Native Software Engineer
Staff Cloud Native Software Engineer

Nscale • Seattle (WA)

On-site
USD 220,000 - 265,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
SVP, Global Data Center Operations
SVP, Global Data Center Operations

Nscale • Seattle (WA)

On-site
USD 250,000 - 350,000
Equity
Flexible PTO
Medical insurance
+4
Director, Supply Chain Engineering
Director, Supply Chain Engineering

Nscale • New York (NY)

On-site
USD 160,000 - 270,000
Medical insurance
Dental insurance
Vision insurance
+3
Staff AI Platform Engineer – Inference & RL Architect
Staff AI Platform Engineer – Inference & RL Architect

Nscale • Seattle (WA)

On-site
USD 220,000 - 293,000