Staff AI Product Engineer

Greenhouse Software, Inc.

New York, Northern (NY, KY)

Hybrid

USD 220,000 - 330,000

Full time

40 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Nscale is seeking a Staff AI Engineer (Specialised) to steer technical direction for parts of the inference platform at the core of our AI cloud. You will influence latency, throughput, and cost of tokens, guiding standards across serving, post-training, and platform teams.

We run state-of-the-art GPU systems and push to close gaps with open-source tooling, mentoring engineers as we go. This role demands deep expertise and hands-on leadership across complex AI workloads.

Qualifications

  • 8–12 years of engineering experience, with significant depth in production AI systems or ML research at scale.
  • 4+ years of hands-on work with LLMs in at least one focus area, in production or research.
  • Deep, demonstrated expertise in at least one focus area.
  • Enough working knowledge of the full serving stack to debug across API, router, scheduler, engine, kernel, and cluster.
  • A track record of setting technical direction and creating standards or tools adopted beyond your own team.
  • Strong Python and PyTorch, with production-grade systems.

Responsibilities

  • Set technical direction for your focus areas and turn it into work that multiple teams can deliver.
  • Lead the resolution of systemic performance and reliability problems across the serving stack, from kernel bottlenecks to fleet-level capacity and multi-tenant isolation.
  • Own the trade-offs between cost, latency, throughput, and model quality in your area, and back them with measurement.
  • Build reusable frameworks, benchmarks, and tooling that make other AI engineers at Nscale more effective.
  • Evaluate emerging serving engines, kernel libraries, RL frameworks, and accelerators, and make clear build/adopt/contribute recommendations.
  • Coach and grow engineers across teams, and raise engineering quality broadly.
  • Work with research, product, and infrastructure leadership so the platform tracks customer demand.
  • Represent your area in cross-team technical reviews and planning.

Skills

Python
PyTorch
LLM Expertise
System Design

Tools

Kubernetes
CUDA
Triton
CUTLASS

Job description

Houston; New York; San Francisco; Seattle

About Nscale

Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As a Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.

About the Role

Nscale is looking for aStaff AI Engineer (Specialised)to set technical direction for part of the inference platform at the core of our AI cloud: dedicated and serverless inference, bring-your-own-model deployments, and the post-training services that sit alongside them.

You’ll be the technical authority for one or more focus areas, working across the teams that build serving, post-training, and platform. Your decisions shape the latency, throughput, and cost of the tokens Nscale serves, and whether we can say with evidence that the models we run are the right ones. You’ll take on open questions, such as how KV cache should move between GPUs, nodes, and storage, or where hand-written kernels beat the open-source defaults, and your answers become the standards others build on.

We run state-of-the-art GPU systems, and much of the open-source ecosystem hasn’t caught up with them yet. This role is for people who want to close that gap hands-on while growing the engineers around them.

How We Work
  • Dog years.We move quickly and compress a lot of learning into a short time.
  • Don’t let perfect be the enemy of good.Ship, measure, iterate.
  • Be relentless.Own the problem end to end and see it through.
  • One team, one mission.Outcomes over process, and no "not my job".
Focus Areas

We’re hiring for depth. You don’t need all of these; we want people who are exceptional inone or more, and we’re deliberately hiring people with different focus areas:

  • KV cache offloading and inference performance:KV cache orchestration across GPU, host memory, and storage; cross-instance cache sharing (LMCache or similar); KV-aware routing; disaggregated prefill/decode; speculative decoding; quantisation (FP8, NVFP4, INT8/4); MoE serving
  • GPU kernels and performance:writing and tuning kernels in CUDA, Triton, CUTLASS or ROCm; attention, MoE and GEMM optimisation; multi-GPU and multi-node communication on NVLink-scale systems
  • Evals and benchmarking:designing evaluation frameworks, turning customer requirements into custom benchmarks, and producing objective model comparisons we can stand behind
  • Post-training and RL:fine-tuning and preference optimisation as a service; RL for LLMs (PPO/GRPO-style, reward modelling, multi-turn and tool-use RL); rollout generation, weight synchronisation, and sharing infrastructure between inference and training
  • Serving platform and APIs:serving engines (vLLM, SGLang, TensorRT-LLM), model onboarding, and OpenAI-compatible APIs that other engineers and customers build on
Responsibilities
  • Set technical direction for your focus areas and turn it into work that multiple teams can deliver
  • Lead the resolution of systemic performance and reliability problems across the serving stack, from kernel bottlenecks to fleet-level capacity and multi-tenant isolation
  • Own the trade-offs between cost, latency, throughput, and model quality in your area, and back them with measurement
  • Build reusable frameworks, benchmarks, and tooling that make other AI engineers at Nscale more effective
  • Evaluate emerging serving engines, kernel libraries, RL frameworks, and accelerators, and make clear build/adopt/contribute recommendations
  • Coach and grow engineers across teams, and raise engineering quality broadly
  • Work with research, product, and infrastructure leadership so the platform tracks customer demand
  • Represent your area in cross-team technical reviews and planning
Requirements
  • 8–12 years of engineering experience, with significant depth in production AI systems or ML research at scale
  • 4+ years of hands-on work with LLMs in at least one of the focus areas above, in production or research
  • Deep, demonstrated expertise in at least one of the focus areas above
  • Enough working knowledge of the full serving stack to debug across it: API, router, scheduler, engine, kernel, and cluster
  • A track record of setting technical direction and creating standards or tools adopted beyond your own team
  • Strong Python and PyTorch, with a track record of maintainable, well-tested, production-grade systems
  • Experience operating AI workloads in containerised, distributed environments (Kubernetes, large GPU clusters)
  • Solid understanding of GPU performance: memory bandwidth, interconnect, and parallelism strategies (data, tensor, pipeline, expert)
  • Deep understanding of transformer and LLM architectures and how they behave under production load
Preferred
  • Contributions to widely used open-source inference, kernel, or RL projects (vLLM, SGLang, TensorRT-LLM, LMCache, FlashInfer, Triton, verl, OpenRLHF, TRL, DeepSpeed, etc.)
  • Experience at an AI lab, hyperscaler AI team, or leading ML infrastructure company
  • Published research or technical writing on inference, kernels, evals, or RL (MLSys, NeurIPS, ICLR, blog posts)
  • Hands-on custom kernel work (CUDA, Triton, CUTLASS) if kernels aren’t already your focus area
  • Experience designing developer-facing APIs and SDKs used by external customers
  • Experience with control plane / data plane separation and cell-based architecture for scale-out and blast-radius isolation

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$220,000 - $330,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Product Engineer
Staff AI Product Engineer

Nscale • Seattle (WA)

On-site
USD 220,000 - 293,000
Staff Cloud Native Software Engineer New Houston; San Francisco; Seattle
Staff Cloud Native Software Engineer New Houston; San Francisco; Seattle

Nscale • Houston (TX), Northern (KY)

On-site
USD 220,000 - 265,000
Medical benefits
Dental benefits
Flexible PTO
Staff AI Product Engineer
Staff AI Product Engineer

Nscale • San Francisco (CA)

On-site
USD 220,000 - 293,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Staff AI Product Engineer
Staff AI Product Engineer

Nscale • New York (NY)

On-site
USD 220,000 - 293,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Senior Network Engineer
Senior Network Engineer

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Staff Cloud Native Software Engineer
Staff Cloud Native Software Engineer

Nscale • Seattle (WA)

On-site
USD 220,000 - 265,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Senior Infrastructure Support Engineer
Senior Infrastructure Support Engineer

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Senior Engineer, Storage Services
Senior Engineer, Storage Services

Nscale • United States

On-site
USD 150,000 - 300,000
Flexible paid time off
Parental leave
Retirement plan participation
Director, Competitive Insights
Director, Competitive Insights

Nscale • Houston (TX)

On-site
USD 220,000 - 380,000
Bonus program
Equity
Retirement plan
Senior Manager, Competitive Insights
Senior Manager, Competitive Insights

Nscale • San Francisco (CA)

On-site
USD 150,000 - 260,000
Medical insurance
Dental insurance
Vision insurance
+3