Staff AI Platform Engineer — Inference & Scale Leader

Greenhouse Software, Inc.

New York, Northern (NY, KY)

Hybrid

USD 220,000 - 330,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Nscale is seeking a Staff AI Engineer (Specialised) to steer technical direction for parts of the inference platform at the core of our AI cloud. You will influence latency, throughput, and cost of tokens, guiding standards across serving, post-training, and platform teams.

We run state-of-the-art GPU systems and push to close gaps with open-source tooling, mentoring engineers as we go. This role demands deep expertise and hands-on leadership across complex AI workloads.

Qualifications

  • 8–12 years of engineering experience, with significant depth in production AI systems or ML research at scale.
  • 4+ years of hands-on work with LLMs in at least one focus area, in production or research.
  • Deep, demonstrated expertise in at least one focus area.
  • Enough working knowledge of the full serving stack to debug across API, router, scheduler, engine, kernel, and cluster.
  • A track record of setting technical direction and creating standards or tools adopted beyond your own team.
  • Strong Python and PyTorch, with production-grade systems.

Responsibilities

  • Set technical direction for your focus areas and turn it into work that multiple teams can deliver.
  • Lead the resolution of systemic performance and reliability problems across the serving stack, from kernel bottlenecks to fleet-level capacity and multi-tenant isolation.
  • Own the trade-offs between cost, latency, throughput, and model quality in your area, and back them with measurement.
  • Build reusable frameworks, benchmarks, and tooling that make other AI engineers at Nscale more effective.
  • Evaluate emerging serving engines, kernel libraries, RL frameworks, and accelerators, and make clear build/adopt/contribute recommendations.
  • Coach and grow engineers across teams, and raise engineering quality broadly.
  • Work with research, product, and infrastructure leadership so the platform tracks customer demand.
  • Represent your area in cross-team technical reviews and planning.

Skills

Python
PyTorch
LLM Expertise
System Design

Tools

Kubernetes
CUDA
Triton
CUTLASS

Job description

Nscale is seeking a Staff AI Engineer (Specialised) to steer technical direction for parts of the inference platform at the core of our AI cloud. You will influence latency, throughput, and cost of tokens, guiding standards across serving, post-training, and platform teams.

We run state-of-the-art GPU systems and push to close gaps with open-source tooling, mentoring engineers as we go. This role demands deep expertise and hands-on leadership across complex AI workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Platform Engineer – Inference & RL Architect
Staff AI Platform Engineer – Inference & RL Architect

Nscale • Seattle (WA)

On-site
USD 220,000 - 293,000
Staff AI Platform Architect for Scalable AI Services
Staff AI Platform Architect for Scalable AI Services

Nscale • San Francisco (CA)

On-site
USD 220,000 - 293,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Staff AI Platform Engineer — Equity & API Leadership
Staff AI Platform Engineer — Equity & API Leadership

Nscale • New York (NY)

On-site
USD 220,000 - 293,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Staff AI Platform Engineer: Scale ML Infra
Staff AI Platform Engineer: Scale ML Infra

DAT Freight Solutions • Seattle (WA)

Hybrid
USD 198,000 - 246,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
Staff AI Product Engineer
Staff AI Product Engineer

Nscale • Seattle (WA)

On-site
USD 220,000 - 293,000
Staff AI Infra Engineer: Scale GPU AI Platforms
Staff AI Infra Engineer: Scale GPU AI Platforms

Seekr • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Equity Ownership – RSUs
Unlimited PTO + 14 paid holidays
Flexible hybrid work environment
+2
Staff AI Product Engineer
Staff AI Product Engineer

Greenhouse Software, Inc. • New York (NY), Northern (KY)

Hybrid
USD 220,000 - 330,000
Staff Cloud-Native Engineer for AI Infrastructure
Staff Cloud-Native Engineer for AI Infrastructure

Nscale • Houston (TX), Northern (KY)

Hybrid
USD 220,000 - 265,000
Medical benefits
Dental benefits
Flexible PTO
Staff Platform Infra Engineer — Scale AI Workloads
Staff Platform Infra Engineer — Scale AI Workloads

Showcify • United States

Remote
USD 180,000 - 240,000
Lead Architect, Scaled AI Inference
Lead Architect, Scaled AI Inference

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package