Staff Software Engineer - Fleet Management

Nscale

United States

On-site

USD 220,000 - 320,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision
Flexible PTO
Parental leave
Retirement plan participation

Job summary

Nscale is seeking a Staff Software Engineer to build Fleet Manager, the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale. You will own domain-level architecture for distributed Python services and drive reliability and observability across the fleet.

The role blends infrastructure automation, workflow orchestration, and hardware lifecycle management, with opportunities to leverage AI tools to accelerate delivery while maintaining

Qualifications

  • Extensive experience designing, building, and operating distributed systems in production.
  • Strong proficiency in Python — Fleet Manager is built entirely in Python.
  • Strong understanding of event-driven and workflow architecture, including reliable delivery, idempotency, retries, replay, and failure handling.
  • Track record of delivering automation systems from ambiguous requirements to production.
  • Ability to lead ambiguous technical work across team boundaries and drive domain-level delivery.

Responsibilities

  • Own domain-level technical direction for a Fleet Manager domain.
  • Design and build production-grade automation in Python.
  • Engineer for reliability, idempotency, resumability, and auditability.
  • Integrate Fleet Manager with DCIM/NetBox and bare-metal provisioning tools.
  • Establish patterns and runbooks for scalable, maintainable software.
  • Own production outcomes with strong observability and incident response.
  • Mentor teammates and promote best practices.
  • Use AI to accelerate delivery while maintaining architectural coherence.

Skills

Python
Distributed systems
Event-driven
Leadership
AI tooling

Tools

NetBox
MAAS
Ironic
PXE
Kubernetes

Job description

Staff Software Engineer - Fleet Management

US

.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the role

Nscale is hiring a Staff Software Engineer to build Fleet Manager — the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale.

This role sits at the intersection of distributed systems, infrastructure automation, and physical hardware. You'll own domain-level architecture within Fleet Manager: Python-based systems that manage the entire operational lifecycle of our compute infrastructure, from initial device enrollment through multi-day burn-in testing to ongoing health monitoring and automated remediation. The problems are challenging and the stakes are high — the software you design and build determines how quickly and how reliably Nscale scales its GPU fleet to meet demand.

This is an opportunity to shape a foundational platform early, setting the patterns and standards that engineers across Fleet Manager build on.

What you'll work on
  • Device provisioning and enrollment: automation that takes bare-metal GPU nodes and network switches from first power-on to production-ready — BMC configuration, DHCP reservations, and provisioning state machines.
  • Burn-in and validation: multi-day testing workflows that qualify hardware before it enters, and re-enters, the fleet.
  • Workflow orchestration: durable, event-driven state machines that span multiple days, survive crashes, resume from checkpoints, support human-in-the-loop approval gates, and let thousands of concurrent idempotent workflows run without stepping on each other.
  • GPU health monitoring and self-healing: detection, diagnosis, and automated remediation workflows that keep nodes healthy in production.
  • Network configuration: switch lifecycle automation and network state management across the fleet.
  • Integrations: keeping Fleet Manager consistent with datacenter inventory tooling (DCIM, NetBox), bare-metal provisioning systems (MAAS, Ironic, IPMI), credential stores, and monitoring infrastructure.
  • Observability: structured logging, metrics, distributed tracing, and tooling that lets operators troubleshoot effectively.
Responsibilities
  • Domain-level technical direction. Own the architecture for a major Fleet Manager domain — such as provisioning, validation, or remediation — influencing engineers across the team and adjacent squads.
  • Design and build production-grade automation. Implement device provisioning, burn-in testing, network configuration, and hardware health validation workflows in Python.
  • Engineer for reliability and auditability. Treat idempotency, resumability, checkpointing, retries, replay, and failure handling as first‑class design concerns.
  • Integrate broadly. Connect Fleet Manager with datacenter infrastructure management systems, cloud orchestration platforms, and bare‑metal provisioning tools.
  • Create leverage through standards. Establish shared patterns, libraries, conventions, and operational runbooks that other engineers build on.
  • Own production outcomes. Operate what you build with strong observability, alerting, incident response, and day‑2 operational discipline.
  • Mentor and influence. Raise the bar through design reviews, implementation guidance, and operational best practices.
  • Use AI to accelerate delivery while maintaining architectural coherence.
Requirements
  • Extensive experience designing, building, and operating distributed systems in production, ideally in infrastructure automation, workflow tooling, or platform engineering.
  • Strong proficiency in Python — Fleet Manager is built entirely in Python.
  • Strong understanding of event-driven and workflow architecture, including reliable delivery, idempotency, retries, replay, and failure handling.
  • Track record of delivering automation systems from ambiguous requirements to production, with hands‑on day‑2 operations experience (monitoring, incident response, performance optimization).
  • Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority.
  • You are driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement.
  • You use AI tools like Claude or Cursor as a core part of your development workflow to create leverage, increase quality, and accelerate delivery.
  • Excellent communication skills to build consensus with stakeholders, both internally and externally, in a fast‑paced, high‑agency environment.
Preferred
  • Experience with workflow orchestration tools like Temporal, Airflow, Prefect, or similar
  • Hands‑on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems
  • Bare‑metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation
  • Experience building hardware lifecycle automation: provisioning, validation, testing, or remediation workflows
  • GPU infrastructure experience: health monitoring, burn‑in testing, or cluster management
  • HPC and networking: datacenter topology, high‑performance interconnects (InfiniBand, RoCE)
  • Deep knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP
  • Open‑source contributions in infrastructure automation or cloud‑native tooling

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$220,000 - $320,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineering Manager - Fleet Management
Senior Software Engineering Manager - Fleet Management

Socket.dev • Seattle (WA)

On-site
USD 300,000 - 350,000
Equity
Bonus potential
Flexible work policy
+1
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Equity
Base salary + equity
Career progression
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • Houston (TX)

On-site
USD 150,000 - 215,000
Base salary + equity
Annual reviews
Growth opportunities
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • New York (NY)

On-site
USD 150,000 - 215,000
Equity
Medical insurance
Dental insurance
+4
Infrastructure Software Engineer, Fleet & Automation Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation Houston; New York; San Francisco; Seattle

Nscale • Northern (KY)

Hybrid
USD 150,000 - 215,000
Base + equity
Equity
Growth opportunities
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Equity
Competitive compensation
Benefits package (medical, dental, V)
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Socket.dev • Houston (TX)

On-site
USD 150,000 - 200,000
Base + equity
Fast-growing startup
Growth/progression plan
Senior Software Engineering Manager - Fleet Management
Senior Software Engineering Manager - Fleet Management

nscaleoperationsukltd • Seattle (WA)

On-site
USD 180,000 - 240,000
Director, Deployment Engineering
Director, Deployment Engineering

Nscale • Houston (TX)

On-site
USD 240,000 - 320,000