Get more replies from employers
Send a job-specific resume in minutes.
Nscale is seeking a Director of HPC Systems Engineering to drive strategy, architecture, and execution for a major HPC platform. You will lead a manager-led organization delivering reliable, scalable HPC services that align with business priorities and platform strategy.
You will shape the org, define priorities, and ensure continuity of engineering excellence while managing budgets and cross-functional partnerships across engineering, product, and infrastructure teams.
The Director, HPC Systems Engineering is accountable for the strategy, execution, and organisational effectiveness of Nscale's HPC systems engineering function. This role operates at organisation scope and is responsible for ensuring that teams deliver reliable, scalable, and supportable HPC platform capabilities aligned with business priorities and long-term platform strategy.
This role combines organisational leadership with strong technical and operational judgement. You will shape the strategy for a major HPC platform area, build and lead the teams responsible for delivering it, and ensure that architecture, execution, and service quality remain aligned as the organisation grows.
As a Director, your impact comes through leadership of managers, senior engineers, and team‑of‑teams structures. You are accountable for building a healthy, high‑performing organisation with the right structure, talent, operating model, and priorities to package and operate high‑quality HPC systems as services.
Define and evolve the strategy for a major HPC systems engineering area, aligned with business goals, customer needs, platform priorities, and long‑term service direction
Translate company and engineering strategy into clear priorities, investment areas, and execution plans for the organisation
Ensure HPC platform architecture, operational maturity, and delivery priorities remain aligned across teams
Partner with Principal and Senior Staff engineers to ensure technical strategy is coherent, actionable, and executed effectively
Balance short‑term delivery pressure against long‑term platform health, team sustainability, and organisational capability
Lead a major HPC systems engineering organisation through managers, senior technical leaders, and team‑of‑teams structures
Own delivery outcomes, execution quality, service quality, and team health across the area
Shape team topology, ownership boundaries, support models, and operating mechanisms to support effective execution at scale
Establish the right planning, prioritisation, review, and escalation mechanisms to keep the organisation effective and accountable
Remove systemic blockers that reduce engineering effectiveness or create avoidable delivery or operational risk
Maintain sufficient technical depth to challenge designs, evaluate trade-offs, and support strong engineering judgement across the organisation
Ensure teams are building on sound software, systems, and operational foundations for Slurm‑based and adjacent HPC services
Sponsor major cross‑team HPC platform initiatives that require coordination, investment, or organisational change
Ensure reliability, performance, supportability, and operational maturity are treated as first‑class requirements
Build a strong partnership model with adjacent cloud‑native software engineering teams so shared patterns are reused effectively
Hire, develop, and retain strong engineering leaders, managers, and senior HPC systems engineers
Raise the leadership bar through coaching, feedback, succession planning, and strong performance management
Build a healthy partnership model between engineering managers and senior ICs
Strengthen organisational capability over time by addressing skill gaps, clarifying expectations, and improving how the organisation operates
Foster a culture of ownership, clarity, urgency, strong engineering standards, and calm operational leadership
Serve as a trusted partner to senior engineering, product, infrastructure, security, and business stakeholders
Represent the organisation in strategic planning, executive discussions, architecture reviews, customer conversations, and partner discussions where needed
Own budget, headcount planning, and organisational investment decisions for the area
Ensure the HPC systems engineering organisation is positioned to support broader company goals and changing business needs
Where relevant, support Nscale's technical position in the broader HPC and cloud‑native ecosystem
Proven experience leading a significant HPC systems, platform engineering, or infrastructure‑heavy engineering organisation
Strong technical credibility in software engineering for HPC systems, distributed infrastructure, Linux‑based platforms, and production operations
Demonstrated ability to operate at org scope, translating business and engineering strategy into team structure, priorities, and execution
Strong track record of leading through managers and senior technical leaders rather than relying primarily on individual contribution
Experience building high‑performing engineering organisations, including hiring, leadership development, organisation design, and per