Get more replies from employers
Send a job-specific resume in minutes.
Nscale, a Seattle-based GPU cloud for AI, is hiring a Software Engineering Manager to lead the Fleet Manager team that provisions, tests, and remediates GPU nodes and network switches at scale. You will own delivery, people leadership, and end-to-end lifecycle of our compute infrastructure.
You’ll stay hands-on with design reviews, incident deep-dives, and coaching, partnering with Principal and Staff engineers. Strong background in Python & distributed systems is essential.
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
We're hiring a Software Engineering Manager to lead the team building Fleet Manager - the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale. Reporting to the Director of Software Engineering, Fleet Management based in EMEA, you'll lead the engineers who build the Python-based systems managing the entire operational lifecycle of our compute infrastructure: device enrollment, burn-in testing, network configuration, GPU health monitoring, and automated remediation. You'll own delivery and team health end to end - planning, execution, hiring, and career development - while partnering closely with Principal and Staff engineers who drive technical architecture. We're a hands‑on engineering culture: you'll stay close to the systems through design reviews, incident deep‑dives, and code, with people leadership as your primary craft.
8+ years of software engineering experience building and operating production systems, including 2+ years directly managing software engineers . Strong technical foundation in Python and distributed systems : credible in design and code reviews, and able to guide trade‑offs in infrastructure automation or workflow tooling. Track record of delivering complex projects from ambiguous requirements to production, with hands‑on day‑2 operations experience (monitoring, incident response, performance optimization).
Proven people leadership: hiring, coaching, performance management, and developing engineers toward senior and staff levels. You are driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement. You use AI tools like Claude or Cursor as a core part of your workflow and know how to raise a whole team's leverage with them. You stay effective while context-switching between technical depth, delivery judgment calls, and people leadership - reviewing a design, unblocking an engineer, and running a hiring debrief in the same morning. Excellent communication skills to build consensus with stakeholders, both internally and externally.