Infrastructure Software Engineer, Fleet & Automation

nscaleoperationsukltd

Houston (TX)

On-site

USD 120,000 - 160,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Base + equity

Job summary

Nscale is seeking an Infrastructure Software Engineer for Fleet & Automation to join the AI Infrastructure Operations team. You will build and maintain the control plane, tooling, and automation that support GPU HPC environments and AI workloads.

You’ll own end-to-end workflows for device provisioning, validation, and remediation at scale, design workflow orchestration for hardware lifecycles, and collaborate across product, design, and SRE to drive reliable, observable, and cost-efficient

Qualifications

  • Bachelor's degree in CS, CE or related field or equivalent practical experience.
  • 5+ years building large-scale infrastructure applications.
  • Proficiency in C, C++, Java, and Python for API design and testing.
  • Deep Linux knowledge, networking basics (TCP/IP, BGP), and config management tools like Ansible or Terraform.
  • Experience with distributed systems, stateful/stateless services, and infrastructure tooling (DCIMs, NetBox, OpenStack, MAAS/Ironic/IPMI).

Responsibilities

  • Architect and implement workflow automation systems with scalability in mind.
  • Own end-to-end device provisioning, validation, testing, and remediation at scale.
  • Design workflow orchestration for hardware lifecycle, GPUs and switches.
  • Collaborate with Infra, Platform, and SRE to translate ops needs into automation.
  • Establish reliability and observability standards across services.
  • Build production-grade Python systems for hardware automation and explore AI-driven improvements.
  • Collaborate with cross-functional teams to ensure maintainable automated systems.

Skills

C
C++
Java
Python

Education

Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience

Tools

Ansible
Terraform
NetBox
OpenStack
MAAS
Ironic
IPMI

Job description

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

Overview

As an Infrastructure Software Engineer for Fleet & Automation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering principles, you will focus on building and maintaining the control plane, tooling, and automation that supports Fleet Operations, Network Operations, and Observability functions. Your work will directly translate into higher system availability and reduced operational costs.

Key Responsibilities
  • Perform technical architecture, roadmap and implementation for workflow automation systems, driving architecture decisions that balance automation complexity, reliability, and maintainability. Identify and resolve performance and scalability issues. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership.
  • Own end-to-end delivery of device provisioning, validation, testing, and remediation workflows at scale.
  • Design and build workflow orchestration systems for hardware lifecycle management, including GPU nodes and network switches.
  • Partner with Infrastructure, Platform, and SRE teams to translate operational needs into robust, scalable automation.
  • Establish engineering standards for reliability, observability, and operational excellence across all services. Help set up engineering best practices in collaboration with the broader engineering team.
  • Build production‑grade Python systems for hardware lifecycle automation, leveraging AI tools to accelerate delivery. Assess impact to team software stack from new hardware product programs and explore AI driven process improvement and automation.
  • Collaborate with cross‑functional teams (product, design, operations, infrastructure) to build efficient, interoperable, and maintainable automated systems.
Required Qualifications
  • Education: Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • Experience: 5+ years relevant experience building large‑scale infrastructure applications or similar experience.
  • Programming: Experience in utilizing languages such as C, C++, Java, and scripting languages such as Python for API design and unit testing techniques.
  • Systems Expertise: Deep understanding of Linux operating systems, networking fundamentals (TCP/IP, BGP), and familiarity with configuration management tools (e.g., Ansible, Terraform).
  • Distributed Systems: Experience building, running and debugging large‑scale infrastructure, stateful and stateless services for distributed systems or networks, and experience with compute technologies, storage, or hardware architecture. Experience integrating with infrastructure tooling such as: DCIMs, NetBox, OpenStack, bare metal APIs (MAAS, Ironic, IPMI).
Preferred Qualifications
  • Master's degree or PhD in Engineering, Computer Science, or a related technical field.
  • Experience designing, analyzing and improving efficiency, scalability, and performance of various system resources.
  • Direct experience with AI/HPC infrastructure , including NVIDIA GPUs, InfiniBand or high‑speed Ethernet fabrics, and related management software (e.g., NCCL, SLURM).
  • Experience with advanced observability and monitoring systems (Prometheus, Grafana, OpenTelemetry) for complex, high‑cardinality telemetry data.
  • Familiarity with cloud‑native technologies (Kubernetes, Docker) and infrastructure‑as‑code principles.
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
  • Familiarity with SLOs/metrics measurement, logs/telemetry/metrics integration with tools for enhanced operator experience.
What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core. Highly competitive package (base + equity) with reviews every 12 months. Join the fastest-growi

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Socket.dev • Houston (TX)

On-site
USD 150,000 - 200,000
Base + equity
Fast-growing startup
Growth/progression plan
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Equity
Competitive compensation
Benefits package (medical, dental, V)
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • New York (NY)

On-site
USD 150,000 - 215,000
Equity
Medical insurance
Dental insurance
+4
Infrastructure Software Engineer, Fleet & Automation Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation Houston; New York; San Francisco; Seattle

Nscale • Northern (KY)

Hybrid
USD 150,000 - 215,000
Base + equity
Equity
Growth opportunities
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Equity
Base salary + equity
Career progression
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • Houston (TX)

On-site
USD 150,000 - 215,000
Base salary + equity
Annual reviews
Growth opportunities
Senior Software Engineering Manager - Fleet Management
Senior Software Engineering Manager - Fleet Management

nscaleoperationsukltd • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Software Engineering Manager - Fleet Management
Senior Software Engineering Manager - Fleet Management

Socket.dev • Seattle (WA)

On-site
USD 300,000 - 350,000
Equity
Bonus potential
Flexible work policy
+1
Infrastructure Operations Engineer Greensboro, NC
Infrastructure Operations Engineer Greensboro, NC

Nscale • Winston-Salem (NC)

Hybrid
USD 100,000 - 160,000
Competitive package (base + equity)
Flexible workplace and support for personal growth
Medical, dental, and vision benefits