Fleet Automation Engineer – HPC & GPU Compute

NorthMark Strategies

Dallas (TX)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Company-Paid Lunch Stipend
Company-Paid Benefits: Medical, Dental
401(k) matching up to 6%
Flexible/Optional Employee Benefits
25 days PTO plus holidays

Job summary

NorthMark Strategies LLC is seeking a Software Engineer for the Fleet Automation team in the HPC & Infrastructure organization in Dallas. You will design and implement automation platforms, internal services, and APIs to provision, configure, and maintain large-scale compute fleets, including GPU and CPU nodes.

You will write production-grade Go, C#, and TypeScript, collaborate across Infrastructure, Operations, and Research teams, and help evolve observability with Prometheus and Grafana while

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 5+ years of software engineering experience building production backend services or infrastructure automation tooling.
  • Proficiency in Go, C#, or TypeScript.
  • Experience designing and working with relational and NoSQL databases to support stateful automation workflows and internal platform services.
  • Solid understanding of Linux systems — networking, storage, process management, and debugging on Ubuntu or RHEL variants.
  • Experience building and maintaining CI/CD pipelines and observability stacks (Prometheus, Grafana, Alertmanager, ELK) in a production environment.
  • Familiarity with GPU compute infrastructure and NVIDIA tooling (DCGM, nvidia-smi) is a strong plus.
  • Exposure to event-driven architectures or messaging platforms (e.g. Kafka) is a plus for teams building automation workflows across distributed services.
  • Strong communication skills and a collaborative mindset — comfortable navigating ambiguity, taking initiative, and working across Infrastructure, Operations, and Research teams.

Responsibilities

  • Design, build, and maintain fleet automation services and internal platforms for provisioning, configuration, and lifecycle management of large-scale GPU and CPU compute nodes.
  • Develop APIs and service integrations that enable Infrastructure and Operations teams to deploy, image, validate, and decommission hardware with minimal manual intervention.
  • Build and maintain backend services in Go, C#, and TypeScript with a strong focus on reliability, testability, and long-term maintainability.
  • Design and evolve data models and persistent state for automation workflows, working across relational and NoSQL databases as appropriate.
  • Build and maintain CI/CD pipelines that gate configuration changes, run automated hardware validation tests, and promote changes safely across environments.
  • Instrument systems for observability — designing metrics, alerts, and dashboards in Prometheus and Grafana that provide real-time fleet health visibility to on-call teams.
  • Participate in on-call rotations; own incident response, post-mortems, and follow-through on reliability improvements across the fleet.
  • Identify systemic gaps in fleet reliability and efficiency and champion engineering solutions that reduce operational toil at scale.

Skills

Go
C#
TypeScript
Linux
CI/CD
Observability
Communication

Education

Bachelor’s Degree in CS/Software Engineering or equivalent

Tools

Prometheus
Grafana
Alertmanager
ELK
Kafka
NVIDIA DCGM/nvidia-smi

Job description

NorthMark Strategies LLC is seeking a Software Engineer for the Fleet Automation team in the HPC & Infrastructure organization in Dallas. You will design and implement automation platforms, internal services, and APIs to provision, configure, and maintain large-scale compute fleets, including GPU and CPU nodes.

You will write production-grade Go, C#, and TypeScript, collaborate across Infrastructure, Operations, and Research teams, and help evolve observability with Prometheus and Grafana while

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Fleet Automation Engineer – HPC Infra
Fleet Automation Engineer – HPC Infra

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 150,000
Fleet Automation Engineer – HPC Infrastructure
Fleet Automation Engineer – HPC Infrastructure

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lunch stipend
Medical benefits (HDHP)
401(k) match
Senior Fleet Automation Engineer (Go/C#/TypeScript)
Senior Fleet Automation Engineer (Go/C#/TypeScript)

NorthMark Strategies • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Lunch stipend
Employer-paid health & dental & vision
Parental leave 16 weeks
+4
Fleet Automation Engineer — HPC Infrastructure
Fleet Automation Engineer — HPC Infrastructure

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 110,000 - 160,000
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 150,000
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 110,000 - 160,000
GPU HPC Systems Engineer — Fleet Reliability & Automation
GPU HPC Systems Engineer — Fleet Reliability & Automation

OpenAI • California (MO)

On-site
USD 180,000 - 260,000
Fleet & Automation Infrastructure Engineer — AI/HPC
Fleet & Automation Infrastructure Engineer — AI/HPC

Nscale • Houston (TX)

On-site
USD 150,000 - 215,000
Base salary + equity
Annual reviews
Growth opportunities
Senior HPC Hardware Engineer: Lead the AI Compute Fleet
Senior HPC Hardware Engineer: Lead the AI Compute Fleet

NorthMark Strategies • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Lunch stipend
Medical benefits (employer-paid)
Parental leave 16 weeks
+2
Senior HPC Hardware Architect - AI & GPU Compute
Senior HPC Hardware Architect - AI & GPU Compute

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 150,000 - 200,000