Software Engineer - Fleet Automation

Career Techniques

Dallas (TX)

Hybrid

USD 120,000 - 160,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Career Techniques in Dallas, TX seeks an experienced backend engineer to design, build, and maintain fleet automation services and internal platforms for provisioning, configuration, and lifecycle management of large-scale GPU and CPU compute nodes. You will develop APIs and service integrations, build and maintain backend services, and instrument systems with Prometheus and Grafana to ensure reliability.

Proficiency in Go, C#, or TypeScript is required with 5+ years of production backend

Qualifications

  • Bachelor’s degree or equivalent practical experience in CS, software engineering, or related field.
  • 5+ years of software engineering experience building production backend services or infrastructure automation tooling.
  • Proficiency in Go, C#, or TypeScript.
  • Experience with relational and NoSQL databases to support stateful automation workflows.
  • Solid understanding of Linux systems and debugging on Ubuntu or RHEL.
  • Experience building and maintaining CI/CD pipelines and observability stacks (Prometheus, Grafana, ELK).
  • Exposure to GPU compute infrastructure and NVIDIA tooling is a strong plus.

Responsibilities

  • Design, build, and maintain fleet automation services and internal platforms for provisioning, configuration, and lifecycle management of large-scale GPU and CPU compute nodes.
  • Develop APIs and service integrations to enable deployment, imaging, validation, and decommissioning with minimal manual intervention.
  • Build and maintain backend services in Go, C#, and TypeScript focusing on reliability and long-term maintainability.
  • Design data models for automation workflows across relational and NoSQL databases.
  • Build CI/CD pipelines that gate changes and run automated hardware validation tests.
  • Instrument systems for observability with metrics, alerts, and dashboards (Prometheus, Grafana).
  • Participate in on-call rotations; own incident response and reliability improvements across the fleet.
  • Identify gaps and champion engineering solutions to reduce operational toil at scale.

Skills

Go
C#
TypeScript
Linux
CI/CD
Observability
Communication

Education

Bachelor’s Degree in Computer Science or related field

Tools

Prometheus
Grafana
ELK
Kafka
DCGM
nvidia-smi
NVIDIA Container Toolkit

Job description

RESPONSIBILITIES
  • Design, build, and maintain fleet automation services and internal platforms for provisioning, configuration, and lifecycle management of large-scale GPU and CPU compute nodes.
  • Develop APIs and service integrations that enable Infrastructure and Operations teams to deploy, image, validate, and decommission hardware with minimal manual intervention.
  • Build and maintain backend services in Go, C#, and TypeScript with a strong focus on reliability, testability, and long-term maintainability.
  • Design and evolve data models and persistent state for automation workflows, working across relational and NoSQL databases as appropriate.
  • Build and maintain CI/CD pipelines that gate configuration changes, run automated hardware validation tests, and promote changes safely across environments.
  • Instrument systems for observability — designing metrics, alerts, and dashboards in Prometheus and Grafana that provide real‑time fleet health visibility to on‑call teams.
  • Participate in on‑call rotations; own incident response, post‑mortems, and follow‑through on reliability improvements across the fleet.
  • Identify systemic gaps in fleet reliability and efficiency and champion engineering solutions that reduce operational toil at scale.
REQUIREMENTS
  • Bachelor’s Degree in Computer Science, Software Engineering, or equivalent practical experience.
  • 5+ years of software engineering experience building production backend services or infrastructure automation tooling.
  • Proficiency in Go, C#, or TypeScript.
  • Experience designing and working with relational and NoSQL databases to support stateful automation workflows and internal platform services.
  • Solid understanding of Linux systems — networking, storage, process management, and debugging on Ubuntu or RHEL variants.
  • Experience building and maintaining CI/CD pipelines and observability stacks (Prometheus, Grafana, Alertmanager, ELK) in a production environment.
  • Familiarity with GPU compute infrastructure and NVIDIA tooling (DCGM, nvidia-smi, NVIDIA Container Toolkit) is a strong plus.
  • Exposure to event‑driven architectures or messaging platforms (e.g. Kafka) is a plus for teams building automation workflows across distributed services.
  • Strong communication skills and a collaborative mindset — comfortable navigating ambiguity, taking initiative, and working across Infrastructure, Operations, and Research teams.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NMC2 • Dallas (TX), Northern (KY)

On-site
USD 120,000 - 150,000
Senior Infrastructure Software Engineer
Senior Infrastructure Software Engineer

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 190,000
Software - Fleet Automation
Software - Fleet Automation

Glocomms • Dallas (TX)

Hybrid
USD 250,000 - 300,000
Senior Fleet Automation Engineer (Go/C#/TypeScript)
Senior Fleet Automation Engineer (Go/C#/TypeScript)

NorthMark Strategies • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Lunch stipend
Employer-paid health & dental & vision
Parental leave 16 weeks
+4
Senior Software Engineer, Fleet Intelligence Backend
Senior Software Engineer, Fleet Intelligence Backend

NVIDIA AI • Seattle (WA)

On-site
USD 140,000 - 200,000
Equity
Benefits
Senior Fleet Automation Backend Engineer (Go/C#/TS)
Senior Fleet Automation Backend Engineer (Go/C#/TS)

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 160,000
Software Engineer, Fleet Hardware Health
Software Engineer, Fleet Hardware Health

Cloudjobs • San Francisco (CA)

On-site
USD 140,000 - 180,000
Senior Software Engineer, Fleet Intelligence Agent Systems
Senior Software Engineer, Fleet Intelligence Agent Systems

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 140,000 - 200,000
Equity
Benefits
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Software Engineer, Fleet Infrastructure
Software Engineer, Fleet Infrastructure

Cloudjobs • New York (NY)

On-site
USD 140,000 - 210,000
Relocation assistance