Senior Go Engineer – GPU Infrastructure Automation

GTN Technical Staffing

Dallas (TX)

Hybrid

USD 150,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Relocation assistance
Hybrid work model

Job summary

GTN Technical Staffing seeks a Senior Go Engineer to join the Fleet Automation team supporting large-scale HPC and GPU infrastructure. Hybrid work arrangement with three days onsite and relocation assistance available.

The role focuses on building backend services, automating deployment, and managing hardware lifecycle across a Linux environment. Candidates should have 5+ years in backend/backend automation and strong Go/C#/TypeScript experience.

Qualifications

  • 5+ years of software engineering experience building production backend services or automation tooling.
  • Experience designing APIs, backend services, and distributed/stateful systems.
  • Experience with relational and/or NoSQL databases.
  • Ubuntu and/or RHEL environments.
  • Experience building and maintaining CI/CD pipelines.

Responsibilities

  • Design and build fleet automation platforms for provisioning, configuration, validation, and lifecycle management of GPU and CPU compute nodes.
  • Develop internal services and APIs to automate hardware deployment, imaging, remediation, and decommissioning.
  • Build reliable backend services using Go, C#, and/or TypeScript.
  • Design data models and persistent state for automation workflows using relational and NoSQL databases.
  • Develop and maintain CI/CD pipelines for infrastructure and configuration changes.
  • Automate hardware validation and testing across large-scale compute environments.
  • Build observability, monitoring, dashboards, and alerting using Prometheus, Grafana, Alertmanager, and ELK.
  • Partner with Infrastructure, Network, Operations, and Research teams to automate repeatable processes.
  • Participate in on-call rotations and incident response.
  • Identify systemic infrastructure issues and improve fleet reliability and scalability.

Skills

Go
C#
TypeScript

Education

Bachelor's degree

Tools

Prometheus
Grafana
Alertmanager
ELK
Kafka

Job description

Work Model: Hybrid, 3 days onsite

Relocation: Available

Compensation: base + bonus

Employment Type: Direct Hire

Overview

Our client is seeking a Senior Go Engineer to join its Fleet Automation team supporting large-scale HPC and GPU infrastructure.

This team builds the software, automation, and internal platforms used to provision, configure, monitor, and manage hundreds of high-performance GPU and CPU compute nodes. The environment sits at the intersection of software engineering, infrastructure, and distributed systems, with a strong focus on eliminating manual operational work through scalable automation.

This is a hands‑on engineering role for someone who enjoys building backend services while also understanding how those services interact with Linux systems, physical hardware, networking, storage, and GPU infrastructure.

What You’ll Do
  • Design and build fleet automation platforms for provisioning, configuration, validation, and lifecycle management of GPU and CPU compute nodes.
  • Develop internal services and APIs that automate hardware deployment, imaging, remediation, and decommissioning.
  • Build reliable backend services using Go, C#, and/or TypeScript.
  • Design data models and persistent state for automation workflows using relational and NoSQL databases.
  • Develop and maintain CI/CD pipelines for infrastructure and configuration changes.
  • Automate hardware validation and testing across large-scale compute environments.
  • Build observability, monitoring, dashboards, and alerting using tools such as Prometheus, Grafana, Alertmanager, and ELK.
  • Partner closely with Infrastructure, Network, Operations, and Research teams to identify operational pain points and automate repeatable processes.
  • Participate in on‑call rotations and support incident response, root‑cause analysis, and post‑incident reliability improvements.
  • Identify systemic infrastructure issues and develop engineering solutions that improve fleet reliability, efficiency, and scalability.
Required Experience
  • 5+ years of software engineering experience building production backend services, infrastructure platforms, or automation tooling.
  • Strong development experience with at least one of the following:
  • Go
  • C#
  • TypeScript
  • Experience designing APIs, backend services, and distributed or stateful systems.
  • Strong experience with relational and/or NoSQL databases.
  • Networking
  • Storage
  • Process management
  • Ubuntu and/or RHEL environments
  • Experience building and maintaining CI/CD pipelines.
  • Hands‑on experience with production monitoring and observability platforms such as:
  • Grafana
  • Alertmanager
  • ELK
  • Strong troubleshooting and problem‑solving skills across both software and infrastructure environments.
Preferred Experience
  • Experience supporting GPU, HPC, AI/ML, or large‑scale compute infrastructure.
  • Familiarity with NVIDIA technologies such as:
  • DCGM
  • nvidia-smi
  • NVIDIA Container Toolkit
  • Experience with bare‑metal provisioning and hardware lifecycle automation.
  • Exposure to event‑driven architectures and messaging platforms such as Kafka.
  • Experience working with infrastructure, network, SRE, or platform engineering teams.
  • Bachelor’s degree in Computer Science, Software Engineering, or equivalent practical experience.
Why Consider This Opportunity?
  • Work directly on large‑scale GPU and HPC infrastructure supporting advanced AI workloads.
  • Build automation platforms that have direct impact on infrastructure reliability and scalability.
  • Highly technical environment combining software engineering, systems engineering, and infrastructure automation.
  • Relocation assistance available for candidates moving to Dallas.
  • Hybrid schedule with three days per week onsite.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Go Engineer - GPU Fleet Automation
Senior Go Engineer - GPU Fleet Automation

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 150,000 - 190,000
Relocation assistance
Hybrid work model
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 150,000
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Software - Fleet Automation
Software - Fleet Automation

Glocomms • Dallas (TX)

Hybrid
USD 250,000 - 300,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Calance • Costa Mesa (CA)

Hybrid
USD 180,000 - 240,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 200,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits
HPC Developer
HPC Developer

Autonomai Recruitment • Chicago (IL)

On-site
USD 110,000 - 190,000
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Cloudjobs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Restricted Stock Units
Health insurance
Vision insurance
+15