Senior Infrastructure Software Engineer

GTN Technical Staffing

Dallas (TX)

On-site

USD 140,000 - 190,000

Full time

43 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

GTN Technical Staffing seeks a Senior Infrastructure Software Engineer to build software, automation, and internal platforms used to provision, configure, monitor, and manage large-scale GPU and CPU compute infrastructure. This software-first role operates at the intersection of backend development, infrastructure automation, Linux systems, distributed systems, and high-performance compute.

The ideal candidate combines strong software engineering with infra depth to enable scalable, reliable

Qualifications

  • 5+ years building backend services, infra platforms, or automation tooling.
  • Strong Go, C#, or TypeScript backend experience.
  • Experience designing APIs and stateful/distributed systems.
  • Proficient with relational and NoSQL databases.
  • Deep Linux knowledge: networking, storage, and troubleshooting.
  • CI/CD pipelines experience and production monitoring.
  • Experience with Ubuntu and/or RHEL environments.

Responsibilities

  • Design and build automation platforms for provisioning, configuration, validation, and lifecycle management of GPU and CPU compute nodes.
  • Develop backend services and APIs for hardware deployment, imaging, remediation, and decommissioning.
  • Build reliable software using Go, C#, TypeScript, or similar backend languages.
  • Design data models and persistent state for infrastructure automation workflows.
  • Develop and maintain CI/CD pipelines for infrastructure and configuration changes.
  • Automate hardware validation and testing across large compute environments.
  • Build monitoring, observability dashboards, and alerting using Prometheus, Grafana, Alertmanager, ELK.
  • Partner with Infrastructure, Network, Operations, and Research teams to automate operational workflows.
  • Participate in incident response, root-cause analysis, and reliability improvement efforts.
  • Identify systemic infrastructure issues and develop software solutions that improve scalability and reliability.

Skills

Go
C#
TypeScript
APIs
Distributed systems
Linux
CI/CD
Monitoring

Education

Bachelor's degree in CS/SE

Tools

Prometheus
Grafana
Alertmanager
ELK
Ubuntu
RHEL
NVIDIA Container Toolkit

Job description

Compensation: Competitive Base Salary + Performance Bonus

Overview

Our client is seeking a Senior Infrastructure Software Engineer to build the software, automation, and internal platforms used to provision, configure, monitor, and manage large-scale GPU and CPU compute infrastructure.

This is a software-first engineering role operating at the intersection of backend development, infrastructure automation, Linux systems, distributed systems, and high-performance compute.

The ideal candidate enjoys building production-quality services and APIs while also understanding how software interacts with physical hardware, networking, storage, and GPU infrastructure.

Key Responsibilities
  • Design and build automation platforms for provisioning, configuration, validation, and lifecycle management of GPU and CPU compute nodes.
  • Develop backend services and APIs for hardware deployment, imaging, remediation, and decommissioning.
  • Build reliable software using Go, C#, TypeScript, or similar backend languages.
  • Design data models and persistent state for infrastructure automation workflows.
  • Develop and maintain CI/CD pipelines for infrastructure and configuration changes.
  • Automate hardware validation and testing across large compute environments.
  • Build monitoring, observability, dashboards, and alerting using Prometheus, Grafana, Alertmanager, ELK, or similar tools.
  • Partner with Infrastructure, Network, Operations, and Research teams to automate operational workflows.
  • Participate in incident response, root-cause analysis, and reliability improvement efforts.
  • Identify systemic infrastructure issues and develop software solutions that improve scalability and reliability.
Required Qualifications
  • 5+ years of software engineering experience building backend services, infrastructure platforms, or automation tooling.
  • Strong development experience with Go, C#, TypeScript, or another modern backend language.
  • Experience designing APIs, backend services, and distributed or stateful systems.
  • Strong experience with relational and/or NoSQL databases.
  • Strong Linux knowledge, including networking, storage, process management, and system troubleshooting.
  • Experience with Ubuntu and/or RHEL environments.
  • Experience building and maintaining CI/CD pipelines.
  • Hands-on experience with production monitoring and observability platforms.
  • Strong troubleshooting and problem-solving skills across software and infrastructure environments.
Preferred Experience
  • GPU, HPC, AI/ML, or large-scale compute infrastructure.
  • NVIDIA technologies such as DCGM, nvidia-smi, or NVIDIA Container Toolkit.
  • Bare-metal provisioning and hardware lifecycle automation.
  • Kafka or other event-driven architectures.
  • Experience working with infrastructure, SRE, network, or platform engineering teams.
  • Bachelor's degree in Computer Science, Software Engineering, or equivalent practical experience.
Ideal Candidate

The ideal candidate is a software engineer first with strong backend development skills and enough infrastructure depth to build automation for complex, large-scale compute environments.

The strongest profiles will combine software engineering, Linux systems, distributed infrastructure, automation, and observability, with GPU or HPC experience considered a strong plus.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer - Fleet Automation
Software Engineer - Fleet Automation

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 160,000
Senior GPU & HPC Infrastructure Engineer
Senior GPU & HPC Infrastructure Engineer

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 190,000
GPU Systems Infrastructure Engineer
GPU Systems Infrastructure Engineer

Blue Signal Search • Fremont (CA)

On-site
USD 120,000 - 170,000
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NMC2 • Dallas (TX), Northern (KY)

On-site
USD 120,000 - 150,000
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
HPC Developer
HPC Developer

Autonomai Recruitment • Chicago (IL)

On-site
USD 110,000 - 190,000
HPC Developer
HPC Developer

Autonomai Recruitment • New York (NY)

On-site
USD 130,000 - 195,000