Senior Software Engineer, HPC Scheduling

GTN Technical Staffing

Dallas (TX)

On-site

USD 170,000 - 250,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance
Performance bonus
Company-paid benefits

Job summary

GTN Technical Staffing is seeking a Senior Software Engineer, HPC Scheduling to design, build, and maintain large-scale scheduling software supporting HPC, AI, and research workloads. You will work on distributed systems, backend services, APIs, and automation, with Armada as a core project, using Go and cloud-native tech.

The role emphasizes production-ready code, code reviews, and operating in Kubernetes and Linux environments.

Qualifications

  • Strong software engineering fundamentals and production-grade coding practices.
  • Experience building backend services, APIs, distributed systems, or infrastructure software in production.

Responsibilities

  • Design, write, test, and review high-quality production code, primarily in Go.
  • Build and maintain scalable backend services, APIs, and distributed systems supporting high-demand workloads.
  • Contribute to Armada and related internal scheduling, orchestration, and platform services.
  • Develop tooling and automation that improves platform reliability, developer productivity, and operational efficiency.
  • Apply strong software architecture principles to ensure systems are maintainable, correct, and scalable.

Skills

Go
Distributed systems
Backend services
Kubernetes
Linux
PostgreSQL
Cloud environments (AWS/GCP/Azure)
Testing practices
Code review
Observability

Tools

Prometheus
Grafana

Job description

Senior Software Engineer, HPC Scheduling

Type: Direct Hire

Relocation: Available for non-local candidates

Compensation

Base salary: $170,000 – $250,000 + performance bonus

Overview

GTN is seeking a Senior Software Engineer, HPC Scheduling to help design, build, and maintain large-scale scheduling software that supports demanding HPC, AI, research, and production workloads.

This role sits on a highly technical scheduling team responsible for developing distributed systems, backend services, APIs, tooling, and automation that keep a high-scale compute platform reliable, performant, and maintainable.

Much of the work centers around Armada, an open-source project built and maintained by the team, along with internal services and platform tooling written primarily in Go. This is a hands-on engineering role focused on writing clean, well-tested code, reviewing designs, solving complex distributed systems problems, and owning production-quality software.

The ideal candidate is a strong software engineer with excellent coding fundamentals, experience building backend or distributed systems, and a practical understanding of how software runs in cloud, Linux, Kubernetes, and production infrastructure environments.

Key Responsibilities
Software Engineering & Platform Development
  • Design, write, test, and review high-quality production code, primarily in Go
  • Build and maintain scalable backend services, APIs, and distributed systems supporting high-demand workloads
  • Contribute to Armada and related internal scheduling, orchestration, and platform services
  • Develop tooling and automation that improves platform reliability, developer productivity, and operational efficiency
  • Apply strong software architecture principles to ensure systems are maintainable, correct, and scalable
Distributed Systems & Infrastructure
  • Build services that operate reliably across large-scale HPC and AI infrastructure environments
  • Work with Kubernetes-based orchestration, containerized services, and modern deployment workflows
  • Develop and debug software in Linux environments using command-line and system-level tooling
  • Apply networking fundamentals to troubleshoot, optimize, and improve platform connectivity and performance
  • Independently diagnose and resolve complex issues across software and infrastructure layers
Data, Reliability & Operations
  • Manage and optimize data interactions across relational and non-relational data stores, with emphasis on PostgreSQL
  • Contribute to CI/CD pipelines, automated testing, observability, and engineering best practices
  • Use monitoring, logging, and runtime tools such as Prometheus, Grafana, or similar platforms
  • Think critically about correctness, edge cases, performance, and failure modes
  • Stay current with emerging technologies and apply new approaches where they improve platform outcomes
Required Experience
  • Strong software engineering fundamentals, including data structures, algorithms, system design, and maintainable code practices
  • Proficiency in Go or another statically typed language, with the ability to quickly ramp into Go-based codebases
  • Experience building backend services, APIs, distributed systems, or infrastructure software in production environments
  • Familiarity with cloud environments such as AWS, GCP, or Azure
  • Experience with Linux-based development and debugging
  • Familiarity with Kubernetes, containers, or modern deployment pipelines
  • Experience with PostgreSQL or similar relational databases
  • Understanding of observability practices, including monitoring, logging, metrics, and alerting
  • Strong testing mindset with focus on correctness, reliability, and failure scenarios
  • Ability to work independently, review code thoughtfully, and contribute in a collaborative engineering team
Preferred Experience
  • Experience with HPC, AI infrastructure, batch scheduling, workload orchestration, or large-scale compute platforms
  • Hands-on experience with Kubernetes scheduling, multi-cluster systems, or distributed job orchestration
  • Contributions to open-source projects or experience working in open-source engineering environments
  • Experience with non-relational databases, message queues, event-driven systems, or high-throughput platforms
  • Familiarity with performance optimization, reliability engineering, or production platform operations
Ideal Profile

The ideal candidate is a hands-on software engineer who enjoys building infrastructure software that operates at scale. They write clean, tested code, understand distributed systems tradeoffs, and are comfortable working close to production infrastructure. They do not need to come directly from an HPC background, but they should have strong backend engineering fundamentals and an interest in solving complex scheduling, orchestration, and platform reliability challenges.

Why This Role
  • Work on high-scale HPC and AI infrastructure supporting demanding production workloads
  • Contribute to Armada, an open-source scheduling platform
  • Join a senior, collaborative engineering team with real ownership over technical direction
  • Build software that directly impacts platform reliability, performance, and scalability
  • Competitive compensation, performance bonus, relocation support, and 100% company-paid benefits
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Scheduling Engineer (Go)
Senior HPC Scheduling Engineer (Go)

GTN Technical Staffing • Dallas (TX)

On-site
USD 170,000 - 250,000
Relocation assistance
Performance bonus
Company-paid benefits
Software Engineer, HPC Scheduling
Software Engineer, HPC Scheduling

NMC2 • Dallas (TX)

On-site
USD 100,000 - 130,000
Software Engineer, HPC Scheduling
Software Engineer, HPC Scheduling

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 90,000 - 120,000
Software Engineer, HPC Scheduling
Software Engineer, HPC Scheduling

NorthMark Compute and Cloud LLC • Dallas (TX)

On-site
USD 100,000 - 130,000
Company-Paid Lunch Stipend
Employer-Paid Medical Benefits
Paid Parental Leave
+2
Software Engineer, HPC Scheduling
Software Engineer, HPC Scheduling

NorthMark Strategies • Dallas (TX)

On-site
USD 90,000 - 130,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical and Benefits
401(k) with company match
+2
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Senior Software Engineer - GPUaaS
Senior Software Engineer - GPUaaS

Armada • Bellevue (WA)

On-site
USD 158,000 - 197,000
Equity
Subsidized benefits
Medical, dental, and vision
+3
Backend Software Engineer
Backend Software Engineer

Glint Tech Solutions LLC • San Francisco (CA)

Hybrid
USD 250,000
Competitive equity package
Comprehensive benefits
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Colossus Technologies Group • Boston (MA)

Hybrid
USD 180,000 - 220,000
Health & wellness benefit
Competitive equity package
Hybrid work flexibility
Senior Software Engineer - HPC
Senior Software Engineer - HPC

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits