HPC Systems Engineer

NorthMark Strategies

United States

On-site

USD 140,000 - 175,000

Full time

31 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Lunch stipend
Medical, dental, vision
Parental leave
401(k) match
Life insurance
Disability coverage

Job summary

NorthMark Strategies in Dallas, TX is hiring an HPC Systems Engineer to join its infrastructure engineering organization at the Victory Commons office. This role focuses on designing, integrating, and delivering an end-to-end HPC platform across compute, storage, networking, Kubernetes, automation, and data center infrastructure.

The Systems Engineer will collaborate across domain teams to resolve cross-functional dependencies, anticipate integration risks, and ensure performance, resiliency,

Qualifications

  • Bachelor's degree in Engineering, Computer Science, or equivalent experience.
  • 8+ years of experience in systems engineering, HPC, cloud infrastructure, or data center infrastructure.
  • Deep expertise in at least one infrastructure domain with broad cross-domain knowledge.
  • Familiarity with GPU compute, InfiniBand, Lustre/GPFS, Kubernetes or Slurm.
  • Experience with Terraform, Ansible, and CI/CD pipelines.
  • Strong understanding of system-level dependencies and architectural tradeoffs.
  • Proven ability to lead technical discussions across teams without direct authority.
  • Excellent communication skills and ability to present technical concepts.

Responsibilities

  • Drive cross-domain technical alignment across compute, storage, networking, Kubernetes, automation, and data center infrastructure teams.
  • Lead integrated design reviews for new HPC deployments and major infrastructure initiatives.
  • Identify design gaps, scalability constraints, and integration risks throughout delivery.
  • Define end-to-end architecture with engineering standards and readiness criteria.
  • Ensure designs meet performance, resiliency, scalability, and long-term support needs.
  • Develop system-level specifications for interoperation of components.
  • Participate in failure analysis and resiliency reviews ahead of production turnover.
  • Collaborate with automation teams to improve deployment consistency and lifecycle management.
  • Serve as escalation point for cross-team integration issues and arbitrate tradeoffs.
  • Influence standards and direction through design authorship and mentorship.

Skills

HPC infrastructure
Kubernetes
Automation
CI/CD pipelines
Communication skills

Education

Bachelor's degree in Engineering, Computer Science, or equivalent

Tools

Terraform
Ansible
CI/CD tooling

Job description

The Company

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision‑making, accelerating discovery and driving faster innovation.

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision‑making, accelerating discovery and driving faster innovation.

THE POSITION

NMC² is hiring an HPC Systems Engineer to join its infrastructure engineering organization in Dallas, TX, based out of the Victory Commons office. This is a cross-domain technical role responsible for ensuring that HPC infrastructure is designed, integrated, and delivered as a cohesive end-to-end platform rather than a collection of independently optimized parts. The Systems Engineer works horizontally across the compute, storage, networking, Kubernetes, automation, and data center infrastructure teams, driving technical alignment, resolving cross-functional dependencies, and reducing integration risk across large-scale AI and HPC deployments.

This role is built for engineers who think beyond a single technology domain. Great performance here looks like anticipating the failure modes that live between teams — the fabric assumption that breaks a storage design, the power and cooling constraint that reshapes a rack layout, the scheduler behavior that undermines a networking decision — and resolving them in design rather than in production.

The Systems Engineer translates business objectives into scalable, reliable infrastructure solutions and holds the system-level view that no individual domain team owns on its own.

The role partners closely with domain engineering leads across compute, storage, networking, and data center infrastructure, as well as with automation and platform teams, operations, and hardware and technology vendors. Influence here is earned through engineering depth, collaboration, and sound judgment rather than positional authority, and the Systems Engineer is expected to shape technical direction across teams accordingly.

Responsibilities
  • Drive cross-domain technical alignment across the compute, storage, networking, Kubernetes, automation, and data center infrastructure teams.
  • Lead integrated design reviews for new HPC deployments, platform expansions, and major infrastructure initiatives.
  • Identify technical dependencies, design gaps, scalability constraints, and integration risks throughout the delivery lifecycle.
  • Partner with domain engineering teams to define end-to-end architecture, engineering standards, and implementation readiness criteria.
  • Ensure infrastructure designs meet requirements for performance, resiliency, scalability, and long-term operational supportability.
  • Develop system-level engineering specifications that define how infrastructure components integrate and interoperate.
  • Participate in failure analysis, resiliency reviews, and operational readiness assessments ahead of production turnover.
  • Collaborate with automation teams to improve deployment consistency, validation, lifecycle management, and operational efficiency.
  • Serve as a technical point of escalation for cross-team integration issues, arbitrating architectural tradeoffs where domain priorities conflict.
  • Influence engineering direction and platform standards across the organization through design authorship, review, and technical mentorship.
Requirements
  • Bachelor's degree in Engineering, Computer Science, or equivalent experience.
  • 8+ years of experience in systems engineering, product engineering, HPC, cloud infrastructure, distributed systems, or data center infrastructure.
  • Deep expertise in at least one infrastructure domain — compute, networking, storage, Kubernetes, or data center infrastructure — with broad working knowledge across adjacent domains.
  • Working familiarity with the components of a modern AI/HPC platform, such as GPU compute (NVIDIA H100/H200 class), high-speed fabrics (InfiniBand NDR/HDR, RoCEv2), scale-out storage (VAST Data, WekaFS, Lustre, or GPFS), and Kubernetes or Slurm-based scheduling.
  • Experience with infrastructure automation and lifecycle tooling such as Terraform, Ansible, and CI/CD pipelines.
  • Strong understanding of system-level dependencies, architectural tradeoffs, and downstream operational impacts.
  • Proven ability to lead technical discussions and drive alignment across multiple engineering teams without direct authority.
  • Excellent communication skills, with the ability to present complex technical concepts to both engineering and leadership audiences.
  • Ability to thrive in a fast-paced environment with evolving requirements and aggressive delivery timelines.

It is impossible to list every requirement for, or responsibility of, any position. Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company’s needs may change over time. Therefore, the above job description is not comprehensive or exhaustive. The Company reserves the right to adjust, add to or eliminate any aspect of the above description. The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs.

Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

Benefits & Perks
  • Company-Paid Lunch Stipend: Lunch is provided via GrubHub
  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families, 16 weeks of Paid Parental Leave, Employee Assistance Program, Life insurance, Short-Term Disability and Long-Term Disability
  • 401(k): Company will match 100% of your contributions up to 6%
  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.
  • Time Off: 25 days of Paid Time Off plus 12 company holidays
EQUAL OPPORTUNITY EMPLOYER

NORTHMARK STRATEGIES LLC IS AN EQUAL EMPLOYMENT OPPORTUNITY EMPLOYER. THE COMPANY'S POLICY IS NOT TO DISCRIMINATE AGAINST ANY APPLICANT OR EMPLOYEE BASED ON RACE, COLOR, RELIGION, NATIONAL ORIGIN, GENDER, AGE, SEXUAL ORIENTATION, GENDER IDENTITY OR EXPRESSION, MARITAL STATUS, MENTAL OR PHYSICAL DISABILITY, AND GENETIC INFORMATION, OR ANY OTHER BASIS PROTECTED BY APPLICABLE LAW. THE FIRM ALSO PROHIBITS HARASSMENT OF APPLICANTS OR EMPLOYEES BASED ON ANY OF THESE PROTECTED CATEGORIES.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Hardware Engineer
Senior HPC Hardware Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 140,000 - 220,000
Lunch stipend
Medical benefits - employer-paid
Dental & Vision benefits
+6
HPC Orchestration Architect
HPC Orchestration Architect

NorthMark Strategies • United States

On-site
USD 180,000 - 240,000
Lunch stipend
Medical insurance
Parental leave
+2
HPC Systems Engineer
HPC Systems Engineer

NorthMark Compute and Cloud LLC • Dallas (TX)

On-site
USD 140,000 - 210,000
Lunch stipend
Medical, dental and vision coverage
Parental leave 16 weeks
+6
HPC Orchestration Architect
HPC Orchestration Architect

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 180,000 - 260,000
Lunch stipend
Employer-paid medical benefits
Parental leave (16 weeks)
+2
HPC Storage Engineer
HPC Storage Engineer

NorthMark Strategies • United States

On-site
USD 120,000 - 160,000
Lunch stipend
Employer-paid medical
Dental and vision
+7
HPC Network Engineer
HPC Network Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 120,000 - 150,000
Lunch stipend
Medical/dental/vision benefits
Parental leave (16 weeks)
+3
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 120,000 - 180,000
Lunch stipend
Medical benefits
Paid parental leave
+3
HPC Storage Architect
HPC Storage Architect

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 140,000 - 190,000
Lunch stipend
Medical benefits (employer-paid)
Dental and Vision benefits
+6
Director, Strategic Infrastructure Partnerships
Director, Strategic Infrastructure Partnerships

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 180,000 - 230,000
Lunch stipend
Employer-paid medical, dental, vision
401(k) match
+4
Emerging Network Architect
Emerging Network Architect

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 140,000 - 230,000
Lunch stipend
Medical benefits
Dental & Vision
+5