Compute Platform Engineer

NorthMark Compute & Cloud

Dallas (TX)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NorthMark Compute & Cloud is seeking a Compute Platform Engineer to own the reliability and performance of our HPC compute platforms that support research and production workloads. You will ensure CPU/GPU nodes run at peak efficiency, coordinate with hardware vendors, and implement automation and IaC practices to scale operations while maintaining security and compliance across the data center.

Join a team that values engineering excellence, proactive problem solving, and cross-team

Qualifications

  • 3+ years of hands-on experience supporting large-scale compute platforms.
  • Proficiency with HPE server infrastructure (ProLiant, Apollo) and NVIDIA GPUs (A100, H200).
  • Strong understanding of server architecture (UEFI/BIOS, PCIe, iLO, BMC).
  • Ability to resolve complex hardware issues and manage vendor relationships.
  • Experience with automation tools (Ansible, Terraform) and CI/CD.
  • Working knowledge of Linux in HPC/latency-sensitive environments.
  • Familiarity with basic networking (DNS, DHCP, VLANs, switching, routing).
  • Basic knowledge of Kubernetes and OpenStack (preferred).
  • Experience in data center operations and process adherence.
  • Excellent communication and coordination with cross-functional teams and external partners.

Responsibilities

  • Design and manage a high-performance compute infrastructure of CPU/GPU nodes.
  • Manage firmware/BIOS lifecycle across HPC/AI fleet, baselines to rollout.
  • Troubleshoot hardware components and drive fixes with vendors.
  • Monitor performance and implement improvements for reliability.
  • Automate health checks and onboarding workflows for safe deployments.
  • Collaborate with vendors on firmware issues with clear repro cases.
  • Suggest process, tooling, and architectural improvements.
  • Perform capacity planning for scale-out.
  • Collaborate with teams to integrate hardware improvements.
  • Apply security hardening of platform and related systems.
  • Mentor junior engineers and foster continuous learning.
  • Provide SME guidance for infrastructure issues.
  • Leverage IaC for scalable management.

Skills

Linux
Vendor management
Automation tools
Networking basics
Cross-team communication
Infrastructure as Code

Tools

Ansible
Terraform
CI/CD
Kubernetes
OpenStack

Job description

The Company

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation.

The Position

The Compute Platform Engineer role is responsible for the day-to-day reliability, performance, and operational health of our high-performance compute platforms that support critical research and production workloads. This position focuses on maintaining and troubleshooting CPU and GPU infrastructure, coordinating with vendors, and ensuring systems operate consistently at scale. Working closely with platform, infrastructure, and operations teams, the role plays a key part in sustaining a stable compute environment.

We are seeking a highly skilled and motivated Engineer to join our Compute Platform Management team. In this role, you will take ownership of the reliability and operational excellence of our high-performance computing infrastructure, which underpins our firm’s research and production workloads.

As a Compute Platform Engineer, you will be responsible for identifying and resolving hardware issues, coordinating with vendors and ensuring compute nodes (CPU and GPU) maintain peak performance. This contract role is ideal for someone who thrives in technically demanding environments and is eager to contribute to the continuous evolution of our compute platform.

Responsibilities
  • Designing, configuring, and manage a High performance compute infrastructure made up of GPU and CPU nodes
  • Manage the full firmware/BIOS lifecycle across our HPC/AI fleet – from baselines and validation through rollout and compliance.
  • Troubleshoot hardware components (CPU, GPU, DPU, NVSwitch, NICs, memory, PSU, BMC) and guide replacement or configuration changes. Diagnose and automate recurring hardware issues to improve reliability and reduce recovery time.
  • Work on the latest AI platforms from day one (e.g., NVL72 / Grace Blackwell), ensuring they are stable, performant, and ready for production use.
  • Monitoring hardware performance, identifying areas for improvement, and implementing solutions
  • Automate health checks and onboarding workflows to accelerate safe deployment.
  • Collaborate with vendors on firmware issues – providing clear repro cases, logs, and impact to drive fixes and improvements.
  • Recommend process, tooling, and architectural improvements to strengthen platform operations.
  • Performing diagnostics, tuning, and capacity planning to ensure smooth scale-out
  • Performing analysis of existing hardware lifecycle processes and providing recommendations for improvement and optimization
  • Collaborating with various teams to integrate hardware improvements and align with organizational goals
  • Implementing best practices for security hardening of the platform and associated systems
  • Mentoring junior engineers and fostering a culture of continuous learning and improvement
  • Acting as a subject matter expert, providing guidance and support for infrastructure-related issues
  • Leveraging Infrastructure as Code (IaC) methodologies to ensure efficient and scalable infrastructure management
Requirements
  • 3+ years of hands-on experience supporting large-scale compute platforms
  • Proficiency with HPE server infrastructure, such as ProLiant and Apollo, and NVIDIA GPUs, including A100 and H200
  • Solid understanding of server architecture, including UEFI/BIOS, PCIe devices and out-of-band management systems, such as iLO and BMC)
  • Proven ability to resolve complex hardware issues and manage vendor relationships
  • Familiarity with automation tools such as Ansible, Terraform and CI/CD systems
  • Working knowledge of Linux in high-performance or latency-sensitive environments
  • Working knowledge of basic network concepts, such as DNS, DHCP, VLANs, switching and routing
  • Basic working knowledge of Kubernetes and Openstack technologies (preferred but not required)
  • Experience with data center operations and process adherence
  • Excellent communication and coordination skills with cross-functional teams and external partners
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Compute Platform Engineer – HPC & GPU Systems
Compute Platform Engineer – HPC & GPU Systems

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 120,000 - 180,000
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 110,000 - 160,000
Software Engineer, Fleet Automation
Software Engineer, Fleet Automation

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 150,000
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Infrastructure Hardware Specialist
Infrastructure Hardware Specialist

Blue Signal Search • Austin (TX)

On-site
USD 90,000 - 130,000
Competitive base salary plus equity
Health/dental/vision coverage
Retirement benefits
+1
Infrastructure Hardware Specialist
Infrastructure Hardware Specialist

Blue Signal Search • Phoenix (AZ)

On-site
USD 90,000 - 130,000
Equity
Health/Dental/Vision
Retirement benefits
+2
Compute Engineer, Deployment
Compute Engineer, Deployment

Insight Global • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Infrastructure Engineer
Infrastructure Engineer

Calance • Illinois

On-site
USD 110,000 - 150,000