Remote GPU Infra NOC Engineer — Automation & AI Ops

Orionplacement

Pittsburgh (Allegheny County)

On-site

USD 75,000 - 140,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Bonus and equity opportunities
Medical, dental, and vision insurance
401(k)

Job summary

Orionplacement is seeking a dedicated NOC analyst to monitor and support cutting-edge GPU and AI infrastructure powering enterprise workloads. This fully remote role offers 24/7 rotating shifts, opportunities for automation, and hands-on exposure to GPU/HPC, networking, and data center operations.

You'll triage incidents, build tooling with Python or Bash, and collaborate with OEMs and data center partners to maintain SLA commitments and improve reliability across global deployments.

Qualifications

  • 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment.
  • Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments.
  • Hands-on Python or Bash scripting experience.
  • Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms.
  • Strong troubleshooting and incident-response skills.
  • Experience working with network, compute, storage, or data center infrastructure.
  • Ability to understand technical issues quickly and communicate effectively during incidents.
  • Demonstrated interest in automation, scripting, tool-building, and continuous operational improvement.
  • Must be comfortable working rotating 24/7 shifts, including nights and weekends.

Responsibilities

  • Monitor live GPU cluster health, power, cooling, networking, and infrastructure status across production deployments.
  • Triage, troubleshoot, and resolve incidents while maintaining SLA requirements.
  • Identify infrastructure issues early and take proactive action before they become customer-impacting incidents.
  • Escalate appropriate issues to OEMs, data center operators, network providers, or other Tier 3 partners.
  • Build and improve internal monitoring tools, scripts, and automation to reduce repetitive manual work.
  • Use Python, Bash, or similar scripting tools to automate monitoring, triage, reporting, and operational workflows.
  • Explore and implement AI-enabled workflows that improve NOC speed, accuracy, and efficiency.
  • Create, maintain, and continuously improve technical runbooks and standard operating procedures.
  • Participate in incident reviews and turn recurring problems into permanent tooling, process, or automation improvements.
  • Track SLA and incident metrics and identify opportunities to improve reliability and response times.
  • Communicate clearly and proactively with customers and internal stakeholders during incidents.
  • Coordinate with data center operators, OEMs, network providers, and other third parties to resolve customer-impacting issues.
  • Support a 24/7 rotating operations schedule, including nights, weekends, and other assigned shifts.

Skills

NOC experience
GPU/HPC infra
Python or Bash scripting
Monitoring tools

Tools

Datadog
Grafana
PagerDuty

Job description

Orionplacement is seeking a dedicated NOC analyst to monitor and support cutting-edge GPU and AI infrastructure powering enterprise workloads. This fully remote role offers 24/7 rotating shifts, opportunities for automation, and hands-on exposure to GPU/HPC, networking, and data center operations.

You'll triage incidents, build tooling with Python or Bash, and collaborate with OEMs and data center partners to maintain SLA commitments and improve reliability across global deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote GPU Infra NOC Engineer | Automation & AI Ops
Remote GPU Infra NOC Engineer | Automation & AI Ops

Orion Placement • Pittsburgh

On-site
USD 75,000 - 140,000
Dental insurance
Paid time off
Retirement plan
+2
NOC Engineer
NOC Engineer

Axe Compute • Miami (FL)

On-site
USD 65,000 - 90,000
Remote GPU Infrastructure Deployment Program Lead
Remote GPU Infrastructure Deployment Program Lead

Orionplacement • Pittsburgh

On-site
USD 120,000 - 200,000
401(k)
Dental insurance
Paid time off
+2
NOC Engineer - AI-Driven GPU Incident Response
NOC Engineer - AI-Driven GPU Incident Response

Axe Compute • Miami (FL)

On-site
USD 65,000 - 90,000
Remote GPU Infra Deployment Program Lead
Remote GPU Infra Deployment Program Lead

Orion Placement • Pittsburgh

On-site
USD 120,000 - 200,000
401(k)
Dental insurance
Paid time off
+4
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
Network Operations Center Technician II
Network Operations Center Technician II

Cirrascale Corporation • Austin (TX)

On-site
USD 55,000 - 90,000
Remote AI Platform Engineer — Kubernetes & GPU, Equity
Remote AI Platform Engineer — Kubernetes & GPU, Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 250,000 - 300,000
Meaningful equity
Fully remote across North America
Full insurance coverage for you and你的依
Remote AI Infrastructure Engineer: GPU Clusters & MLOps
Remote AI Infrastructure Engineer: GPU Clusters & MLOps

Bonfirevc • United States

On-site
USD 120,000 - 150,000
Competitive salary
Stock options
Health/dental/vision insurance
+2