Infrastructure Ops Specialist

Massed Compute

Northern (KY)

Hybrid

USD 90,000 - 130,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote stipend
Flexible PTO policy
Medical, dental, vision insurance for
401(k) retirement plan
Parental leave

Job summary

Massed Compute is building a modern GPU cloud platform for AI workloads. We seek an Infrastructure Operations Specialist to optimize asset placement, tracking from PO to production, and the warranty/RMA lifecycle while collaborating with procurement, data center leasing, and engineering teams.

You will design repeatable processes, build tooling, and drive data-backed decisions to scale deployments across multiple sites and partners in a remote, fast-paced environment.

Qualifications

  • Experience in supply chain management, analytics, or logistics.
  • Working fluency in data center fundamentals — racks, power, cooling, cabling.
  • Experience coordinating work across organizations that do not report to you.
  • The instinct to build a repeatable process the second time you do something.
  • Strong personal organization and comfort holding many parallel threads.
  • Strong written communication, especially in making status legible to non‑technical audiences.
  • High ownership and follow‑through; you stay with a deployment through acceptance.
  • Fluency with AI tools and related workflow automation tools.

Responsibilities

  • Decide where equipment goes across our site portfolio.
  • Track everything from PO to production and know where things are at all times.
  • Run deployments and redeployments across facilities and teams.
  • Own warranty, RMA, and servicing processes to resolution.
  • Own the fleet system of record and ensure accurate reconciliation.
  • Build the forward view and reporting for capacity and risk.
  • Support solutions engineering with truth on commitments and timelines.
  • Manage vendor and remote hands relationships and SLAs.

Skills

Supply chain management
Data center basics
Cross-functional collaboration
Process improvement
Organizational skills
Written communication
Ownership
AI tooling

Education

Bachelor's degree or equivalent

Job description

Build the GPU cloud that AI teams actually want to use .

Think it. Build it. Ship it.

We own and run our own NVIDIA infrastructure, so the people who build here work close to the hardware and close to the customer. That is the promise we make to customers, and it is how the work happens here.

Infrastructure Ops Specialist
Location

Remote - United States

Employment Type

Full time

Location Type

Remote

Department

Strategy & Operations

About Massed Compute

Massed Compute is building a modern GPU cloud platform for AI and high-performance compute workloads. Customers choose us when they need high-performance infrastructure with more flexibility, visibility, and support than they get from traditional options. Our work sits at the intersection of AI infrastructure, product design, systems thinking, and customer obsession.

About the role

Massed Compute is looking for an Infrastructure Operations Specialist to make sure every GPU we buy ends up in the right place, in service, and earning — and that we always know where all of it is

You’ll be the connective tissue between procurement, data center leasing, site operations, and solutions engineering. Those four functions each hold part of the picture: what we’ve bought, where we have space and power, what is actually racked and running, and what we’ve promised customers. That means deciding where new equipment goes across our site portfolio, tracking every shipment from purchase order to production, running deployments and redeployments, driving warranty and RMA work to resolution, and building the process that tells us what state our infrastructure is in now and what it will be in ninety days.

This role involves extensive logistics and coordination that will directly drive utilization and operational efficiency, and transparency that informs decisions and prioritization for multiple teams across the company. Additionally, you’ll define what has to be tracked and how, and partner with engineering and use coding agents to build out the tooling behind the monitoring and analytics.

This is a builder role. We’re looking for someone who can both run today’s deployments personally and build the tracking, process, and vendor discipline that let us run ten times as many without ten times the people.

What success looks like

In your first 8–12 months, you will have built a single, trusted view of what hardware we own and what state it is in, cut the time between equipment arriving and equipment earning, established a warranty and RMA process that resolves failures without anyone chasing them, and given procurement, capacity, and the commercial team a forward view of the fleet they can plan against.

Critically, you will have replaced a picture that lives in a handful of people’s heads and a dozen spreadsheets with a system that stays true on its own — so that "where is it, what state is it in, and when is it available" stops being a question anyone has to ask around to answer.

Responsibilities

Decide where equipment goes. Own placement across our site portfolio, balancing power and cooling headroom, rack density, customer adjacency, deployment speed, and what each site should be used for over the next year.

Track everything from PO to production. Maintain the chain of custody for every shipment — ordered, in transit, received, staged, racked, cabled, burned in, accepted — and know at any moment what is where and what is late.

Run deployments and redeployments. Coordinate the sequence across facilities, networking, storage, and platform teams, and move capacity between sites when demand or economics change.

Own warranty, RMA, and servicing. Drive hardware failures through vendor and depot processes to resolution, track what we are owed, and hold vendors to their commitments.

Own the fleet system of record. Select, implement, and maintain the asset and inventory system that tells us what we own, where it is, and what state it is in — including the reconciliation discipline that keeps it accurate

Build the forward view. Design the reporting that shows what is landing when, what capacity frees up when, and where we are at risk so leadership and the commercial team can plan against reality rather than intent

Support solutions engineering. Be the source of truth on what we can credibly commit to a customer and by when, and flag early when a promise is at risk

Manage vendors and remote hands. Own the relationships, SLAs, and escalation paths with the integrators, remote-hands providers, and logistics partners we depend on

Close the loop on failures. When a deployment slips or hardware goes missing, find the actual cause and change the process, rather than absorbing it and moving on

Bring leverage through AI. Use AI tools aggressively to automate tracking, reporting, vendor follow-up, and documentation — this role has more repeatable, automatable surface area than almost any other on the team

Qualifications

You may be a good fit if you have…

Experience in supply chain management, analytics, or logistics, and a track record of tracking physical assets accurately at scale, and specific opinions about where tracking breaks down

Working fluency in data center fundamentals — racks, power, cooling, cabling, density — enough to make placement decisions and be credible with site teams

Experience coordinating work across organizations that do not report to you, including vendors who have their own priorities

The instinct to build a repeatable process the second time you do something, rather than the tenth

Strong personal organization and comfort holding many parallel threads without dropping any, and a system for it that is not memory

Strong written communication, especially in making status legible to people who do not want a status meeting

High ownership and follow-through; you stay with a deployment through acceptance rather than handing it off at delivery

Fluency with AI tools and excitement about using them to automate the repetitive parts of operations work

Desired background (we encourage you to apply even if you don't check every box)

Typically 3–8+ years of professional experience in data center operations, infrastructure deployment, hardware or IT asset management, supply chain and logistics operations, technical program management, or a similarly high-responsibility role

Experience at a high-growth technology company, rapidly scaling startup, infrastructure operator, or similarly demanding environment

Hands‑on experience with hardware deployment lifecycles — receiving, installation, commissioning, failure, RMA, redeployment, decommissioning

Experience implementing or materially rebuilding an asset management, inventory, or DCIM system that other teams then relied on

A bachelor's degree or equivalent combination of education, training, and experience

Have worked with GPU infrastructure specifically, and understand what is different about it — density, cooling, failure modes, RMA timelines, and how long a dead node actually costs you

Have run deployments across multiple sites or colocation providers simultaneously

Have managed warranty and RMA processes at volume with OEMs, integrators, or resellers

Have built lightweight internal tooling or automation yourself to make an operations process work without more headcount

Are fluent with AI tools such as Claude, Claude Code, and Codex, and related workflow automation tools

We are a lean, ambitious team operating in one of the most important technology markets in the world. If you want ownership, pace, and the opportunity to do career-defining work alongside practical, high-agency teammates, MassedCompute is a great place to build.

Market competitive compensation

  • We are a remote-first company. You’ll get a stipend to dial in your setup and additional opportunities to connect with your peers during periodic in person team gatherings
  • Flexible PTO policy with the expectation that you take the time you need to recharge
  • 100% coverage of medical, dental, and vision insurance for employees, 30% of premiums for dependents, plus a tax-advantaged Flexible Spending Account
  • Company-facilitated 401(k) retirement plan
  • Parental leave for all new parents

Massed Compute is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, veteran status, or any other characteristic protected by law.

Why the work matters
AI teams need compute today, and we give it to them.

Most GPU providers ask for a long contract, a support queue, and a bill that takes a spreadsheet to understand. We built the friendlier option. On demand GPUs by the hour, bare metal for production workloads, and full clusters for serious training runs, all on hardware we own and run ourselves.

Real infrastructure

We are an NVIDIA Preferred Partner running current data center GPUs in Tier III facilities, with SOC 2, GDPR, and HIPAA compliance behind the platform.

Real customers

Researchers, studios, startups, and enterprise teams run production work on Massed Compute every day. What you build reaches them the week you ship it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Business Operations Manager
Business Operations Manager

Massed Compute • Northern (KY)

Hybrid
USD 120,000 - 210,000
Remote-first culture
Market-competitive compensation
Flexible PTO
+3
Founding GTM Leader
Founding GTM Leader

Massed Compute • Northern (KY)

Hybrid
USD 150,000 - 210,000
Remote-friendly startup culture
Setup stipend for home office
Flexible PTO policy
+1
Data Center Tech - Kansas City
Data Center Tech - Kansas City

Massed Compute • Kansas City (MO), Northern (KY)

Hybrid
USD 41,000 - 55,000
Solutions Engineer
Solutions Engineer

Massed Compute • United States

On-site
USD 110,000 - 180,000
Market competitive compensation
Remote-first with setup stipend
Flexible PTO
+3
Founding GTM Leader
Founding GTM Leader

Massed Compute • United States

On-site
USD 180,000 - 280,000
Remote-first company with stipend for
In-person team gatherings
Flexible PTO
+3
Infrastructure Engineer, Perimeter Compute
Infrastructure Engineer, Perimeter Compute

Montauk Capital • New York (NY)

On-site
USD 120,000 - 160,000
Competitive compensation
Equity options
Studio support from Montauk Capital
Data Center Tech - Columbus, IN
Data Center Tech - Columbus, IN

Massed Compute • Columbus (IN)

On-site
USD 34,000 - 55,000
Competitive hourly pay
Steady on-site work at data center
Remote GPU Infrastructure Operations Specialist
Remote GPU Infrastructure Operations Specialist

Massed Compute • Northern (KY)

Hybrid
USD 90,000 - 130,000
Remote stipend
Flexible PTO policy
Medical, dental, vision insurance for
+2
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime-Intellect • United States

On-site
USD 150,000 - 300,000
Equity incentives
Open frontier AI environment
Growth opportunities