Data Center & Lab Technician - AI Accelerator Infrastructure - Contract

Entrada Ventures

Santa Clara (CA)

On-site

USD 80,000 - 110,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

d-Matrix is seeking a Data Center & Lab Technician to maintain on-prem and colocation lab infrastructure, rack servers, and hardware bring-up for AI accelerator systems. This one-year contract emphasizes independent work, precise documentation, and collaboration with SRE and silicon validation teams.

Candidates should have 3+ years in data center operations, hands-on server experience, and DCIM/tool proficiency. Location is Santa Clara, CA.

Qualifications

  • 3+ years in data center operations, hardware validation, or systems technician roles.
  • Proven rack-and-stack experience with servers, PDUs, patch panels and cabling.
  • Hardware bring-up: BIOS/UEFI, firmware updates, OS install across Linux.
  • Linux CLI proficiency to run diagnostics and runbooks.
  • Asset management with DCIM or inventory tools and up-to-date records.
  • Strong written communication; document actions and write SOPs.
  • Ability to work independently in a fast-moving startup environment.

Responsibilities

  • Rack, stack, cable, and decommission servers, PDUs, network gear, and storage in on-premises labs and colocation facilities.
  • Execute hardware bring-up for d-Matrix AI-accelerated systems and validation test benches, including BIOS configuration, firmware validation, and OS installation (Linux primary).
  • Replace and upgrade components: PCIe cards, NICs, storage, memory, and accelerator hardware across multiple server generations.
  • Maintain lab spaces in organized, ESD-compliant, and audit-ready condition at all times.
  • Own accurate asset tracking and inventory records using DCIM tools (NetBox, Jira, or equivalent).
  • Coordinate equipment moves between lab, staging, and colocation; manage shipping, receiving, and RMA workflows.
  • Track hardware lifecycle status: warranty, EOL, refresh schedules, and spare parts inventory.
  • Diagnose and resolve hardware failures, power and thermal issues, and network connectivity problems.
  • Serve as the first physical responder for infrastructure incidents requiring hands-on intervention.
  • Document all troubleshooting actions in ticketing/knowledge bases; contribute to runbooks.
  • Write and maintain SOPs, rack diagrams, and cabling guides for new technicians.
  • Communicate issues and status across hardware, software, SRE, and silicon validation teams.
  • Support parallel workstreams as silicon development programs evolve; operate independently when needed.

Skills

Rack-and-stack
Linux command-line
Documentation
Independent work
Written communication

Tools

NetBox
Jira

Job description

Data Center & Lab Technician - AI Accelerator Infrastructure - Contract

As a DC & lab technician, you are the hands and feet of the unified infrastructure team: racking servers, running cables, executing hardware bring-ups, and keeping lab environments in the precise, audit-ready state that high-velocity silicon and software development demands.

About the Role

This is a one-year contract with potential for a full-time conversion. This is an ownership role, not a ticket executor role. You operate independently, document what you build, and escalating with precision when something needs engineering attention.

What You Will Do
Physical Infrastructure & Hardware Bring-Up
  • Rack, stack, cable, and decommission servers, PDUs, network gear, and storage in on-premises labs and colocation facilities.

  • Execute hardware bring-up for d-Matrix AI-accelerated systems and validation test benches, including BIOS configuration, firmware validation, and OS installation (Linux primary).

  • Replace and upgrade components: PCIe cards, NICs, storage, memory, and accelerator hardware across multiple server generations.

  • Maintain lab spaces in organized, ESD-compliant, and audit-ready condition at all times.

Asset Management & Inventory
  • Own accurate asset tracking and inventory records using DCIM tools (NetBox, Jira, or equivalent); every piece of hardware is accounted for with the current configuration state.

  • Coordinate equipment moves between lab, staging, and colocation; manage shipping, receiving, and RMA workflows.

  • Track hardware lifecycle status: warranty, EOL, refresh schedules, and spare parts inventory.

Troubleshooting & Incident Support
  • Diagnose and resolve hardware failures, power and thermal issues, and network connectivity problems — escalating to SRE engineers with clear documentation and reproduction steps.

  • Serve as the first physical responder for infrastructure incidents requiring hands‑on intervention in lab or colo environments.

  • Document all troubleshooting actions in the team’s ticketing and knowledge base systems, contributing to runbooks that reduce repeat escalations.

Documentation & Collaboration
  • Write and maintain SOPs, rack diagrams, and cabling guides clear enough for a new technician to execute without shadowing.

  • Communicate issues and status across hardware, software, SRE, and silicon validation teams professionally and with precision.

  • Support parallel workstreams as silicon development programs evolve; operate independently when SREs are heads‑down on engineering work.

What You Will Bring
Minimum Qualifications
  • 3+ years in data center operations, hardware validation, or systems technician roles, hands‑on with real servers.

  • Proven rack‑and‑stack experience: physical server installation, PDU and patch panel cabling, rack power planning, and cable management to professional standards.

  • Hardware bring‑up experience: BIOS/UEFI configuration, firmware updates, component replacement, and OS installation across Linux distributions.

  • Linux command‑line proficiency: enough to run diagnostics, inspect logs, and execute runbooks independently.

  • Asset management discipline: experience with DCIM or inventory tools and a track record of accurate, up‑to‑date records.

  • Strong written communication: you document what you do, escrow with context, and write SOPs others can follow.

  • Comfortable operating independently in a fast‑moving startup environment.

Preferred Qualifications
  • Prior colocation data center experience (Equinix, CoreSite, Digital Realty, or similar).

  • Experience with AI/ML or GPU hardware: NVIDIA, AMD, or Intel accelerator cards in a lab or production environment.

  • Basic Ansible exposure — able to run playbooks and understand what they do.

  • Familiarity with high‑speed interconnects: InfiniBand, RoCE, SFP/QSFP optics.

  • Python or Bash scripting for operational task automation.

Why This Role

You will be a critical part of the infrastructure that powers d-Matrix’s AI hardware development programs. In a small, high‑ownership team, your work is immediately visible; when you bring up a system cleanly, the engineers depending on it notice. If you take pride in physical infrastructure done right and want to work at the cutting edge of AI silicon development, this is the role for you.

Equal Opportunity Employment Policy

d-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - AI Accelerator Infrastructure - Contract
Site Reliability Engineer - AI Accelerator Infrastructure - Contract

d-Matrix • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 270,000
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
Site Reliability Engineer - AI Accelerator Infrastructure - Contract

Entrada Ventures • Santa Clara (CA)

On-site
USD 120,000 - 180,000
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract

d-Matrix • Santa Clara (CA)

On-site
USD 250,000 - 350,000
AI Hardware Lab & Data Center Technician (1-Year Contract)
AI Hardware Lab & Data Center Technician (1-Year Contract)

Entrada Ventures • Santa Clara (CA)

On-site
USD 80,000 - 110,000
Hardware Qualification Engineer, Senior Staff
Hardware Qualification Engineer, Senior Staff

d-Matrix inc. • Santa Clara (CA)

Hybrid
USD 130,000 - 160,000
Data Center Technician L2 - 3 Openings
Data Center Technician L2 - 3 Openings

Covestic Inc • Hubbard (TX)

On-site
USD 60,000 - 90,000
Senior Runtime Systems Engineer
Senior Runtime Systems Engineer

Entrada Ventures • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Data Center Technician L3
Data Center Technician L3

Covestic Inc • Town of Texas (WI), Northern (KY)

On-site
USD 90,000 - 130,000
Principal Software Engineer, Kernels
Principal Software Engineer, Kernels

D-Matrix Corp. • Santa Clara (CA), Northern (KY)

Hybrid
USD 250,000 - 350,000