Lab Reliability Engineer: Automation & Platform Expert

VAST Data

New York (NY)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

VAST Data in New York State seeks a senior Lab Operations Engineer to own the reliability of our lab clusters. You will drive automation, tooling, and infrastructure-as-code practices, while serving as the escalation point for the toughest technical problems in the lab.

You will collaborate with a growing team across engineering, pre-sales, and services to validate solutions and keep the lab production-quality.

Qualifications

  • 4+ years in systems engineering, storage engineering or related role.
  • Hands-on with enterprise storage systems at operational level.
  • Strong Linux admin skills: networking, storage, filesystems, systemd, CLI.
  • Programming/scripting: Python or Bash to build automation.
  • Operational experience with Docker and Kubernetes.
  • Infra-as-code or config management tools such as Ansible.
  • Solid networking foundations and troubleshooting ability.
  • Experience with virtualization platforms (VMware vSphere, ESXi, Proxmox).
  • Methodical troubleshooting across storage, network, compute layers.

Responsibilities

  • Own operational reliability of VAST clusters in the lab; monitor health and plan upgrades.
  • Serve as primary tech escalation for complex cluster issues; collaborate with engineering.
  • Shape automation and tooling strategy for the lab, including provisioning and dashboards.
  • Establish standards for infrastructure-as-code and config management in the lab.
  • Reproduce issues in controlled labs and provide diagnostic data for engineering.
  • Own lab infrastructure including virtualization, compute, networking, and storage.
  • Mentor and support other lab operations team members.
  • Collaborate with pre-sales, services, and engineering to validate lab scenarios.

Skills

Systems engineering
Storage engineering
Customer support engineering
Linux administration
Python
Bash
Docker
Kubernetes
Ansible
VMware vSphere
Proxmox
Networking

Tools

Docker
Kubernetes
Ansible
VMware vSphere
Proxmox
Grafana
Prometheus

Job description

VAST Data in New York State seeks a senior Lab Operations Engineer to own the reliability of our lab clusters. You will drive automation, tooling, and infrastructure-as-code practices, while serving as the escalation point for the toughest technical problems in the lab.

You will collaborate with a growing team across engineering, pre-sales, and services to validate solutions and keep the lab production-quality.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Lab Reliability Engineer
Senior Lab Reliability Engineer

VAST Data • New York (NY)

On-site
USD 140,000 - 190,000
Lab Ops Engineer — Build the Lab Operating System
Lab Ops Engineer — Build the Lab Operating System

Light Labs • Ann Arbor (MI)

On-site
USD 90,000 - 130,000
Lab Infrastructure & Automation Engineer
Lab Infrastructure & Automation Engineer

Insight Global • Austin (TX)

On-site
USD 100,289,000 - 120,347,000
Senior Lab Automation Engineer – Platform Architect
Senior Lab Automation Engineer – Platform Architect

100 Eli Lilly and Company • New York (NY)

On-site
USD 160,000 - 230,000
401(k)
Pension
Vacation benefits
+4
Senior Lab Automation Engineer: Reliability & Innovation
Senior Lab Automation Engineer: Reliability & Innovation

Guardant Health, Inc. • Redwood City (CA)

Hybrid
USD 127,000 - 175,000
Hybrid work model
Platform Reliability Lead: Production Systems Expert
Platform Reliability Lead: Production Systems Expert

Cielo Projects • New York (NY)

On-site
USD 170,000 - 220,000
DataOps & Reliability Engineer
DataOps & Reliability Engineer

Compunnel, Inc. • Erie

On-site
USD 110,000 - 160,000
Lead Site Reliability Engineer - Architect & Own Production
Lead Site Reliability Engineer - Architect & Own Production

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000
Lead Data Center Hardware Reliability Engineer
Lead Data Center Hardware Reliability Engineer

OpenAI • United States

Remote
USD 150,000 - 210,000
Associate Director, Service Reliability & Engineering
Associate Director, Service Reliability & Engineering

Natera • Austin (TX)

On-site
USD 180,000 - 240,000
Competitive benefits