Senior Lab Reliability Engineer

VAST Data

New York (NY)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

VAST Data in New York State seeks a senior Lab Operations Engineer to own the reliability of our lab clusters. You will drive automation, tooling, and infrastructure-as-code practices, while serving as the escalation point for the toughest technical problems in the lab.

You will collaborate with a growing team across engineering, pre-sales, and services to validate solutions and keep the lab production-quality.

Qualifications

  • 4+ years in systems engineering, storage engineering or related role.
  • Hands-on with enterprise storage systems at operational level.
  • Strong Linux admin skills: networking, storage, filesystems, systemd, CLI.
  • Programming/scripting: Python or Bash to build automation.
  • Operational experience with Docker and Kubernetes.
  • Infra-as-code or config management tools such as Ansible.
  • Solid networking foundations and troubleshooting ability.
  • Experience with virtualization platforms (VMware vSphere, ESXi, Proxmox).
  • Methodical troubleshooting across storage, network, compute layers.

Responsibilities

  • Own operational reliability of VAST clusters in the lab; monitor health and plan upgrades.
  • Serve as primary tech escalation for complex cluster issues; collaborate with engineering.
  • Shape automation and tooling strategy for the lab, including provisioning and dashboards.
  • Establish standards for infrastructure-as-code and config management in the lab.
  • Reproduce issues in controlled labs and provide diagnostic data for engineering.
  • Own lab infrastructure including virtualization, compute, networking, and storage.
  • Mentor and support other lab operations team members.
  • Collaborate with pre-sales, services, and engineering to validate lab scenarios.

Skills

Systems engineering
Storage engineering
Customer support engineering
Linux administration
Python
Bash
Docker
Kubernetes
Ansible
VMware vSphere
Proxmox
Networking

Tools

Docker
Kubernetes
Ansible
VMware vSphere
Proxmox
Grafana
Prometheus

Job description

Responsibilities
  • Own the operational reliability of VAST clusters in the lab environment, including proactive health monitoring, upgrade planning, and issue resolution
  • Serve as the primary technical escalation point for complex cluster issues, working hands‑on-keyboard to resolve them and partnering with VAST engineering when deeper investigation is needed
  • Shape the automation and tooling strategy for the lab environment, including provisioning scripts, CLI utilities, monitoring dashboards, and internal tooling that the rest of the team builds on
  • Help establish standards for infrastructure‑as‑code and configuration management (Ansible or equivalent) across the lab environment
  • Reproduce and isolate difficult issues in controlled lab environments, producing high‑quality diagnostic data and reports for engineering
  • Own broader systems infrastructure supporting the lab, including virtualization platforms (VMware vSphere, Proxmox), compute, networking, and storage
  • Mentor and support other members of the lab operations team, raising the collective technical bar
  • Partner with pre‑sales SEs, professional services, and engineering to reproduce customer‑relevant scenarios and validate solutions in the lab
Required Qualifications
  • 4+ years of professional experience in a systems engineering, storage engineering, customer support engineering, or related role
  • Deep hands‑on experience with enterprise storage systems (VAST, Pure, NetApp, Isilon, Ceph, or similar) at an operational or reliability level
  • Strong Linux systems administration skills, including networking, storage, filesystems, systemd, and CLI tooling
  • Strong scripting/programming experience (Python, Bash, or similar), sufficient to build and maintain automation and tooling other engineers rely on
  • Experience with Docker and Kubernetes in operational environments
  • Experience with infrastructure‑as‑code or configuration management tools (Ansible or equivalent)
  • Solid networking foundations (VLANs, routing, subnetting, network troubleshooting)
  • Experience with virtualization platforms (VMware vSphere, ESXi, Proxmox, or equivalent)
  • Methodical approach to troubleshooting complex issues across storage, networking, and compute layers
  • Comfort working across timezones with distributed team members
  • Excellent written and verbal communication, including ability to produce clear technical documentation and reports
Preferred Qualifications
  • Existing hands‑on experience with VAST Data clusters
  • Prior experience as a Customer Support Engineer at a storage or infrastructure company OR Reliability Engineer / SRE supporting a wide variety of infrastructure services
  • Familiarity with observability platforms (Grafana, Prometheus, Elasticsearch)
  • Network switch administration experience (Cumulus, Arista, Cisco)
  • Experience with high‑performance computing environments or ML/AI training workloads

VAST is building out a Presales Platform team responsible for the tooling, automation, and infrastructure that enables our field engineering organization to demonstrate value at scale. Our lab is a core asset for evaluations, demos, and internal enablement, and we're building a team whose job is to keep it running at production‑quality reliability.

This role holds the senior technical ownership on that team. You'll own the reliability of the VAST clusters in our lab environment, influence the direction for the automation, tooling, and infrastructure‑as‑code practices the rest of the team builds on and executes within, and serve as the escalation point for the hardest technical problems in the lab. You'll partner with a growing team of lab and platform engineers, as well as the internal teams who rely on these tools day to day, to understand their needs and raise the operational bar of everything we run.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Customer Support Engineer
Customer Support Engineer

VAST Data • New York (NY)

On-site
USD 110,000 - 190,000
Customer Support Engineer
Customer Support Engineer

VAST Data • San Francisco (CA)

On-site
USD 110,000 - 160,000
Customer Support Engineer
Customer Support Engineer

VAST Data • San Jose (CA)

On-site
USD 80,000 - 100,000
Customer Support Engineer
Customer Support Engineer

VAST Data • Lehi (UT)

On-site
USD 80,000 - 100,000
Lab Reliability Engineer: Automation & Platform Expert
Lab Reliability Engineer: Automation & Platform Expert

VAST Data • New York (NY)

On-site
USD 140,000 - 190,000
Professional Services Engineer
Professional Services Engineer

VAST Data • Raleigh (NC)

On-site
USD 90,000 - 120,000
Resident Customer Support Engineer (Onsite - Dallas, TX)
Resident Customer Support Engineer (Onsite - Dallas, TX)

Drive Capital • Town of Texas (WI)

On-site
USD 90,000 - 130,000
Customer Support Engineer
Customer Support Engineer

VAST Data • Seattle (WA)

On-site
USD 100,000 - 120,000
Senior Systems Engineer - Higher Ed and SLED - TOLA
Senior Systems Engineer - Higher Ed and SLED - TOLA

VAST Data • Austin (TX)

On-site
USD 100,000 - 130,000
Senior Systems Engineer
Senior Systems Engineer

VAST Data • New York (NY)

On-site
USD 100,000 - 130,000