Site Reliability Engineer (SRE)

Veriipro

Jersey City (NJ)

On-site

USD 130,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veriipro in Jersey City is seeking an experienced Site Reliability Engineer focused on Microsoft Hyper-V & Private Cloud to operate and continually improve private cloud and VDI platforms. You will apply SRE practices to boost reliability, scalability, and automation while partnering with security, networking, and application teams.

The role requires hands-on Hyper-V administration, strong automation via PowerShell and IaC, and a track record of incident response, RCA, and capacity planning.

Qualifications

  • 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Hyper-V administration.
  • Experience in mission-critical enterprise infrastructure with high availability.
  • Automation to reduce operational overhead and improve reliability.
  • Experience with private cloud and VDI environments.
  • Incident response, root cause analysis, and continuous service improvement.
  • Microsoft certifications such as Windows Server Hybrid Administrator Associate desirable.
  • Banking/financial services experience advantageous.

Responsibilities

  • Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
  • Optimize, and support highly available VDI environments on Hyper-V.
  • Improve platform reliability, availability, scalability, and resiliency using SRE principles.
  • Disaster recovery, backup, patch management, and business continuity strategies.
  • Define and maintain SLIs/SLOs and operational metrics.
  • Automate infrastructure provisioning and configuration management using PowerShell and IaC.
  • Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
  • Develop proactive monitoring, alerting, logging, and observability capabilities.
  • Lead incident response, perform root cause analysis (RCA), and implement preventive actions.
  • Perform capacity planning, performance tuning, and resource optimization.
  • Collaborate with Security, Networking, Platform Engineering, and Application teams.
  • Develop and maintain documentation and runbooks.
  • Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.

Skills

PowerShell
IaC
Observability
SRE practices
VDI platforms

Tools

Microsoft Hyper-V
Windows Server
SCOM
Prometheus/Grafana

Job description

Overview

We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.

The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.

Key Responsibilities
  • Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
  • Optimize, and support highly available VDI environments on Hyper-V.
  • Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
  • Disaster recovery, backup, patch management, and business continuity strategies.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
  • Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
  • Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
  • Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
  • Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
  • Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
  • Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
  • Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
  • Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
  • Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.
Experience & Qualifications
  • 6+ years of infrastructure engineering experience with at least 4+ years of hands‑on Microsoft Hyper‑V administration.
  • Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
  • Proven experience implementing automation to reduce operational overhead and improve service reliability.
  • Experience supporting enterprise private cloud and VDI environments.
  • Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
  • Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
  • Experience in Banking or Financial Services environments is advantageous.
Preferred Skills
  • Windows Server 2016/2019/2022 administration.
  • Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper‑V Replica.
  • Exposure to hybrid cloud and private cloud platforms.
  • Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
  • Experience supporting enterprise VDI environments.
  • Understanding of ITIL Incident, Problem, Change, and Release Management.
  • Experience working in regulated industries such as Banking or Financial Services.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) - Microsoft Hyper-V | Private Cloud | Azure
Senior Site Reliability Engineer (SRE) - Microsoft Hyper-V | Private Cloud | Azure

Veriipro • Jersey City (NJ)

On-site
USD 120,000 - 180,000
Senior Hyper‑V SRE for Private Cloud & VDI Reliability
Senior Hyper‑V SRE for Private Cloud & VDI Reliability

Veriipro • Jersey City (NJ)

On-site
USD 120,000 - 180,000
Hyper-V SRE: Private Cloud & VDI Reliability Lead
Hyper-V SRE: Private Cloud & VDI Reliability Lead

Veriipro • Jersey City (NJ)

On-site
USD 130,000 - 185,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Deerfield (IL)

On-site
USD 120,000 - 150,000
SRE/Platform Engineer
SRE/Platform Engineer

Stash Talent Services • Virginia (MN)

Remote
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mikealbert • Cincinnati (OH)

Hybrid
USD 100,000 - 130,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

TechDigital Group • Georgia

Hybrid
USD 140,000 - 180,000