HCI Sr. Compute Engineer (Red Hat OpenShift)

Stefanini EMEA

Town of Poland (NY)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Stefanini Group is seeking a Senior Compute Engineer specialised in Red Hat OpenShift to strengthen our Compute Operations team and provide Level 3 expert support for enterprise workloads on container and virtualization platforms.

This role focuses on day-to-day operations, stability, upgrades, patching, and platform modernization, including VMware-to-OpenShift transformations, with strong troubleshooting and automation expectations.

Qualifications

  • Strong hands-on OpenShift administration and operations.
  • Solid Linux background (RHEL), with troubleshooting, networking and storage.
  • Good understanding of Kubernetes fundamentals (pods, deployments, RBAC, etcd).
  • Experience in production environments with uptime and SLA commitments.

Responsibilities

  • Act as the L3 escalation point for complex issues in OpenShift clusters.
  • Own major incidents and perform root cause analysis.
  • Plan and execute OpenShift lifecycle activities (upgrades, patches).
  • Support VMware-to-OpenShift virtualization transformation projects.
  • Develop and maintain runbooks, SOPs, and automation to improve stability.

Skills

OpenShift
Linux
Kubernetes
SRE
Scripting

Tools

Ansible
GitOps (ArgoCD)
Python
Bash

Job description

Stefanini Group is seeking a Senior Compute Engineer specialised in Red Hat OpenShift to strengthen our Compute Operations team and provide Level 3 (L3) expert support for enterprise customers running critical workloads on container and virtualization platforms. This role is a key technical position focused on day-to-day operations, stability, and continuous improvement of OpenShift-based platforms. The engineer will act as the highest escalation point for complex incidents and problems, support platform lifecycle activities (upgrades, patching, performance tuning), and contribute to platform modernization initiatives – including VMware-to-OpenShift virtualization transformation programs. The ideal candidate combines strong troubleshooting skills, deep infrastructure understanding, and hands‑on OpenShift expertise, with the ability to work in a structured operational environment (ITIL/managed services), while also supporting automation and standardisation.

Job Responsibilities
Level 3 Operations & Technical Escalation (Core Responsibility)
  • Act as the L3 escalation point for complex technical issues related to:
  • Red Hat OpenShift clusters (control plane, worker nodes, networking, storage, authentication)
  • OpenShift Virtualization (KubeVirt) and VM-based workloads hosted on OpenShift
  • Linux OS level issues impacting cluster stability or workloads.
  • Own and drive resolution of:
  • Major Incidents (P1/P2) with deep technical investigation and rapid recovery focus
  • Recurring incidents through Problem Management (root cause analysis and permanent fixes).
  • Lead deep troubleshooting activities:
  • cluster degradation, node failures, API instability, etcd performance issues
  • networking issues (ingress, routes, DNS, CNI, service connectivity)
  • storage issues (persistent volumes, performance bottlenecks, CSI failures)
  • workload failures (pods, operators, deployments, stateful applications).
  • Provide clear technical updates during incidents, including impact assessment, recovery plan / workaround, risks and next steps.
Platform Lifecycle Management (Upgrades, Patching, Stability)
  • Plan and execute OpenShift lifecycle activities such as version upgrades (cluster upgrades and operator upgrades), patching and security hardening and certificate management and renewal processes.
  • Validate platform readiness before changes: capacity, compatibility, performance, known issues.
  • Maintain high availability and resilience: backup/restore strategy support (including etcd backup practices), disaster recovery readiness and operational runbooks.
  • Ensure operational compliance with defined maintenance windows and change governance.
VMware-to-OpenShift Virtualization Transformation Support
  • Support enterprise modernization initiatives involving migration from traditional virtualization platforms (VMware) to OpenShift Virtualization
  • Contribute to:
    • migration approach definition and technical design support
    • workload onboarding, validation, and stabilization on OpenShift
    • performance tuning and operational model definition for VM-based workloads on OpenShift.
  • Ensure production‑grade operational readiness: monitoring, alerting, backup, patching and support model aligned with managed services standards.
Standardization, Automation & Operational Improvement
  • Develop and maintain operational documentation, including troubleshooting guides, standard operating procedures (SOPs), build standards and reference architectures, operational runbooks for recurring tasks.
  • Support automation initiatives using tools such as: Ansible / Automation Platform (preferred), GitOps practices (ArgoCD) where applicable and scripting (Bash / Python) to reduce manual operations.
  • Proactively identify improvements to increase platform stability, recovery speed (MTT), repeatability and reduction of human error.
Monitoring, Observability & Performance Management
  • Support and improve observability across the platform, including:
    • OpenShift monitoring stack (Prometheus / Alertmanager / Grafana)
    • log management (e.g., EFK / Loki or enterprise logging platforms).
  • Troubleshoot performance issues related to compute resource constraints, scheduling and resource requests/limits and cluster scaling and capacity planning.
  • Work with customer stakeholders and internal teams to define alert thresholds, reduce noise and false positives and improve operational dashboards and health reporting.
Security & Compliance Support
  • Ensure the platform is operated in a secure manner aligned with enterprise expectations:
    • RBAC best practices
    • integration with enterprise identity providers (LDAP / AD / SSO)
    • secure cluster configuration and segregation
    • Support vulnerability remediation and platform hardening initiatives
    • Collaborate with Security teams for audits, compliance requests, and evidence collection.
Job Requirements
Mandatory Technical Skills
  • Strong hands‑on experience with Red Hat OpenShift administration and operations.
  • Strong Linux background (RHEL preferred), including troubleshooting OS performance, services, networking, and storage.
  • Solid understanding of Kubernetes fundamentals: pods, deployments, services, ingress, namespaces, RBAC, operators.
  • Experience troubleshooting infrastructure-related issues across compute, network, storage, and platform services.
  • Experience working in production environments with uptime and SLA commitments.
Mandatory Professional Skills
  • Proven ability to operate as Level 3 support, including deep troubleshooting, structured root cause analysis and ownership until resolution.
  • Ability to communicate clearly with customers (technical and non‑technical stakeholders) and internal teams (L1/L2/architects/project teams).
  • Strong documentation discipline and operational mindset.
Preferred/ Nice-to-Have Skills
  • Experience with OpenShift Virtualization (KubeVirt) and VM-based workloads.
  • Experience supporting VMware environments and understanding virtualization concepts: vSphere architecture, clusters, HA/DRS, storage/datastores, VM lifecycle.
  • Experience with automation tools:
    • Ansible / Red Hat Ansible Automation Platform
    • GitOps tools (ArgoCD)
    • Infrastructure as Code practices.
  • Experience with enterprise storage and CSI integrations.
  • Experience with enterprise networking topics (DNS, routing, firewall constraints, load balancing).
  • Experience with public cloud OpenShift deployments (optional): ROSA / ARO / OCP on AWS/Azure/GCP.
Certifications (Preferred)
  • Red Hat Certified Specialist in OpenShift Administration (preferred)
  • Red Hat Certified Engineer (RHCE) (strong advantage)
  • Kubernetes certifications (CKA/CKAD) (nice to have).
Working Model & Operational Expectations
  • Work in an operational environment following ITIL practices (Incident / Problem / Change Management) and managed services delivery model and SLA commitments.
  • Participate in on-call rotation, planned maintenance windows and technical escalation duty as required.
  • Provide clear handovers and updates to ensure continuity across shifts/regions.
Diversity & Inclusion

Here at the Stefanini Group, we value plurality and equity, regardless of race, sexual orientation, disability, age, ancestry, religion, gender, and nationality. We understand and encourage the importance of being you!

About Us

We are the Stefanini group, a global tech consulting company of Brazilian origin that believes in the power of people to transform businesses through technology.

We are present in over 40 countries and operate with the purpose of co‑creating solutions TOGETHER WITH OUR CLIENTS that accelerate results and improve the experience of people and organizations.

Here, we like to say that technology is not the end, but the means: what really matters are the people who drive it all.

Our mindset is AI First, meaning we invest in cutting‑edge technology in everything we do, focusing on results for our clients.

We are a company, A GROUP, that breathes collaboration and offers a dynamic environment where you will learn by doing, grow alongside the team, and have space to contribute with ideas and projects.

More than just talking about digital transformation, we believe in real transformation that starts with people and impacts real businesses.

If you are looking for a place to develop, innovate, and be part of something bigger, the Stefanini Group is your place.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

OpenShift DevOps Engineer
OpenShift DevOps Engineer

Capgemini • Charlotte (NC)

On-site
USD 110,000 - 130,000
Healthcare including dental and vision
401(k) and Employee Share Ownership Plan
Paid time off and holidays
+2
Principal Engineer Software
Principal Engineer Software

Palo Alto Networks • Santa Clara (CA)

On-site
USD 147,000 - 237,500
Competitive salary
Comprehensive health, dental, and vision insurance
401(k) retirement plan with company match
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Red River • Raleigh (NC)

On-site
USD 119,000 - 196,000
Comprehensive medical, dental, and vis
401(k) with employer match
Paid time off and holidays
+1
Openshift Engineer
Openshift Engineer

Cloud BC Labs • Dallas (TX)

Hybrid
USD 110,000 - 160,000
Specialist Solution Architect, Cloud Services
Specialist Solution Architect, Cloud Services

Socket.dev • North Carolina

Hybrid
USD 151,000 - 241,000
Medical coverage
401(k) match
PTO
+3
Infrastructure Architect with Openshift
Infrastructure Architect with Openshift

Crossvale • Fort Worth (TX), Arlington (TX), Dallas (TX)

Remote
USD 100,000 - 130,000
Competitive salary and OTE
100% remote work
15 days PTO + 8 paid holidays
+2
OpenShift Cloud Engineer #2792
OpenShift Cloud Engineer #2792

Genius Road, LLC • Dallas (TX)

On-site
USD 120,000 - 150,000
Competitive pay
Professional growth opportunities
Collaborative environment
Senior Openshift Engineer
Senior Openshift Engineer

Apollo Solutions • United States

On-site
USD 130,000 - 190,000
OpenShift Virtualization Senior Consultant - Top Secret Clearance Required
OpenShift Virtualization Senior Consultant - Top Secret Clearance Required

Red River • United States

On-site
USD 130,000 - 215,000
Comprehensive benefits
401(k) with employer match
Paid time off and holidays
+3
Site Reliability Engineer
Site Reliability Engineer

Stefanini, Inc • Dearborn (MI)

Hybrid
USD 84,000 - 91,000