DevOps Engineer - Generative AI & Enterprise Web Platforms

Hewlett Packard Enterprise Company

Friday Harbor (WA)

Hybrid

USD 140,000 - 190,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Hewlett Packard Enterprise is seeking a hands-on DevOps Engineer for Generative AI and Enterprise Web Platforms. The role blends CI/CD, container orchestration, and cloud operations to deliver secure, scalable platforms for AI-powered enterprise apps.

You will collaborate with many teams to automate operations, monitor performance, and improve release reliability across on-prem and AWS environments, with hybrid work and an office presence required.

Qualifications

  • Bachelor's degree in Computer Science, Machine Learning, Artificial Intelligence, or a related discipline.
  • 4–6 years of relevant experience in DevOps, Site Reliability Engineering, platform engineering, cloud operations, or production engineering.

Responsibilities

  • CI/CD and release engineering: Design, maintain , and improve automated build, test, security-scan, deployment, rollback, and release pipelines across development and production environments.
  • Container platforms: Operate Docker-based workloads and Kubernetes or k3s clusters, including on-premises developer sandboxes, configuration, upgrades, capacity, and troubleshooting.
  • Cloud and infrastructure automation: Provision and manage secure AWS infrastructure using repeatable automation, sound identity and access controls, networking, storage, compute, and environment configuration practices.
  • Observability and reliability: Implement metrics, logs, traces, dashboards, alerts, service-level indicators, and operational runbooks; support incident response, root-cause analysis, capacity planning, and resilience improvements.
  • Automation and validation: Build reusable Python and pytest utilities and integrate API, performance, reliability, security, and end-to-end validation into CI/CD workflows.
  • GenAI platform operations: Support deployment, tracing, regression checks, safety controls, latency monitoring, failure recovery, and cost visibility for LLM and agent-based services using tools such as LangSmith , LangGraph , Langfuse , and MCP.
  • Collaboration and continuous improvement: Partner with engineering, quality, security, and networking teams, participate in design and operational reviews, mentor less-experienced engineers, and drive practical improvements to platform standards and developer experience.

Skills

CI/CD
Python
Git workflows
Automation
Container orchestration
AWS
Observability
APIs testing
Networking basics
LangSmith
LangGraph
Langfuse
MCP

Education

Bachelor's degree in Computer Science/ML/AI or related field

Tools

Docker
Kubernetes
k3s
Datadog
Postman
JMeter
SoapUI

Job description

DevOps Engineer - Generative AI & Enterprise Web Platforms

This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today's complex world. Our culture thrives on finding new and better ways to accelerate what's next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description
Job Family Definition

Designs, develops, troubleshoots and debugs software programs for software enhancements and new products. Develops software including operating systems, compilers, routers, networks, utilities, databases and Internet-related tools. Determines hardware compatibility and/or influences hardware design.

Management Level Definition

Contributions have visible technical impact on a product or major subcomponent. Applies in-depth professional knowledge and innovative ideas to solve complex problems. Visible contributions improve time-to-market, achieve cost reductions, or satisfy current and future unmet customer needs. Recognized internal authority on key technology area applying innovative principles and ideas. Provides technical leadership for significant project/program work. Leads or participates in cross-functional initiatives and contributes to mentorship and knowledge sharing across the organization.

Role Overview

We are seeking a hands-on DevOps Engineer with experience to build, automate, and operate secure, reliable delivery platforms for enterprise applications powered by Generative AI, large language models, APIs, distributed systems, and modern web technologies.

The role combines CI/CD engineering, container orchestration, cloud operations, infrastructure automation, observability, security, performance, and production support. The successful candidate will collaborate with development, quality, networking, and platform teams to improve release reliability, troubleshoot cross-layer issues, automate repeatable operations, and strengthen service resilience across on-premises and AWS environments.

Responsibilities
  • CI/CD and release engineering: Design, maintain , and improve automated build, test, security-scan, deployment, rollback, and release pipelines across development and production environments.
  • Container platforms: Operate Docker-based workloads and Kubernetes or k3s clusters, including on-premises developer sandboxes, configuration, upgrades, capacity, and troubleshooting.
  • Cloud and infrastructure automation: Provision and manage secure AWS infrastructure using repeatable automation, sound identity and access controls, networking, storage, compute, and environment configuration practices.
  • Observability and reliability: Implement metrics, logs, traces, dashboards, alerts, service-level indicators, and operational runbooks; support incident response, root-cause analysis, capacity planning, and resilience improvements.
  • Automation and validation: Build reusable Python and pytest utilities and integrate API, performance, reliability, security, and end-to-end validation into CI/CD workflows.
  • GenAI platform operations: Support deployment, tracing, regression checks, safety controls, latency monitoring, failure recovery, and cost visibility for LLM and agent-based services using tools such as LangSmith , LangGraph , Langfuse , and MCP.
  • Collaboration and continuous improvement: Partner with engineering, quality, security, and networking teams, participate in design and operational reviews, mentor less-experienced engineers, and drive practical improvements to platform standards and developer experience.
Education and Experience Required
  • Bachelor's degree in Computer Science, Machine Learning, Artificial Intelligence, or a related discipline.
  • 4-6 years of relevant experience in DevOps, Site Reliability Engineering, platform engineering, cloud operations, or production engineering
Knowledge and Skills
  • CI/CD and automation: Strong hands-on experience with pipeline engineering, Git-based workflows, Python, pytest , scripting, artifact management, automated testing, diagnostics, and deployment automation.
  • Containers and orchestration: Production experience with Docker, Kubernetes, and k3s, including cluster operations, workload deployment, configuration, scaling, upgrades, and troubleshooting.
  • AWS and infrastructure: Hands-on knowledge of AWS services, identity and access management, compute, storage, networking, monitoring, infrastructure automation, security, and cost-aware operations.
  • Observability and operations: Experience with production monitoring, logging, tracing, alerting, dashboards, incident response, root-cause analysis, high availability, disaster recovery, and tools such as Datadog.
  • APIs and performance: Experience validating REST or gRPC services and using tools such as JMeter, SoapUI, and Postman for functional, integration, performance, scale, reliability, and security testing.
  • Networking and protocols: working knowledge of networking protocols like BGP, MPLS, EVPN, VXLAN, NETCONF, RESTCONF, gRPC , SNMP, LLDP and traffic tools such as Ixia or Spirent.
  • GenAI operations: Familiarity with LLM and agent evaluation, groundedness , hallucination and safety checks, regression metrics, tracing, human review, LangSmith , LangGraph , Langfuse , and MCP.
  • Problem-solving and collaboration: Ability to analyse architecture, diagnose cross-layer failures, assess operational risk, communicate root causes clearly, collaborate across teams, and guide junior engineers; bachelor's degree or equivalent practical experience.
  • Platform and infrastructure: Experience with infrastructure as code, GitOps , secrets management, and multi-environment or hybrid-cloud operations.
  • AI quality and observability: Experience with agent evaluation, prompt regression, model comparison, tracing, production monitoring, and operational controls for AI services.
  • Performance and resilience: Proficiency in API and browser performance testing, cloud-scale resilience, chaos testing, failover validation, and disaster-recovery exercises.
  • Security and complex platforms: Knowledge of OWASP risks, threat modelling, vulnerability management, adversarial validation, and multi-tenant or distributed enterprise architectures.
  • Credentials: Relevant AWS, Kubernetes, DevOps, security, SRE, or networking certifications.
What We Can Offer You:
Health & Wellbeing
  • We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.
Personal & Professional Development
  • We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have - whether you want to become a knowledge expert in your field or apply your skills to another division.
Unconditional Inclusion
  • We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.
Job:

Engineering

Job Level:

TCP_03

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer.

We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need.

Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together.

Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer – Generative AI & Enterprise Web Platforms
DevOps Engineer – Generative AI & Enterprise Web Platforms

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 120,000 - 150,000
DevOps Engineer - Generative AI & Enterprise Web Platforms
DevOps Engineer - Generative AI & Enterprise Web Platforms

Hewlett Packard Enterprise Company in • San Juan (PR)

Hybrid
USD 100,000 - 150,000
DevOps Engineer – Generative AI & Enterprise Web Platforms
DevOps Engineer – Generative AI & Enterprise Web Platforms

Hewlett Packard Enterprise Development LP • San Juan (PR)

On-site
USD 110,000 - 170,000
Software Engineer II – Generative AI & LLM Platforms
Software Engineer II – Generative AI & LLM Platforms

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 70,000 - 95,000
Health & Wellbeing
Professional Development
Unconditional Inclusion
Software Engineer II - Generative AI & LLM Platforms
Software Engineer II - Generative AI & LLM Platforms

Hewlett Packard Enterprise Company • Friday Harbor (WA)

Hybrid
USD 90,000 - 130,000
Principal Engineer – Generative AI & LLM Platforms
Principal Engineer – Generative AI & LLM Platforms

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 180,000 - 280,000
Software Engineer II - Generative AI & LLM Platforms
Software Engineer II - Generative AI & LLM Platforms

Hewlett Packard Enterprise Company in • San Juan (PR)

Hybrid
USD 90,000 - 120,000
Health & Wellbeing programs
Career development support
Inclusive culture and teams
Senior Quality Engineer - Generative AI & Enterprise Web Platforms
Senior Quality Engineer - Generative AI & Enterprise Web Platforms

Hewlett Packard Enterprise Company • Friday Harbor (WA)

Hybrid
USD 140,000 - 190,000
Full-Stack AI Engineer
Full-Stack AI Engineer

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 120,000 - 160,000
Health benefits
Software Engineering Manager – Generative AI & Enterprise Platforms
Software Engineering Manager – Generative AI & Enterprise Platforms

Hewlett Packard Enterprise Development LP • San Juan (PR)

On-site
USD 190,000 - 260,000
Health & Wellbeing
Career Development Programs
Unconditional Inclusion