DevOps Engineer – Generative AI & Enterprise Web Platforms

Hewlett Packard Enterprise

San Juan (PR)

Hybrid

USD 120,000 - 150,000

Full time

11 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Hewlett Packard Enterprise in Puerto Rico seeks a hands-on DevOps Engineer to design, build, and operate secure, scalable delivery platforms for enterprise apps powered by Generative AI and modern web technologies. The role blends CI/CD, container orchestration, cloud operations, and observability across on‑premises and AWS environments.

You will collaborate with development, quality, networking, and platform teams to automate repeatable operations, improve release reliability, and strengthen

Qualifications

  • Bachelor's degree in Computer Science, Machine Learning, AI, or related field.
  • 4–6 years of DevOps, SRE, platform or cloud operations experience.

Responsibilities

  • Design, maintain, and improve CI/CD pipelines across development and production environments.
  • Operate Docker-based workloads and Kubernetes/k3s clusters, including on‑prem developer sandboxes.
  • Provision and manage secure AWS infrastructure with IAM, networking, storage, compute, and monitoring.
  • Implement metrics, logs, traces, dashboards, alerts, and runbooks; support incident response and RCAs.
  • Build reusable Python utilities and integrate API, performance, reliability, and security validation into CI/CD workflows.
  • Support GenAI platform operations including deployment, tracing, and cost visibility for LLM/agent services.

Skills

CI/CD automation
Python
Docker & Kubernetes
AWS
Observability
Security awareness
Team collaboration

Education

Bachelor's degree in CS / ML / AI

Tools

Docker
Kubernetes
k3s
Datadog
Postman
JMeter

Job description

This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description

Job Family Definition: Designs, develops, troubleshoots and debugs software programs for software enhancements and new products. Develops software including operating systems, compilers, routers, networks, utilities, databases and Internet-related tools. Determines hardware compatibility and/or influences hardware design.

Management Level Definition: Contributions have visible technical impact on a product or major subcomponent. Applies in-depth professional knowledge and innovative ideas to solve complex problems. Visible contributions improve time-to-market, achieve cost reductions, or satisfy current and future unmet customer needs. Recognized internal authority on key technology area applying innovative principles and ideas. Provides technical leadership for significant project/program work. Leads or participates in cross-functional initiatives and contributes to mentorship and knowledge sharing across the organization.

Role Overview

We are seeking a hands‑on DevOps Engineer with experience to build, automate, and operate secure, reliable delivery platforms for enterprise applications powered by Generative AI, large language models, APIs, distributed systems, and modern web technologies.

The role combines CI/CD engineering, container orchestration, cloud operations, infrastructure automation, observability, security, performance, and production support. The successful candidate will collaborate with development, quality, networking, and platform teams to improve release reliability, troubleshoot cross-layer issues, automate repeatable operations, and strengthen service resilience across on‑premises and AWS environments.

Responsibilities
  • CI/CD and release engineering: Design, maintain, and improve automated build, test, security‑scan, deployment, rollback, and release pipelines across development and production environments.
  • Container platforms: Operate Docker‑based workloads and Kubernetes or k3s clusters, including on‑premises developer sandboxes, configuration, upgrades, capacity, and troubleshooting.
  • Cloud and infrastructure automation: Provision and manage secure AWS infrastructure using repeatable automation, sound identity and access controls, networking, storage, compute, and environment configuration practices.
  • Observability and reliability: Implement metrics, logs, traces, dashboards, alerts, service‑level indicators, and operational runbooks; support incident response, root‑cause analysis, capacity planning, and resilience improvements.
  • Automation and validation: Build reusable Python and pytest utilities and integrate API, performance, reliability, security, and end‑to‑end validation into CI/CD workflows.
  • GenAI platform operations: Support deployment, tracing, regression checks, safety controls, latency monitoring, failure recovery, and cost visibility for LLM and agent‑based services using tools such as LangSmith, LangGraph, Langfuse, and MCP.
  • Collaboration and continuous improvement: Partner with engineering, quality, security, and networking teams, participate in design and operational reviews, mentor less‑experienced engineers, and drive practical improvements to platform standards and developer experience.
Education And Experience Required
  • Bachelor's degree in Computer Science, Machine Learning, Artificial Intelligence, or a related discipline.
  • 4-6 years of relevant experience in DevOps, Site Reliability Engineering, platform engineering, cloud operations, or production engineering.
Knowledge And Skills
  • CI/CD and automation: Strong hands‑on experience with pipeline engineering, Git‑based workflows, Python, pytest, scripting, artifact management, automated testing, diagnostics, and deployment automation.
  • Containers and orchestration: Production experience with Docker, Kubernetes, and k3s, including cluster operations, workload deployment, configuration, scaling, upgrades, and troubleshooting.
  • AWS and infrastructure: Hands‑on knowledge of AWS services, identity and access management, compute, storage, networking, monitoring, infrastructure automation, security, and cost‑aware operations.
  • Observability and operations: Experience with production monitoring, logging, tracing, alerting, dashboards, incident response, root‑cause analysis, high availability, disaster recovery, and tools such as Datadog.
  • APIs and performance: Experience validating REST or gRPC services and using tools such as JMeter, SoapUI, and Postman for functional, integration, performance, scale, reliability, and security testing.
  • Networking and protocols: working knowledge of networking protocols like BGP, MPLS, EVPN, VXLAN, NETCONF, RESTCONF, gRPC, SNMP, LLDP and traffic tools such as Ixia or Spirent.
  • GenAI operations: Familiarity with LLM and agent evaluation, groundedness, hallucination and safety checks, regression metrics, tracing, human review, LangSmith, LangGraph, Langfuse, and MCP.
  • Problem‑solving and collaboration: Ability to analyse architecture, diagnose cross‑layer failures, assess operational risk, communicate root causes clearly, collaborate across teams, and guide junior engineers; bachelor’s degree or equivalent practical experience.
  • Platform and infrastructure: Experience with infrastructure as code, GitOps, secrets management, and multi‑environment or hybrid‑cloud operations.
  • AI quality and observability: Experience with agent evaluation, prompt regression, model comparison, tracing, production monitoring, and operational controls for AI services.
  • Performance and resilience: Proficiency in API and browser performance testing, cloud‑scale resilience, chaos testing, failover validation, and disaster‑recovery exercises.
  • Security and complex platforms: Knowledge of OWASP risks, threat modelling, vulnerability management, adversarial validation, and multi‑tenant or distributed enterprise architectures.
  • Credentials: Relevant AWS, Kubernetes, DevOps, security, SRE, or networking certifications.
What We Can Offer You
Health & Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal & Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have — whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

Let's Stay Connected

Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.

#puertorico

#networking

Job

Engineering

Job Level

TCP_03

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together.

Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer - Generative AI & Enterprise Web Platforms
DevOps Engineer - Generative AI & Enterprise Web Platforms

Hewlett Packard Enterprise Company in • San Juan (PR)

Hybrid
USD 100,000 - 150,000
AI Ops Engineer
AI Ops Engineer

Hewlett Packard Enterprise Company in • San Juan (PR)

Hybrid
USD 110,000 - 160,000
Relocation support
Health & wellbeing benefits
Professional development programs
AI Ops Engineer
AI Ops Engineer

Hewlett Packard Enterprise Company • Friday Harbor (WA)

Hybrid
USD 120,000 - 160,000
Relocation support
Hybrid work model
AI Ops Engineer
AI Ops Engineer

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 90,000 - 150,000
Hybrid work model (2 days in-office)
Relocation support
Software Engineering Manager - Generative AI & Enterprise Platforms
Software Engineering Manager - Generative AI & Enterprise Platforms

Hewlett Packard Enterprise Company • Friday Harbor (WA)

Hybrid
USD 180,000 - 240,000
Software Engineering Manager - Generative AI & Enterprise Platforms
Software Engineering Manager - Generative AI & Enterprise Platforms

Hewlett Packard Enterprise Company in • San Juan (PR)

Hybrid
USD 150,000 - 230,000
Principal Engineer - Generative AI & LLM Platforms
Principal Engineer - Generative AI & LLM Platforms

Hewlett Packard Enterprise Company • Friday Harbor (WA)

Hybrid
USD 210,000 - 270,000
Software Engineering Manager – Generative AI & Enterprise Platforms
Software Engineering Manager – Generative AI & Enterprise Platforms

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 140,000 - 200,000
Health & Wellbeing
Career Development
Unconditional Inclusion
Principal Engineer - Generative AI & LLM Platforms
Principal Engineer - Generative AI & LLM Platforms

Hewlett Packard Enterprise Company in • San Juan (PR)

Hybrid
USD 180,000 - 240,000
Health & Wellbeing
Professional Development
Inclusive culture
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 120,000 - 160,000
Comprehensive benefits suite
Personal & professional development opportunities
Unconditional inclusion in the workplace