Lead Infrastructure Engineer - Storage

Next Frontier Capital

Houston (TX)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

JPMorganChase is seeking a Lead Infrastructure Engineer to advance our Enterprise Technology Infrastructure Platforms. You will apply deep infrastructure knowledge across storage, Linux, and automation to deliver resilient, secure, AI-assisted operations.

You will collaborate with cross-functional teams to maintain high availability and drive automated, auditable processes. The role emphasizes incident management, SRE fundamentals, and implementations of AI-enabled observability across on-prem

Qualifications

  • Formal training or certification on infrastructure engineering concepts and 5+ years applied experience.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment.
  • Ability to review and validate AI-assisted recommendations before implementation.
  • Strong knowledge of storage fundamentals (RAID/erasure coding, replication, snapshots, IOPS/latency).
  • Hands-on experience with major storage ecosystems (NetApp, Dell EMC, Pure, Ceph, etc.).
  • Solid Linux fundamentals and kernel/storage-stack concepts.
  • Strong scripting in Python/Go/Bash.
  • Experience with observability stacks (Prometheus, Grafana, ELK/OpenSearch, Datadog).
  • Proven incident management skills and on-call rotation ability.
  • Practical AI/data skills for operations (anomaly detection, CI/CD integration).

Responsibilities

  • Uses enterprise AI capabilities to accelerate infrastructure analysis and design documentation.
  • Applies AI-assisted practices to identify recurring issues with traceability and security alignment.
  • Owns and improves SLOs/SLIs, error budgets, on-call readiness, and storage operational excellence.
  • Leads incident response for storage outages and drives RCAs and preventative actions.
  • Creates runbooks, escalation paths, and standardized procedures.
  • Operates and enhances block/file/object storage across on-prem and cloud.
  • Performs capacity planning, lifecycle management, and resiliency testing.
  • Partners with cross-functional teams to meet workload reliability targets.
  • Builds automation for provisioning, patches, upgrades, replication, backups, and compliance.
  • Implements AI-driven observability and safe LLM-enabled workflows.

Skills

AI capabilities
Incident management
SLOs/SLIs
Linux basics
Python/Go/Bash
Observability stacks
Automation scripting

Education

Formal infrastructure training

Tools

NetApp
Dell EMC PowerStore/Isilon
Pure Storage
Ceph
Terraform
Ansible
Kubernetes (storage CSI)
Kafka
Prometheus/Grafana
OpenTelemetry

Job description

Assume a vital position as a key member of a high-performing team that delivers infrastructure and performance excellence. Your role will be instrumental in shaping the future at one of the world's largest and most influential companies.

As a Lead Infrastructure Engineer at JPMorganChase within the Enterprise Technology, Infrastructure Platforms team, you apply deep knowledge of software, applications, and technical processes within the infrastructure engineering discipline. Continue to evolve your technical and cross-functional knowledge outside of your aligned domain of expertise.

Job responsibilities
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis and design documentation, validating outputs and handling operational data according to sensitivity and security requirements.
  • Applies reuse-first, AI-assisted practices within delivery and automation routines to identify recurring issues and validate remediation options, ensuring changes are traceable/auditable and aligned to resiliency and security expectations.
  • Own and continuously improve SLOs/SLIs, error budgets, on-call readiness, and operational excellence for storage services.
  • Lead incident response for storage outages/performance degradations; drive RCAs and implement preventative actions.
  • Create and maintain runbooks, escalation paths, and standardized operational procedures.
  • Operate and enhance block/file/object storage platforms across on-prem and/or cloud environments.
  • Perform performance tuning, capacity planning, lifecycle management, and resiliency testing (failover/DR validation).
  • Partner with infrastructure, network, OS, database, and application teams to meet workload requirements and reliability targets.
  • Build automation for provisioning, patching, upgrades, replication, backup/restore, and compliance checks.
  • Implement AI-driven observability/AIOps (telemetry correlation, anomaly/regression detection, LLM-assisted incident/runbook workflows) with accuracy, auditability, and safe rollout.
Required qualifications, capabilities, and skills
  • Formal training or certification on infrastructure engineering concepts and 5+ years applied experience
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity.
  • Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations.
  • Strong knowledge of storage fundamentals (RAID/erasure coding, replication, snapshots, tiering/caching, IOPS/latency, multipathing, SAN/NAS, object semantics).
  • Hands-on experience with at least one major storage ecosystem (e.g., NetApp, Dell EMC PowerStore/Isilon, Pure, Hitachi, Ceph, IBM, or cloud storage services).
  • Solid Linux fundamentals, including system performance, networking basics, and kernel/storage-stack concepts.
  • Strong scripting/programming in one or more of Python, Go, Bash.
  • Experience with observability stacks (e.g., Prometheus/Grafana, ELK/OpenSearch, Splunk, Datadog, OpenTelemetry).
  • Proven incident management skills and ability to operate effectively in an on-call rotation.
  • Practical AI/data skills for operations (anomaly detection/forecasting/correlation/classification; feature extraction and evaluation; integrating AI into production tooling/CI/CD; safe LLM use with guardrails and human-in-the-loop review).
Preferred qualifications, capabilities, and skills
  • Kubernetes storage (CSI), stateful workloads, and container platform operations.
  • Infrastructure as Code (Terraform/CloudFormation) and configuration management (Ansible/Chef/Puppet).
  • Streaming/queue tooling for telemetry and event pipelines (e.g., Kafka).
  • Experience with ITSM/event management platforms (e.g., ServiceNow).
  • Backup/DR products and strategy design, including RPO/RTO tradeoffs.
  • Security controls for data platforms (KMS/HSM, secrets management, key rotation).
  • Experience building/operating controlled self-service platforms with guardrails to reduce toil at scale.

JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans

Our Global Technology Infrastructure group is a team of innovators who love technology as much as you do. Together, you’ll use a disciplined, innovative and a business focused approach to develop a wide variety of high-quality products and solutions. You’ll work in a stable, resilient and secure operating environment where you—and the products you deliver—will thrive.

High Risk Roles (HRR) are sensitive roles within the technology organization that require high assurance of the integrity of staff by virtue of 1) sensitive cybersecurity and technology functions they perform within systems or 2) information they receive regarding sensitive cybersecurity or technology matters. Users in these roles are subject to enhanced pre-hire screening which includes both criminal and credit background checks (as allowed by law). The enhanced screening will need to be successfully completed prior to commencing employment or assignment.

Carry out critical infrastructure engineering solutions across multiple technical areas as an integral part of an agile team

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Vice President - Lead Infrastructure Engineer | Storage
Vice President - Lead Infrastructure Engineer | Storage

Fairygodboss • Houston (TX)

On-site
USD 140,000 - 190,000
Lead Infrastructure Engineer - Linux and VMWare System Admin
Lead Infrastructure Engineer - Linux and VMWare System Admin

JPMorganChase • Plano (TX)

On-site
USD 120,000 - 160,000
Senior Lead Software Engineer - Full Stack/Infrastructure
Senior Lead Software Engineer - Full Stack/Infrastructure

Fairygodboss • Plano (TX)

On-site
USD 180,000 - 240,000
Health care
On-site health centers
Retirement plan
+2
Senior Lead Software Engineer - Network Test Automation Platform
Senior Lead Software Engineer - Network Test Automation Platform

JPMorganChase • Plano (TX)

On-site
USD 150,000 - 230,000
Infrastructure Engineer III - Storage SRE / Storage Platform Engineer
Infrastructure Engineer III - Storage SRE / Storage Platform Engineer

JPMorganChase • Plano (TX)

On-site
USD 120,000 - 180,000
Associate - Infrastructure Engineer | Storage SRE / Storage Platform
Associate - Infrastructure Engineer | Storage SRE / Storage Platform

Fairygodboss • Plano (TX)

On-site
USD 110,000 - 160,000
Senior Lead Software Engineer- Platform / Linux Engineering
Senior Lead Software Engineer- Platform / Linux Engineering

JPMorganChase • New York (NY)

On-site
USD 180,000 - 240,000
Infrastructure Engineer III
Infrastructure Engineer III

JPMorganChase • Houston (TX)

On-site
USD 110,000 - 170,000
Senior Lead Infrastructure Engineer - Network Engineer
Senior Lead Infrastructure Engineer - Network Engineer

JPMorganChase • Plano (TX)

On-site
USD 140,000 - 190,000
Health coverage
On-site wellness centers
Retirement plan
+4
Lead Infrastructure Engineer-Network Engineer
Lead Infrastructure Engineer-Network Engineer

Next Frontier Capital • United States

On-site
USD 170,000 - 230,000