Application Support Engineer [Multiple Positions Available]

JPMorgan Chase & Co.

Plano (TX)

On-site

USD 120,000 - 180,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

JPMorgan Chase & Co. is seeking a Senior Site Reliability Engineer to define reliability targets, design monitoring, and drive proactive incident management.

You will lead cross-team initiatives to improve resilience, automate failover, and support scalable platforms across cloud regions. Applicants should have a 5-year track record in SRE/DevOps, strong skills in Kubernetes, AWS, Terraform, and observability, and a Bachelor’s degree in a related field.

Qualifications

  • Bachelor's degree plus 5 years of relevant experience in reliability engineering or related roles.
  • Experience with Sev1/Sev2 incidents, RCA, postmortems, and MTTR improvement.
  • Strong focus on observability, automation, and scalable fault-tolerant systems.

Responsibilities

  • Define and enforce measurable reliability targets for critical environments.
  • Architect and implement advanced monitoring and alerting across stacks.
  • Lead incident management, root cause analysis, and long-term improvements.
  • Automate failover, recovery, database and service resiliency workflows.
  • Oversee deployment practices with CI/CD, blue/green, canaries, and rollbacks.

Skills

On-call and incident management
Observability and monitoring
CI/CD automation
Kubernetes deployments
AWS production operations

Education

Bachelor's degree in Information Systems Engineering, Computer Engineering, or related field

Tools

Terraform
Python
Bash
SQL
Linux
Kubernetes
AWS

Job description

DESCRIPTION

Duties: Define and enforce measurable reliability targets for critical application environments, ensuring operational metrics are consistently tracked and achieved. Architect, implement, and refine advanced monitoring and alerting solutions to proactively surface application health and performance issues across complex technology stacks. Collaborate with cross-functional teams to design and support robust, highly available application platforms capable of meeting stringent business and technical demands. Drive adoption of reliability and resilience best practices providing technical leadership to elevate operational standards across support and engineering groups. Develop, automate, and maintain failover and recovery workflows for application services ensuring uninterrupted operations across diverse infrastructure and cloud regions. Create, update, and operationalize incident response documentation and automated remediation mechanisms including rollback and circuit breaker strategies for application failures. Lead critical incident management for application platforms ensuring rapid restoration, comprehensive root cause analysis, and implementation of long-term improvements. Optimize resource allocation and infrastructure costs for large-scale application support, balancing efficiency with reliability, and performance requirements. Facilitate seamless integration and deployment of application changes bridging development and support to ensure smooth operational transitions. Oversee continuous validation processes including pre- and post-deployment monitoring to detect and remediate application drift and performance regressions. Maintain day-to-day operational stability and high availability for application systems, leveraging deep technical expertise in support and troubleshooting. Monitor production environments using advanced diagnostic and observability tools rapidly identifying and resolving anomalies. Escalate and communicate complex technical issues, delivering actionable insights and solutions to both technical and business stakeholders.

QUALIFICATIONS

Minimum education and experience required: Bachelor's degree in Information Systems Engineering, Computer Engineering, or related field of study plus 5 years of experience in the job offered or as Site Reliability Engineer, Data Engineer, Data Analyst, MSSQL Server Developer, Support Engineer, Software Developer, or related occupation.

Skills Required: This position requires five (5) years of experience with the following: Utilizing Sev1/Sev2 on-call including triaging, mitigating, coordinating restoration, RCA, and blameless postmortems; MTTR recurrence reduction; Observability including metrics, logs, traces; low-noise alerts, fast detection; maintaining runbooks and escalations; Automation including Python and Bash; utilizing CI/CD with blue and green or canary, feature flags, pre-deploy validation, and reliable rollback; using IaC including Terraform for public cloud; using Modules, remote state, drift detection, LUT/policy-as-code, and targeted applies; using Kubernetes for deployments, autoscaling, health probes, progressive rollouts and rollbacks, quotas, and cluster troubleshooting; utilizing AWS production operations including VPC networking, IAM policy design, encryption/ KMS, load balancing, multi-AZ resilience, backup and restore, and regional failover; Database reliability including PostgreSQL, MySQL, Oracle backups/PITR, replication and failover, online schema changes, SQL and index tuning under load; Linux and networking including processing memory/IO diagnostics, kernel and sysctl tuning, TLS, TCP/IP, DNS, HTTP, load balancing, and service discovery; Security including least privilege, secrets management, patch, vulnerability management, immutable audit logging and framework alignment. This position requires three (3) years of experience with the following: Utilizing SLI/SLO and error budgets sustaining ≥99.9% SLOs; Gate high-risk changes; Performance and capacity including load testing (JMeter, BlazeMeter), trace hot paths, tail-latency/throughput tuning, and peak demand modeling; Resilience and DR including timeouts, retries, backpressure, circuit breaking, graceful degradation; validated RTO/RPO; using Linux and networking to process memory/IO diagnostics, kernel and sysctl tuning, TLS, TCP/IP, DNS, HTTP, load balancing, and service discovery; Releasing change management including versioned applications, DB migrations, automated quality gates Maxwell, controlled rollouts, and rapid clean NB rollback; Reliability outcomes including sustaining lower bast MT pipeline failure TR, fewer false alarms, higher S, logical improvements, safer deployments, hardened failure modes, and predictable peak scaling.

Job Location: 8181 Communications Pkwy, Plano, TX 75024.

Full-Time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Application Support Engineer
Application Support Engineer

Pipe Recruit • Coppell (TX)

On-site
USD 80,000 - 100,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

System One • Dallas (TX)

On-site
USD 130,000 - 170,000
Application Engineer View role →
Application Engineer View role →

NRnP Technology • Northern (KY)

Hybrid
USD 90,000 - 140,000
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

StradIT • Jersey City (NJ)

Hybrid
USD 120,000 - 160,000
Systems Analyst 3 529601671
Systems Analyst 3 529601671

LMG Technology Services LLC • Austin (TX)

Hybrid
USD 120,000 - 160,000
Application Support Engineer
Application Support Engineer

Matlen Silver • Jacksonville (FL)

Hybrid
USD <1,000
Application Support Engineer
Application Support Engineer

Jobsbridge • Miami (FL)

On-site
USD 100,000 - 130,000
Site Reliability Engineer in Pittsburgh
Site Reliability Engineer in Pittsburgh

Energy Jobline ZR • Pittsburgh

On-site
USD 130,000 - 180,000
Senior Application Support Developer
Senior Application Support Developer

U.S. Legal Support • Houston (TX)

On-site
USD 110,000 - 140,000