SRE Lead: Java, Kafka & Node.js | AI-Driven Reliability

Appathon

Phoenix, Northern (AZ, KY)

Hybrid

USD 76,000 - 83,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Appathon seeks an experienced Site Reliability Engineering Lead in Phoenix, AZ. The role focuses on reliability, scalability, and automation for cloud-native systems using Java, Kafka, and Node.js. Hybrid work is available for local candidates.

Requirements include 12+ years in SRE, strong Java/Node.js skills, Kafka expertise, and cloud and observability tooling proficiency. AI/GenAI exposure is a plus as the team drives innovative operational improvements.

Qualifications

  • 12+ years IT experience with strong SRE/Production Engineering background.
  • Hands-on experience with Java and Node.js development.
  • Strong expertise with Apache Kafka and event-driven architectures.
  • Experience with monitoring and observability tools such as Splunk, Prometheus, Grafana, Datadog, or ELK.
  • Knowledge of cloud platforms (AWS, Azure, or GCP).
  • Experience with Docker, Kubernetes, and containerized environments.
  • Strong understanding of CI/CD pipelines and automation tools.
  • Excellent troubleshooting and incident management skills.
  • Exposure to AI/ML or Generative AI concepts is highly desirable.

Responsibilities

  • Lead SRE initiatives to improve system reliability, scalability, and performance.
  • Design, implement, and support highly available distributed applications.
  • Develop and maintain services using Java and Node.js.
  • Build and manage event-driven architectures using Apache Kafka.
  • Establish observability, monitoring, logging, and alerting frameworks.
  • Drive automation for deployments, incident response, and operational processes.
  • Collaborate with development, infrastructure, and DevOps teams to improve system resilience.
  • Conduct root cause analysis and implement preventive measures.
  • Support CI/CD pipelines and infrastructure automation.
  • Stay informed on emerging AI and Generative AI technologies and identify opportunities for operational improvements.

Skills

Java
Node.js
Apache Kafka
AI concepts
Cloud platforms
Docker
Kubernetes
CI/CD
Prometheus
Grafana
Datadog
Splunk
ELK

Education

Bachelor's Degree

Tools

Docker
Kubernetes
Terraform
Ansible

Job description

Appathon seeks an experienced Site Reliability Engineering Lead in Phoenix, AZ. The role focuses on reliability, scalability, and automation for cloud-native systems using Java, Kafka, and Node.js. Hybrid work is available for local candidates.

Requirements include 12+ years in SRE, strong Java/Node.js skills, Kafka expertise, and cloud and observability tooling proficiency. AI/GenAI exposure is a plus as the team drives innovative operational improvements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineering - Java, Kafka, Node.js & AI
Site Reliability Engineering - Java, Kafka, Node.js & AI

Appathon • Phoenix (AZ), Northern (KY)

Hybrid
USD 76,000 - 83,000
Java + SRE Engineer
Java + SRE Engineer

TechDigital Group • Phoenix (AZ)

On-site
USD 120,000 - 170,000
Java SRE Engineer
Java SRE Engineer

Pacer Group • Phoenix (AZ)

On-site
USD 150,000 - 190,000
Medical insurance
Dental insurance
Vision insurance
+1
Lead Java SRE — Reliability & Observability
Lead Java SRE — Reliability & Observability

KTek Resourcing LLC • Jersey City (NJ)

On-site
USD 140,000 - 170,000
Java SRE Engineer
Java SRE Engineer

EITACIES Inc. • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Lead Network SRE
Lead Network SRE

Talentola • Jersey City (NJ)

On-site
USD 180,000 - 240,000
SRE Lead: Reliability & Automation in AWS | Hybrid
SRE Lead: Reliability & Automation in AWS | Hybrid

Seek Now • Atlanta (GA)

Hybrid
USD 120,000 - 180,000
Competitive salary
Health, dental, and vision coverage
401(k) with company match
+1
Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Access to BetterUp coaching
Competitive compensation plan
Medical, dental, and vision insurance
+3
SRE Engineering Manager — Lead Reliability, Remote Flexible
SRE Engineering Manager — Lead Reliability, Remote Flexible

DevOpsChat • California (MO)

Hybrid
USD 140,000 - 220,000
Health insurance
Professional development opportunities
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops

Arcoro Holdings Corp • Phoenix (AZ), Northern (KY)

Hybrid
USD 200,000 - 220,000
Remote Work
401(k) with Company match
Flexible PTO and Company-paid holidays