Site Reliability Engineering - Java, Kafka, Node.js & AI

Appathon

Phoenix, Northern (AZ, KY)

Hybrid

USD 76,000 - 83,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Appathon seeks an experienced Site Reliability Engineering Lead in Phoenix, AZ. The role focuses on reliability, scalability, and automation for cloud-native systems using Java, Kafka, and Node.js. Hybrid work is available for local candidates.

Requirements include 12+ years in SRE, strong Java/Node.js skills, Kafka expertise, and cloud and observability tooling proficiency. AI/GenAI exposure is a plus as the team drives innovative operational improvements.

Qualifications

  • 12+ years IT experience with strong SRE/Production Engineering background.
  • Hands-on experience with Java and Node.js development.
  • Strong expertise with Apache Kafka and event-driven architectures.
  • Experience with monitoring and observability tools such as Splunk, Prometheus, Grafana, Datadog, or ELK.
  • Knowledge of cloud platforms (AWS, Azure, or GCP).
  • Experience with Docker, Kubernetes, and containerized environments.
  • Strong understanding of CI/CD pipelines and automation tools.
  • Excellent troubleshooting and incident management skills.
  • Exposure to AI/ML or Generative AI concepts is highly desirable.

Responsibilities

  • Lead SRE initiatives to improve system reliability, scalability, and performance.
  • Design, implement, and support highly available distributed applications.
  • Develop and maintain services using Java and Node.js.
  • Build and manage event-driven architectures using Apache Kafka.
  • Establish observability, monitoring, logging, and alerting frameworks.
  • Drive automation for deployments, incident response, and operational processes.
  • Collaborate with development, infrastructure, and DevOps teams to improve system resilience.
  • Conduct root cause analysis and implement preventive measures.
  • Support CI/CD pipelines and infrastructure automation.
  • Stay informed on emerging AI and Generative AI technologies and identify opportunities for operational improvements.

Skills

Java
Node.js
Apache Kafka
AI concepts
Cloud platforms
Docker
Kubernetes
CI/CD
Prometheus
Grafana
Datadog
Splunk
ELK

Education

Bachelor's Degree

Tools

Docker
Kubernetes
Terraform
Ansible

Job description

Site Reliability Engineering - Java, Kafka, Node.js & AI

Location

., ., .

Job Type

55.00 - 60.00

Deadline

Not specified

Education

Bachelor's Degree

Experience

Executive (10+ years)

About the Role

We are seeking an experienced Site Reliability Engineering (SRE) Lead with strong expertise in Java, Apache Kafka, and Node.js along with awareness of AI/Generative AI technologies. The ideal candidate will lead reliability initiatives, enhance platform performance, drive automation, and ensure highly available and scalable systems in a cloud-native environment.

Key Responsibilities
  • Lead SRE initiatives to improve system reliability, scalability, and performance.
  • Design, implement, and support highly available distributed applications.
  • Develop and maintain services using Java and Node.js.
  • Build and manage event-driven architectures using Apache Kafka.
  • Establish observability, monitoring, logging, and alerting frameworks.
  • Drive automation for deployments, incident response, and operational processes.
  • Collaborate with development, infrastructure, and DevOps teams to improve system resilience.
  • Conduct root cause analysis and implement preventive measures.
  • Support CI/CD pipelines and infrastructure automation.
  • Stay informed on emerging AI and Generative AI technologies and identify opportunities for operational improvements.
Requirements
Required Skills
  • 12+ years of IT experience with strong SRE or Production Engineering background.
  • Hands-on experience with Java and Node.js development.
  • Strong expertise with Apache Kafka and event-driven architectures.
  • Experience with monitoring and observability tools such as Splunk, Prometheus, Grafana, Datadog, or ELK.
  • Knowledge of cloud platforms (AWS, Azure, or GCP).
  • Experience with Docker, Kubernetes, and containerized environments.
  • Strong understanding of CI/CD pipelines and automation tools.
  • Excellent troubleshooting and incident management skills.
  • Exposure to AI/ML or Generative AI concepts is highly desirable.
Preferred Qualifications
  • Experience leading SRE or platform engineering teams.
  • Knowledge of Infrastructure as Code tools such as Terraform or Ansible.
  • Familiarity with DevSecOps practices.

Openings: 4

Location: Phoenix, AZ (Hybrid) – Local candidates preferred

Skills

AI and Generative AI technologies Apache Java Node.js

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Lead: Java, Kafka & Node.js | AI-Driven Reliability
SRE Lead: Java, Kafka & Node.js | AI-Driven Reliability

Appathon • Phoenix (AZ), Northern (KY)

Hybrid
USD 76,000 - 83,000
Site Reliability Engineer
Site Reliability Engineer

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Java SRE Engineer
Java SRE Engineer

EITACIES Inc. • Santa Clara (CA)

On-site
USD 120,000 - 160,000