Site Reliability Engineer — AI Platform Observability

Appnovation

Miami (FL)

On-site

USD 120,000 - 180,000

Full time

24 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Appnovation seeks a Site Reliability Engineer to own a shared observability platform for LLM-based applications. The role covers building monitoring, alerting, SLAs, and runbooks while managing Kubernetes on AWS and ClickHouse-based infrastructure.

You will enhance platform reliability and onboarding for internal teams. The platform currently uses Langfuse, Argo CD, Helm, and Kubernetes, with ClickHouse, PostgreSQL, and Redis as core components.

Qualifications

  • Hands-on Kubernetes on AWS (EKS) operations and maintenance.
  • Experience with ClickHouse replication, Keeper quorum and backup/restore.
  • Experience building monitoring, alerting and service levels from scratch.
  • Experience upgrading self-hosted software with rollback plans.
  • PostgreSQL and Redis operations including backpressure and queue behavior.
  • Strong ability to write runbooks and SOPs for on-call scenarios.

Responsibilities

  • Platform Operations: diagnose failures across ClickHouse, PostgreSQL, Redis and Kubernetes.
  • ClickHouse Operations: manage Keeper quorum, replication, shard/topology and S3 tiering.
  • Backup/Restore: rehearse and test restores in outages and failures.
  • Safe Upgrades: plan and rehearse upgrades with rollback for partial migrations.
  • Monitoring/SLAs: build monitoring, alerting and SLAs from scratch.
  • IaC: maintain Kubernetes manifests, Helm values, Argo CD apps.
  • Runbooks/SOPs: write and maintain incident response documentation.
  • Onboarding/Support: enable internal teams relying on platform.
  • Automation: convert recurring tasks into automation.
  • Migration partnerships: coordinate with platform owner on adoption and rollback.

Skills

Kubernetes
AWS EKS
Argo CD
Helm
ClickHouse
PostgreSQL
Redis
Observability
Langfuse
Grafana

Tools

Langfuse
LangSmith
Arize Phoenix

Job description

Appnovation seeks a Site Reliability Engineer to own a shared observability platform for LLM-based applications. The role covers building monitoring, alerting, SLAs, and runbooks while managing Kubernetes on AWS and ClickHouse-based infrastructure.

You will enhance platform reliability and onboarding for internal teams. The platform currently uses Langfuse, Argo CD, Helm, and Kubernetes, with ClickHouse, PostgreSQL, and Redis as core components.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - AI Observability for LLM Apps
Site Reliability Engineer - AI Observability for LLM Apps

Appnovation Technologies • New York (NY), Austin (TX), Dallas (TX), Miami (FL)

On-site
USD 130,000 - 180,000
Contractor role
Site Reliability Engineer — LLM Platform & Observability
Site Reliability Engineer — LLM Platform & Observability

Appnovation • Dallas (TX)

On-site
USD 140,000 - 190,000
Platform Reliability Engineer - Kubernetes & Langfuse
Platform Reliability Engineer - Kubernetes & Langfuse

Appnovation • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation Technologies • New York (NY), Austin (TX), Dallas (TX), Miami (FL)

On-site
USD 130,000 - 180,000
Contractor role
Site Reliability Engineer, AI Platform & Observability
Site Reliability Engineer, AI Platform & Observability

Schonfeld • Northern (KY), New York (NY)

Hybrid
USD 175,000 - 225,000
Performance bonus
Competitive benefits package
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation • Dallas (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation • Miami (FL)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer - Kubernetes Observability
Senior Site Reliability Engineer - Kubernetes Observability

Evlo AI • Denver (CO)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000