Site Reliability Engineer, AI Observability

Appnovation

New York (NY)

On-site

USD 140,000 - 200,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Appnovation seeks a Site Reliability Engineer to build and operate a shared observability platform for LLM-based applications for a global life sciences client. The stack includes Kubernetes on AWS (EKS), ClickHouse, PostgreSQL, Redis, Helm and Argo CD, with Langfuse/LangSmith observability tools.

You will design monitoring, alerting and service levels from scratch, implement infrastructure as code, and create runbooks and SOPs to guide on-call teams and onboarding.

Qualifications

  • Hands-on Kubernetes on AWS (EKS) with Helm and Argo CD.
  • Experience running ClickHouse in production, including backups.
  • Proficient in PostgreSQL and Redis for metadata and queues.
  • Experience building monitoring, alerting and SLAs from scratch.
  • Strong operational writing: runbooks and SOPs.

Responsibilities

  • Platform Operations: diagnose and resolve failures across ClickHouse, PostgreSQL, Redis and Kubernetes.
  • ClickHouse Operations: manage replication, Keeper quorum, shard/topology and S3 tiering.
  • Backup and Restore: rehearse, document and test restores.
  • Safe Upgrades: plan rehearsals with rollback for partial migrations.
  • Monitoring and SLAs: build monitoring and alerting from scratch.
  • Infrastructure as Code: maintain Kubernetes manifests, Helm values and Argo CD.
  • Runbooks/SOPs: write and maintain on-call runbooks.
  • Onboarding and Support: manage internal team onboarding and support path.
  • Automation: convert recurring tasks into automation.
  • Upgrade Partnership: drive migration and rollback alongside platform engineers.

Skills

SRE fundamentals
Automation mindset
Runbooks writing
Incident response

Tools

Kubernetes (AWS EKS)
Helm
Argo CD
ClickHouse
PostgreSQL
Redis
OpenTelemetry
Grafana
Langfuse
LangSmith
Arize Phoenix
GitHub Actions

Job description

About us

Appnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth. Bold ambition. Practical action. Endless possibilities.

We're looking for a Site Reliability Engineer to keep a shared observability platform for LLM-based applications running for a global life sciences client. The platform is built on Langfuse and self-hosted on Kubernetes on AWS, with ClickHouse as the analytical store, PostgreSQL for metadata and Redis as the ingestion queue, all delivered through Argo CD.

Two things need building rather than maintaining. The platform has no monitoring, alerting or defined service levels today, and you will own putting them in place. Infrastructure is defined as code throughout, in Kubernetes manifests, Helm values and Argo CD applications.

Alongside the platform itself, you will own the runbooks that make an incident survivable by someone other than the author, and the onboarding and support path for the internal teams that depend on the platform.

ROLE RESPONSIBILITIES
  • Platform Operations: Diagnose and resolve failures across ClickHouse, PostgreSQL, Redis, the ingestion workers and the Kubernetes layer beneath them, including ingest backpressure from queue depth and worker drain behaviour.
  • ClickHouse Operations: Own ClickHouse under the operator model, including Keeper quorum, replication, shard and replica topology, and S3 storage tiering.
  • Provable Backup and Restore: Rehearse the restore, time it, document it and test it against its failure modes, rather than assuming a successful backup job means a recoverable system.
  • Safe Upgrades: Plan and rehearse upgrades in a lower environment, with a rollback plan that works even if a migration is only partly complete.
  • Monitoring and Service Levels: Build monitoring, alerting and service levels from scratch, so problems are found here before a user reports them.
  • Infrastructure as Code: Maintain Kubernetes manifests, Helm values and Argo CD applications so every change goes through the delivery pipeline.
  • Runbooks and SOPs: Write and maintain runbooks and SOPs that let a colleague resolve an incident without the author present.
  • Onboarding and Support: Run the onboarding and support path for internal teams that depend on the platform, and triage what they bring.
  • Automation: Turn recurring operational work into automation.
  • Upgrade Partnership: Work with the platform engineer who owns what the platform offers. They decide what to adopt and how it is configured; you own the migration and its rollback.
QUALIFICATIONS
  • Hands-on experience with Kubernetes on AWS (managed EKS), with routine work done through Helm values and Argo CD applications.
  • Hands-on experience running ClickHouse in production, including replication and Keeper quorum, shard and replica topology, and backup and restore, ideally run through an operator.
  • Experience building monitoring and alerting from scratch, including service levels that reflect what users actually experience rather than what is easy to measure.
  • Experience upgrading self-hosted software safely, including schema migrations, rehearsal in a lower environment and a rollback plan for a partly completed migration.
  • PostgreSQL and Redis operations deep enough to debug metadata-store and queue problems, including backpressure and worker drain.
  • Strong operational writing: runbooks, SOPs and post-incident reviews that a colleague can follow unaided during an incident.
PREFERRED QUALIFICATIONS
  • Observability engineering, including OpenTelemetry Collector pipelines, alerting design and Grafana dashboards.
  • OIDC or enterprise SSO integration with a corporate identity provider.
  • GitHub Actions for plan and apply pipelines with approval gates.
  • Experience running LLM observability tools such as Langfuse, LangSmith or Arize Phoenix.
  • Experience in pharma, life sciences or another regulated industry.
WHO YOU ARE
  • You don't trust a backup until you've restored from it
  • You want to find problems before users do
  • You write runbooks for the person on call at 3am, not for yourself
  • You stay calm in incidents and focus on fixing the process afterward
  • You automate anything you have to do twice
  • You work well inside a client team and build trust quickly
  • You have prior experience in consulting
  • Prior experience and connections in the Life Sciences industry is preferred

At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation • Dallas (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation • Miami (FL)

On-site
USD 120,000 - 180,000
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation Technologies • New York (NY), Austin (TX), Dallas (TX), Miami (FL)

On-site
USD 130,000 - 180,000
Contractor role
Platform Reliability Engineer
Platform Reliability Engineer

Appnovation Technologies • New York (NY)

On-site
USD 140,000 - 190,000
Platform Reliability Engineer New New York, Austin, Miami, Dallas
Platform Reliability Engineer New New York, Austin, Miami, Dallas

Appnovation • New York (NY)

On-site
USD 140,000 - 200,000
Platform Reliability Engineer - Kubernetes & Langfuse
Platform Reliability Engineer - Kubernetes & Langfuse

Appnovation • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer — AI Platform Observability
Site Reliability Engineer — AI Platform Observability

Appnovation • Miami (FL)

On-site
USD 120,000 - 180,000
DevOps/Cloud Engineer
DevOps/Cloud Engineer

AgileEngine • United States

Remote
MXN 900,000 - 1,300,000
100% remote work
Annual learning budget
Well-being programs
Site Reliability Engineer — LLM Platform & Observability
Site Reliability Engineer — LLM Platform & Observability

Appnovation • Dallas (TX)

On-site
USD 140,000 - 190,000
3510- Site Reliability Engineer II
3510- Site Reliability Engineer II

Innovaccer • Dallas (TX)

On-site
USD 110,000 - 140,000
Generous Paid Time Off: 22 days per year plus company holidays
Best-in-Class Parental Leave
Comprehensive insurance coverage