Platform Reliability Engineer

Appnovation Technologies

New York (NY)

On-site

USD 140,000 - 190,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Appnovation Technologies is building a dedicated delivery pod to develop and run an internal agent platform on AWS Bedrock AgentCore for a global life sciences client. You will implement OpenTelemetry instrumentation, observability integrations, tracing, SLOs, alerting and cost monitoring for token and compute spend, ensuring visibility and reliability across the platform.

We seek a senior SRE/observability engineer with 6+ years of hands‑on AWS experience, OpenTelemetry expertise, and

Qualifications

  • 6+ years in SRE, platform reliability or observability engineering with hands‑on AWS experience.
  • Hands‑on experience with OpenTelemetry and distributed tracing.
  • Experience with Amazon CloudWatch and at least one enterprise observability platform (Datadog, Splunk, Grafana, New Relic or similar).
  • Experience defining and running SLOs, error budgets and alerting strategies.
  • Experience building cost visibility and FinOps reporting on AWS (Cost Explorer, CUR, tagging strategies).
  • Scripting and coding skills in Python, Go or TypeScript.
  • Clear communication and ability to turn data into decisions for technical and non‑technical audiences.

Responsibilities

  • Instrumentation: Build OpenTelemetry instrumentation standards for agents and platform services.
  • Observability integration: Connect AgentCore Observability and CloudWatch with client tools like Datadog, Splunk, Grafana.
  • Tracing: Set up end-to-end tracing across agent steps and model calls.
  • SLOs and alerting: Define SLOs for latency, availability and error rates; build useful alerts.
  • Cost monitoring: Track token usage and compute spend with dashboards, budgets and alerts.
  • Incident readiness: Write runbooks and support post-incident reviews.

Skills

SRE
OpenTelemetry
CloudWatch
Observability tools
SLOs/alerting
Cost monitoring
Python/Go/TypeScript
Communication

Tools

OpenTelemetry SDKs
Datadog
Splunk
Grafana
New Relic

Job description

About us

Appnovation is the AI Product Company. Headquartered in Vancouver, Canada, with teams around the world, Appnovation builds AI products, delivers enterprise AI and digital solutions, and co-invests with clients to bring new AI products to market.

We’re putting together a dedicated delivery pod to build and run an internal agent platform on AWS Bedrock AgentCore for a global life sciences client. The pod works as one team with the client’s engineers to deliver the platform other teams will build their agents on.

In this role you make the platform visible and dependable. You’ll build the OpenTelemetry instrumentation, observability integrations, tracing, SLOs, alerting, and cost monitoring for token and compute spend. When an agent is slow, wrong or expensive, your work is how the team finds out and fixes it. This is a named-team engagement, so the person we propose is the person who starts.

ROLE RESPONSIBILITIES
  • Instrumentation: Build OpenTelemetry instrumentation standards for agents, tools and platform services, and make it easy for teams to adopt.
  • Observability Integration: Connect AgentCore Observability and CloudWatch with the client’s existing observability tools (e.g., Datadog, Splunk, Grafana).
  • Tracing: Set up end-to-end tracing across agent steps, model calls, tool calls and multi-agent handoffs so issues can be traced to their source.
  • SLOs and Alerting: Define SLOs for latency, availability and error rates, and build alerting that is useful and not noisy.
  • Cost Monitoring: Track token usage and compute spend by team, agent and environment, with dashboards, budgets and alerts for unexpected spikes.
  • Incident Readiness: Write runbooks, support incident response and run post-incident reviews that lead to real fixes.
QUALIFICATIONS
  • 6+ years in SRE, platform reliability or observability engineering, with strong hands‑on AWS experience.
  • Hands‑on experience with OpenTelemetry (SDKs, collectors, exporters) and distributed tracing.
  • Experience with Amazon CloudWatch and at least one enterprise observability platform (Datadog, Splunk, Grafana, New Relic or similar).
  • Experience defining and running SLOs, error budgets and alerting strategies.
  • Experience building cost visibility and FinOps reporting on AWS (Cost Explorer, CUR, tagging strategies).
  • Scripting and coding skills in Python, Go or TypeScript.
  • Clear communication and able to turn data into decisions for technical and non‑technical audiences.
PREFERRED QUALIFICATIONS
  • Experience monitoring LLM or agent-based applications, including token usage, latency and quality signals.
  • Familiarity with LLM observability tools (e.g., Langfuse, Arize, LangSmith) or AgentCore Observability.
  • Experience with infrastructure as code (Terraform or AWS CDK) for monitoring resources.
  • Experience in pharma, life sciences or another regulated industry.
WHO YOU ARE
  • You want to know what’s happening in a system before a user tells you
  • You build alerts people trust and dashboards people actually open
  • You treat cost as a reliability concern, not an afterthought
  • You stay calm in incidents and focus on learning afterward
  • You work well inside a client team and build trust quickly
  • You have prior experience in consulting
  • Prior experience and connections in the Life Sciences industry is preferred

Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.

At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how we stay true to our company values as we build better lives through better digital. Accommodations are available upon request throughout the recruitment process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Reliability Engineer New New York, Austin, Miami, Dallas
Platform Reliability Engineer New New York, Austin, Miami, Dallas

Appnovation • New York (NY)

On-site
USD 140,000 - 200,000
AWS Cloud Platform Engineer
AWS Cloud Platform Engineer

Appnovation Technologies • New York (NY)

On-site
USD 140,000 - 190,000
AgentCore Platform Engineer New New York, Austin, Miami, Dallas
AgentCore Platform Engineer New New York, Austin, Miami, Dallas

Appnovation • New York (NY)

On-site
USD 140,000 - 210,000
Accommodations available
Equal Opportunity Employer
AWS Cloud Platform Engineer New New York, Austin, Miami, Dallas
AWS Cloud Platform Engineer New New York, Austin, Miami, Dallas

Appnovation • New York (NY)

On-site
USD 130,000 - 190,000
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation • Dallas (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation • Miami (FL)

On-site
USD 120,000 - 180,000
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation Technologies • New York (NY), Austin (TX), Dallas (TX), Miami (FL)

On-site
USD 130,000 - 180,000
Contractor role
Site Reliability Engineer, AI Observability
Site Reliability Engineer, AI Observability

Appnovation-Technologies • Miami (FL)

On-site
USD 120,000 - 150,000
AgentCore Platform Engineer
AgentCore Platform Engineer

Appnovation Technologies • New York (NY)

On-site
USD 150,000 - 210,000