Senior AI Infra Observability Engineer

Togetherai

San Francisco (CA)

On-site

USD 200,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Startup equity
Remote work flexibility
Competitive salary

Job summary

Together AI is building the AI Acceleration Cloud. We seek an experienced observability and infrastructure engineer to design scalable monitoring, logging, and tracing for a distributed AI platform, enabling real-time insights and reliable operations.

Responsibilities include building automated alerts, SLIs/SLOs, and anomaly detection, plus deploying tools with Go, Python, Terraform, Ansible, and Helm. You will guide best practices and collaborate across teams on incident response and

Qualifications

  • Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry) and cloud-native monitoring services (AWS, GCP, Azure).
  • Strong programming skills in Go, Python, or similar, with proficiency in IaC tools (Terraform, Ansible, Helm).
  • Experience designing, operating, and scaling large distributed systems and pipelines for high-volume ingestion and real-time queries.
  • Deep understanding of containerization (Docker) and orchestration (Kubernetes).
  • Knowledge of microservices, service mesh, CI/CD, and GitOps workflows.
  • Expertise with databases (PostgreSQL, MongoDB, Redis) and time-series DBs with high-cardinality data.

Responsibilities

  • Design and implement a scalable observability platform (metrics, logs, traces) using Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry.
  • Develop automated monitoring, alerting, and anomaly detection systems including SLIs/SLOs and runbooks.
  • Build and deploy observability tooling and IaC using Go, Python, Terraform, Ansible, and Helm.
  • Collaborate with engineering to enhance tracing and monitoring, and lead post-mortem incident analyses.
  • Define observability best practices for the organization.

Skills

Go
Python
Cloud-native
Prometheus
Grafana
OpenTelemetry
Docker
Kubernetes

Tools

ClickStack
Terraform
Ansible
Helm
PostgreSQL
MongoDB
Redis

Job description

Together AI is building the AI Acceleration Cloud. We seek an experienced observability and infrastructure engineer to design scalable monitoring, logging, and tracing for a distributed AI platform, enabling real-time insights and reliable operations.

Responsibilities include building automated alerts, SLIs/SLOs, and anomaly detection, plus deploying tools with Go, Python, Terraform, Ansible, and Helm. You will guide best practices and collaborate across teams on incident response and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability Platform Engineer
Senior Observability Platform Engineer

Together AI • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

NVIDIA Corporation • California (MO)

Hybrid
USD 184,000 - 357,000
Equity compensation
Benefits
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

Socket.dev • North Carolina

Hybrid
USD 184,000 - 357,000
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior AI Infra Engineer: Telemetry, Observability & CMDB
Senior AI Infra Engineer: Telemetry, Observability & CMDB

NVIDIA • United States

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

NVIDIA Corporation • Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior AI Observability Platform Engineer
Senior AI Observability Platform Engineer

Socket.dev • United States

On-site
USD 160,000 - 230,000
Senior AI Infrastructure Engineer Observability & Automation
Senior AI Infrastructure Engineer Observability & Automation

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 356,000
Equity
Benefits
Senior Software Engineer, Observability
Senior Software Engineer, Observability

Together AI • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
AI Observability Engineer — Scale Logs, Dashboards (Equity)
AI Observability Engineer — Scale Logs, Dashboards (Equity)

OpenAI • Los Angeles (CA)

On-site
USD 255,000 - 405,000