Senior Observability Platform Engineer

Together AI

San Francisco (CA)

Hybrid

USD 200,000 - 280,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure.

The storage and observability team is responsible for designing, implementing, and maintaining robust distributed storage solutions and comprehensive observability platforms, providing critical insights into system performance and GPU utilization, and proactively resolving issues.

Qualifications

  • Experience with observability platforms and cloud-native monitoring services across AWS, GCP, and Azure.
  • Proficient in Go or Python and infrastructure-as-code tools like Terraform, Ansible, and Helm.
  • Proven ability to design, operate, and scale large distributed systems and data pipelines.

Responsibilities

  • Design and implement a scalable observability platform (metrics, logs, traces) using Prometheus, Grafana, ClickHouse/ClickStack, and OpenTelemetry, with telemetry pipelines and log aggregation.
  • Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services.
  • Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm.
  • Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis.
  • Define observability best practices.

Skills

Observability platforms
Go
Python
Cloud-native monitoring
CI/CD & GitOps
Distributed systems

Tools

Prometheus
Grafana
ClickStack
OpenTelemetry
AWS
GCP
Azure
Terraform
Ansible
Helm
Docker
Kubernetes
PostgreSQL
MongoDB
Redis

Job description

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure.

The storage and observability team is responsible for designing, implementing, and maintaining robust distributed storage solutions and comprehensive observability platforms, providing critical insights into system performance and GPU utilization, and proactively resolving issues.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Observability Engineer
Senior AI Infra Observability Engineer

Togetherai • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Remote work flexibility
+1
Senior Software Engineer, Observability
Senior Software Engineer, Observability

Together AI • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior Observability Platform Engineer – GPU AI Infra
Senior Observability Platform Engineer – GPU AI Infra

Nscale • Northern (KY)

Hybrid
USD 160,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior AI Observability Platform Engineer
Senior AI Observability Platform Engineer

Socket.dev • United States

On-site
USD 160,000 - 230,000
Senior Observability Platform Engineer - GPU & AI
Senior Observability Platform Engineer - GPU & AI

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
AI-Powered Observability Engineer
AI-Powered Observability Engineer

OpenAI • California (MO)

On-site
USD 140,000 - 200,000
AI-Powered Observability Engineer
AI-Powered Observability Engineer

OpenAI • San Francisco (CA), Mountain View (CA), New York (NY)

On-site
USD 180,000 - 240,000
Senior Observability Platform Engineer - AI-Driven
Senior Observability Platform Engineer - AI-Driven

504 CGCG-US CG Companies Global-US • Town of Charlotte (NY)

On-site
USD 137,000 - 219,000