A leading technology company is seeking an AI & HPC Observability Engineer to join their Managed AI Superclusters (MARS) team in Seattle. The role involves designing and scaling observability platforms and telemetry pipelines. Candidates should have experience with distributed systems and programming skills in Python, Go, or Java. A Bachelor's degree in Computer Science or a related field is required. The position offers a competitive salary range of $152,000 to $287,500 based on experience, along with equity and benefits.
Qualifications
5+ years of experience building backend or distributed systems in production environments.
Strong programming skills in Python, Go, or Java.
Hands-on experience with observability architectures, including metrics, logs, and traces.
Responsibilities
Design and scale observability platforms for high-volume metrics.
Build high-performance backend services for telemetry ingestion.
Develop and extend OpenTelemetry tools for metrics and telemetry.
Skills
Python programming
Go programming
Java programming
Distributed systems knowledge
Observability architectures
Kafka
Spark
Flink
Kubernetes
Prometheus
Education
Bachelor's degree in Computer Science or related field
Tools
OpenTelemetry
Time-series data systems
Job description
A leading technology company is seeking an AI & HPC Observability Engineer to join their Managed AI Superclusters (MARS) team in Seattle. The role involves designing and scaling observability platforms and telemetry pipelines. Candidates should have experience with distributed systems and programming skills in Python, Go, or Java. A Bachelor's degree in Computer Science or a related field is required. The position offers a competitive salary range of $152,000 to $287,500 based on experience, along with equity and benefits.