A leading technology company in Santa Clara is seeking an AI & HPC Observability Engineer to design and develop high-performance observability platforms. You will work on building telemetry pipelines and ensuring system reliability. The ideal candidate has over 5 years in backend systems, is proficient in programming languages like Python, Go, or Java, and has hands-on experience with observability architectures. This role comes with a competitive salary and benefits, emphasizing diversity and equal opportunity.
Qualifications
5+ years of experience building backend or distributed systems in production environments.
Solid experience with PromQL and time-series data systems.
Experience working with cloud-native infrastructure.
Responsibilities
Design and scale observability platforms for high-volume metrics and logs.
Develop and extend OpenTelemetry collectors and instrumentation libraries.
Optimize metrics pipelines using time-series storage systems.
Skills
Programming in Python, Go, or Java
Experience with distributed systems
Observability and telemetry design
Performance tuning and debugging
Distributed data pipelines using Kafka, Spark, or Flink
Education
Bachelor's degree in Computer Science, Computer Engineering, or related field
Tools
OpenTelemetry
Prometheus
Kubernetes
Job description
A leading technology company in Santa Clara is seeking an AI & HPC Observability Engineer to design and develop high-performance observability platforms. You will work on building telemetry pipelines and ensuring system reliability. The ideal candidate has over 5 years in backend systems, is proficient in programming languages like Python, Go, or Java, and has hands-on experience with observability architectures. This role comes with a competitive salary and benefits, emphasizing diversity and equal opportunity.