Server Software Engineer Palo Alto, Bangalore

Aria Networks, Inc.

Palo Alto, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Aria Networks, Inc. seeks a Server Software Engineer for Telemetry & Data Infrastructure to own backend systems ingesting, processing, storing, and delivering high-frequency network telemetry from Aria’s switch deployments at GPU-cluster scale.

You will design and operate pipelines sustaining microsecond granularity across 100,000+ GPU environments, handling vast data with strict latency budgets to support AI workloads and frontend systems.

Qualifications

  • BS/MS in Software Engineering and 5+ years of backend systems engineering with a focus on high-throughput or low-latency data infrastructure.
  • Demonstrated experience building and operating production data pipelines at scale in Go, Rust, or C++.
  • Deep familiarity with streaming architectures — Kafka, gRPC, or equivalent — including partitioning, ordering guarantees, and consumer group management.

Responsibilities

  • Design and implement high-throughput, low-latency telemetry ingestion pipelines capable of sustaining microsecond-granularity data from switch ASICs across large-scale GPU cluster deployments.
  • Architect and optimize time-series storage backends for telemetry workloads with demanding retention, cardinality, and query latency requirements.
  • Define and enforce data schemas, stream partitioning strategies, and backpressure mechanisms across Kafka or gRPC-based ingestion paths.
  • Build processing and aggregation layers that transform raw switch telemetry into structured, queryable signals for AI/ML and operator-facing systems.
  • Profile and optimize critical pipeline components for CPU efficiency, memory pressure, and tail latency at sustained high-cardinality load.
  • Collaborate with switch software engineers to define and evolve telemetry export contracts, sampling rates, and protocol bindings (gNMI, OpenTelemetry, INT, or equivalent).
  • Instrument the pipeline itself — define SLOs, build internal observability, and own reliability and data fidelity in production.
  • Drive technical decisions on storage engine selection, stream processing topology, and data lifecycle management as telemetry scope scales.

Skills

Backend systems
Latency-sensitive design
Performance profiling
Concurrency
Data pipelines

Education

BS/MS in Software Engineering

Tools

Kafka
gRPC
InfluxDB
Prometheus
ClickHouse
OpenTelemetry
gNMI

Job description

As a Server Software Engineer on Telemetry & Data Infrastructure, you will own the backend systems that ingest, process, store, and deliver high-frequency network telemetry from Aria’s switch deployments at GPU-cluster scale. You will design and operate pipelines that sustain microsecond-level granularity across 100,000+ GPU environments where data volume is large, latency budgets are strict, and correctness directly impacts AI workload performance. You will solve hard problems in streaming throughput, time-series storage efficiency, and low-latency query delivery that few systems at this fidelity have tackled in production networking contexts. This role spans the full stack from ingestion protocol to storage schema to the query interfaces consumed by AI/ML and frontend systems.

Responsibilities
  • Design and implement high-throughput, low-latency telemetry ingestion pipelines capable of sustaining microsecond-granularity data from switch ASICs across large-scale GPU cluster deployments
  • Architect and optimize time-series storage backends for telemetry workloads with demanding retention, cardinality, and query latency requirements
  • Define and enforce data schemas, stream partitioning strategies, and backpressure mechanisms across Kafka or gRPC-based ingestion paths
  • Build processing and aggregation layers that transform raw switch telemetry into structured, queryable signals for AI/ML and operator-facing systems
  • Profile and optimize critical pipeline components for CPU efficiency, memory pressure, and tail latency at sustained high-cardinality load
  • Collaborate with switch software engineers to define and evolve telemetry export contracts, sampling rates, and protocol bindings (gNMI, OpenTelemetry, INT, or equivalent)
  • Instrument the pipeline itself — define SLOs, build internal observability, and own reliability and data fidelity in production
  • Drive technical decisions on storage engine selection, stream processing topology, and data lifecycle management as telemetry scope scales
Qualifications
  • BS/MS in Software Engineering and 5+ years of backend systems engineering with a focus on high-throughput or low-latency data infrastructure
  • Demonstrated experience building and operating production data pipelines at scale in Go, Rust, or C++
  • Deep familiarity with streaming architectures — Kafka, gRPC, or equivalent — including partitioning, ordering guarantees, and consumer group management
  • Hands-on experience with time-series or columnar storage systems such as InfluxDB, Prometheus, ClickHouse, or comparable engines
  • Strong fundamentals in concurrency, memory management, and performance profiling in systems-level languages
  • Ability to reason about end-to-end pipeline latency, data loss scenarios, and correctness tradeoffs under load
  • Experience designing internal APIs or query interfaces consumed by downstream ML or application teams
  • Experience with network telemetry protocols including gNMI, OpenTelemetry, SNMP, or sFlow is a significant plus
  • Prior work on observability or monitoring infrastructure at a hyperscaler, cloud networking company, or large-scale distributed systems organization is a significant plus
  • Experience processing telemetry emitted directly from switching ASICs or hardware line cards is a significant plus
  • Familiarity with data center networking concepts — fabric topologies, congestion signals, flow tracking — is preferred
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Switch Software Engineer Palo Alto, Bangalore
Switch Software Engineer Palo Alto, Bangalore

Aria Networks, Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Low-Latency Telemetry Backend Engineer
Low-Latency Telemetry Backend Engineer

Aria Networks, Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Backend/Infra Engineer
Backend/Infra Engineer

Judgment Labs • San Francisco (CA)

On-site
USD 140,000 - 180,000
Backend/Infra Engineer
Backend/Infra Engineer

Judgment Labs Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Backend Engineer
Backend Engineer

AnoSys Technologies, Inc • Los Angeles (CA)

On-site
USD 100,000 - 140,000
Staff AI Observability & Telemetry Engineer
Staff AI Observability & Telemetry Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000
Staff HPC Network Architect
Staff HPC Network Architect

Lambda • United States

Hybrid
USD 180,000 - 280,000
Staff AI Observability & Telemetry Engineer
Staff AI Observability & Telemetry Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 160,000 - 210,000
Software Engineer - Host Networking
Software Engineer - Host Networking

Meta • Bellevue (WA)

On-site
USD 150,000 - 210,000
Staff AI Observability & Telemetry Engineer
Staff AI Observability & Telemetry Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000