Site Reliability Engineer, Commercial

Aalyria

United States

Hybrid

USD 115,000 - 135,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

401(k) plan
Equity
Hybrid remote work

Job summary

Aalyria, a leading aerospace technology company, seeks a senior SRE/Platform Engineer to build the core observability stack for satellite and space systems. You will own metrics, logs, and tracing, defining SLOs/SLIs and error budgets, and collaborating with engineers to implement scalable, reliable tooling.

Role involves on-call duties and guiding the roadmap from cloud-native tools to a robust production-grade platform using Prometheus, OpenTelemetry, Tempo, Terraform, and multi-cloud

Qualifications

  • 4+ years in SRE or platform engineering focusing on observability for large-scale systems.
  • Hands-on expertise building and scaling observability platforms (Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger).
  • Strong production experience with GCP and Kubernetes.
  • Experience with IaC and GitOps (ArgoCD).
  • Proficiency in Go and Python for tooling.
  • Experience defining and managing SLOs/SLIs and error budgets.

Responsibilities

  • Design and build Aalyria's centralized observability platform, scaling metrics, logging, and tracing.
  • Define and manage SLOs/SLIs and error budgets for core products.
  • Collaborate with SWE teams to implement observability best practices and templates.
  • Automate deployment and management of the observability stack using Terraform and GitOps.

Skills

Prometheus
Grafana
OpenTelemetry
Tempo/Jaeger
Kubernetes
Go/Python tooling

Tools

GCP
AWS
Terraform
ArgoCD
GitLab CI

Job description

Aalyria is a leading technology company that supplies laser communications technology and temporospatial software-defined networking platforms to the aerospace industry. With technology acquired from Google, Aalyria is at the forefront of innovation in satellite and airborne mesh networks, as well as cislunar and deep-space communications. We are revolutionizing the orchestration and management of planetary mesh networks using any radio or optical spectrum, any orbit, and any hardware across land, sea, air, and space.

Role Overview:

This isn't a "keep the lights on" SRE role. This is a strategic, high-impact opportunity to build the nervous system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that ensures the reliability of systems critical to the operation of satellite megaconstellations and missions to deep space.

This is a greenfield/brownfield opportunity. You will be a trusted expert, helping to define and implement the strategy and building the tools that empower our engineers. You will support the roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role is for you.

Note: this role includes on-call responsibilities.

Key Responsibilities:
  • Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry).
  • Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready.
  • Partner with SWEs to implement observability best practices, develop standard templates and documentation, and configure tooling (e.g., OpenTelemetry libraries).
  • Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (e.g. Terraform) and GitOps principles (e.g. ArgoCD).
  • Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying GCP and AWS environments.
  • Develop and lead the company's monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems.
Required Qualifications:
  • 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems.
  • Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems.
  • Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes.
  • Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD).
  • Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling.
  • Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems.
Preferred Qualifications:
  • Experience operating a multi-cloud environment, specifically GCP and AWS.
  • Hands-on experience with GitLab CI for CI/CD pipelines.
  • Working knowledge of service mesh technologies such as Istio or Linkerd.
  • Familiarity with instrumenting applications written in Go and C++.
  • An active Secret clearance, or higher, is preferred for this position.
  • Experience with JVM observability (tuning, monitoring) for Java-based applications.
What We Offer:
  • Innovative Environment: Work at a cutting-edge company shaping the future of aerospace communications.
  • Impactful Work: Directly contribute to critical national security programs and initiatives.
  • Growth Opportunities: Expand your career with opportunities for professional development and advancement.
  • Inclusive Culture: Be part of a collaborative, supportive, and inclusive workplace where your contributions matter.
  • Flexibility: Flexible working arrangements including hybrid remote/in-office schedules.
  • Compensation and Equity: Competitive salary, comprehensive benefits (401(k), dental, vision, health, life insurance), paid time off, and equity options.

This position involves access to export-controlled information. To comply with U.S. government export regulations, applicants must meet one of the following criteria:

(A) Qualify as a U.S. person, which includes:

  • U.S. lawful permanent resident (green card holder)
  • Asylee under 8 U.S.C. 1158

(B) Be eligible to access export-controlled information without requiring an export authorization.

(C) Be eligible and reasonably likely to obtain the necessary export authorization from the appropriate U.S. government agency.

The company reserves the right to decline pursuing an export licensing process for legitimate business-related reasons.

Equal Opportunity Employer Statement:

Aalyria is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, color, religion, sex (including pregnancy, gender identity, and sexual orientation), national origin, age, disability status, genetic information, protected veteran status, or any other characteristic protected by law. Qualified applicants from all backgrounds are encouraged to apply.

The pay range for this role is:

115,000 - 135,000 USD per year (Remote (United States))

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Commercial
Site Reliability Engineer, Commercial

aalyria-careers • United States

Hybrid
USD 140,000 - 210,000
401(k)
Health insurance
Equity options
+1
Site Reliability Engineer, USG
Site Reliability Engineer, USG

Aalyria • United States

On-site
USD 125,000 - 150,000
Hybrid/remote flexibility
401(k) plan, health/dental/vision/life
Senior Software Engineer, Infrastructure & Platform
Senior Software Engineer, Infrastructure & Platform

aalyria-careers • United States

Hybrid
USD 160,000 - 210,000
Innovative environment
Impactful work
Growth opportunities
+3
Remote SRE — Observability Platform Lead
Remote SRE — Observability Platform Lead

Aalyria • United States

On-site
USD 125,000 - 150,000
Hybrid/remote flexibility
401(k) plan, health/dental/vision/life
Remote Site Reliability Engineer — Observability Platform
Remote Site Reliability Engineer — Observability Platform

Aalyria • United States

On-site
USD 115,000 - 135,000
401(k) plan
Equity
Hybrid remote work
Software Engineer, Optimization Systems
Software Engineer, Optimization Systems

Aalyria Technologies, Inc. • United States

On-site
USD 135,000 - 185,000
Innovative environment
Impactful work
Growth opportunities
+3
IT Support Specialist
IT Support Specialist

Aalyria Technologies, Inc. • Livermore (CA)

On-site
USD 75,000 - 100,000
Competitive salary
Comprehensive benefits
Flexible working arrangements
Software Engineer, Optimization Systems
Software Engineer, Optimization Systems

Aalyria • United States

On-site
USD 135,000 - 185,000
Innovative environment
Impactful work on national security
Growth opportunities
+3
Senior Software Engineer, Backend Systems
Senior Software Engineer, Backend Systems

Aalyria • United States

On-site
USD 165,000 - 215,000
Hybrid/remote-friendly work options
Equity options
401(k), dental, vision, health, life
Lead Product Security Engineer
Lead Product Security Engineer

Aalyria • United States

On-site
USD 180,000 - 240,000
Innovative environment
Impactful work
Growth opportunities
+3