Senior Engineer, Cloud Infrastructure and Networking

Skylo

United States

Hybrid

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Skylo is reshaping satellite connectivity from Mountain View, CA, with a direct-to-device NTN vRAN platform. We are seeking a Senior, Cloud Infrastructure and Networking to own 24x7 health across hybrid cloud and on‑prem environments, ensuring fault tolerance, observability, and rapid recovery across GKE, Kubernetes, and related OSS tools.

You’ll drive runbooks, incident response, and post-change validation while collaborating with NI and product teams to keep Skylo’s core services highly

Qualifications

  • Own the observability pipeline end-to-end: Prometheus scrape target integrity, VictoriaMetrics retention and query performance, Grafana dashboard coverage and accuracy, OpenTelemetry collector health, and alert routing via Pub/Sub to the OSS.
  • Maintain database reliability: PostgreSQL streaming replication health, backup and restore procedures, failover testing, query performance monitoring; Redis cluster operations, eviction policy management, and persistence configuration.
  • Ensure log aggregation pipeline health (Loki or ELK): ingestion rates, retention policies, query performance, and completeness — the observability stack must be operational before the network events it monitors can be triaged.
  • Partner with NI (Network Implementation & Infrastructure) on all planned infrastructure changes: receive advance notice, validate post-deployment observability, and sign off on operational readiness before the change window closes

Responsibilities

  • Own 24x7 cloud infrastructure health across Skylo's hybrid production environment: GKE cluster node status, namespace and pod health, Persistent Volume Claim availability, network policies, and multi-cluster federation across Skylo's GCP footprint.
  • Own on-premise Kubernetes cluster health: bare-metal node availability, container runtime stability, CNI networking, persistent storage arrays (Ceph/Rook or equivalent), and hyperconverged compute platform operations (Harvester, KubeVirt, or KVM).
  • Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics telemetry, GCP Cloud Monitoring, and Loki log correlation — distinguish transient platform events from systemic infrastructure degradation.
  • Execute and own Cloud Infra runbooks for P2–P4 fault categories: GKE node recovery, pod eviction and rescheduling, PVC repair, database failover execution, Prometheus WAL corruption recovery, ArgoCD drift remediation, and certificate rotation — without requiring engineering involvement for covered fault classes.
  • Own BSS-IIS GKE cluster monitoring and infrastructure health; maintain runbooks that reflect current cluster topology after every infrastructure change.

Skills

Kubernetes
GKE
ArgoCD
Prometheus
Grafana
OpenTelemetry
PostgreSQL
Redis
Ceph
Rook
KubeVirt
Harvester
KVM
GCP

Tools

Harvester
KubeVirt
KVM
Ceph
Rook

Job description

The world still has coverage blind spots. You could help eliminate them at Skylo. Skylo has pioneered a standards-based approach to satellite connectivity. We connect smartphones and IoT devices directly to satellites. No special hardware, no entirely new networks. Just billions of existing devices, suddenly reachable anywhere on Earth. We're not building toward this future. We're already in it.

Our direct-to-device service is live on millions of activated devices across five continents, covering more than 72 million square kilometers, in partnership with leading satellite operators, mobile network operators, Tier-1 chipset makers, and OEMs worldwide. And we're just getting started.

At the heart of it all is Skylo's commercial NTN vRAN: a 3GPP standards-based, cloud-native platform that seamlessly bridges terrestrial and satellite networks. It's the infrastructure that makes true anywhere, anytime connectivity possible.

When you join Skylo, you'll work at the intersection of three markets reshaping how the world stays connected: mass-market consumer devices, automotive, and industrial IoT. Enabling people outdoors and critical workflows in the world's most remote places.

This is a rare chance to work on technology that matters, at a company that's already proving it works

About Skylo

Skylo is a global Non-Terrestrial Network (NTN) service provider based in Mountain View, CA, offering a service that allows smartphone and IoT cellular devices to connect directly over existing satellites.

Skylo's direct-to-device service is live on millions of activated devices across five continents, with more than 60 million square kilometers of coverage, in partnership with multiple satellite operators, mobile network operators (MNOs), Tier-1 chipset makers, and OEMs. Devices connected over satellite are managed and served by Skylo's commercial NTN vRAN — a 3GPP standards-based, cloud-native base station and core. Skylo provides an anywhere, anytime connectivity solution that seamlessly roams between terrestrial and satellite networks. Our focus is on enabling connected services across three main verticals: mass-market consumer devices, automotive, and industrial IoT.

How You Will Impact Skylo

As a Senior, Cloud Infrastructure and Networking, in the Global Product Support & Customer Success organization, you are the Cloud Infrastructure domain authority within Skylo's production NTN network. Everything runs on the infrastructure you keep healthy — RAN NFs, Core NFs, OSS, BSS, and the observability pipeline itself. When a GKE node fails, when ArgoCD drifts, when a Persistent Volume Claim goes unavailable, when a PostgreSQL replica falls behind, when Prometheus WAL corrupts — you own the response.

You operate across Skylo's full hybrid cloud estate: GCP public cloud (GKE clusters, Pub/Sub pipelines, Cloud SQL) and on-premise private cloud infrastructure (bare-metal Kubernetes, hyperconverged compute, software-defined storage). You own 24x7 platform health, the observability pipeline (Prometheus, VictoriaMetrics, Grafana, OpenTelemetry), persistent storage operations (PostgreSQL, Redis), and the operational interface with Network Implementation for all GitOps-driven infrastructure changes.

Key Responsibilities

  • Own 24x7 cloud infrastructure health across Skylo's hybrid production environment: GKE cluster node status, namespace and pod health, Persistent Volume Claim availability, network policies, and multi-cluster federation across Skylo's GCP footprint.
  • Own on-premise Kubernetes cluster health: bare-metal node availability, container runtime stability, CNI networking, persistent storage arrays (Ceph/Rook or equivalent), and hyperconverged compute platform operations (Harvester, KubeVirt, or KVM).
  • Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics telemetry, GCP Cloud Monitoring, and Loki log correlation — distinguish transient platform events from systemic infrastructure degradation.
  • Execute and own Cloud Infra runbooks for P2–P4 fault categories: GKE node recovery, pod eviction and rescheduling, PVC repair, database failover execution, Prometheus WAL corruption recovery, ArgoCD drift remediation, and certificate rotation — without requiring engineering involvement for covered fault classes.
  • Own BSS-IIS GKE cluster monitoring and infrastructure health; maintain runbooks that reflect current cluster topology after every infrastructure change.

Observability Pipeline & Data Platform Operations

  • Own the observability pipeline end-to-end: Prometheus scrape target integrity, VictoriaMetrics retention and query performance, Grafana dashboard coverage and accuracy, OpenTelemetry collector health, and alert routing via Pub/Sub to the OSS.
  • Maintain database reliability: PostgreSQL streaming replication health, backup and restore procedures, failover testing, query performance monitoring; Redis cluster operations, eviction policy management, and persistence configuration.
  • Ensure log aggregation pipeline health (Loki or ELK): ingestion rates, retention policies, query performance, and completeness — the observability stack must be operational before the network events it monitors can be triaged.
  • Partner with NI (Network Implementation & Infrastructure) on all planned infrastructure changes: receive advance notice, validate post-deployment observability, and sign off on operational readiness before the change window closes
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Engineer, Cloud Infrastructure and Networking
Senior Engineer, Cloud Infrastructure and Networking

engineeringjobs.net, Inc. • Town of Montana (WI)

On-site
USD 125,000 - 135,000
Stock options
Medical, dental, vision
Retirement plan
+4
Senior Cloud Networking Engineer
Senior Cloud Networking Engineer

Skylo • Mountain View (CA)

On-site
USD 144,000 - 181,000
Competitive compensation packages
Comprehensive medical, dental, vision benefits
Monthly wellness and education allowances
+3
Director, Engineering Program Manager
Director, Engineering Program Manager

Skylo • Mountain View (CA)

On-site
USD 215,000 - 230,000
Stock option-based equity
Comprehensive health benefits
Retirement plan
+4
Senior Network Reliability Engineer, Incident Management
Senior Network Reliability Engineer, Incident Management

Skylo • Mountain View (CA)

On-site
USD 150,000 - 210,000
Principal Network Engineer, Planning & Performance
Principal Network Engineer, Planning & Performance

Skylo • Mountain View (CA)

On-site
USD 229,000 - 244,000
Stock options
Medical, dental, vision
Retirement plan
+3
Senior Product Manager, Services
Senior Product Manager, Services

Skylo • Mountain View (CA)

Hybrid
USD 185,000 - 195,000
Stock options
Medical benefits
Retirement plan
+3
Solutions Architect – Skylo Government Systems
Solutions Architect – Skylo Government Systems

Skylo • Mountain View (CA)

On-site
USD 180,000 - 225,000
Senior Product Manager, Services
Senior Product Manager, Services

Skylo Technologies • Mountain View (CA)

On-site
USD 185,000 - 195,000
Stock options/equity
Medical, dental, vision
Retirement plan
+4
Senior Cloud Infra & Networking Architect - Hybrid Cloud
Senior Cloud Infra & Networking Architect - Hybrid Cloud

Skylo • United States

Hybrid
USD 150,000 - 210,000
Solutions Engineer (IoT)
Solutions Engineer (IoT)

Skylo • United States

Hybrid
USD 155,000 - 165,000
Stock option equity
Medical benefits
Wellness & education stipends
+2