Site Reliability Engineer

Decisive Point

Pennsylvania

On-site

USD 118,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity packages
401(k) and matching
Medical, dental, and vision insurance
Paid time off

Job summary

Asylon is seeking a Site Reliability Engineer to join their team in Philadelphia. This role entails managing the reliability, availability, and performance of cloud and on-prem systems with a focus on automation and incident response.

The ideal candidate will have experience in SRE and DevOps, be proficient in Kubernetes, and be comfortable with observability tools. An excellent compensation package including benefits is offered.

Qualifications

  • 3+ years of experience in SRE, DevOps, or infrastructure engineering.
  • Proficient in Kubernetes, deploying, operating, and debugging.
  • Experience with observability stacks like Prometheus or Grafana.

Responsibilities

  • Own reliability of production systems across environments.
  • Build and maintain observability infrastructure.
  • Lead incident response and implement remediation.

Skills

Site Reliability Engineering
Python
Kubernetes
CI/CD Pipelines
Linux Systems Administration
Observability Tools (Prometheus, Grafana)

Tools

Terraform
Ansible

Job description

Site Reliability Engineer

Asylon automates security operations using robots, software, and AI to help its customers protect their people, property, and assets. Founded in 2015 by three MIT aerospace engineers, Asylon brings a strong aerospace background to develop high‑quality, reliable robotics solutions for security applications.

About Us

We build field‑deployable infrastructure for automated robots to augment security teams. Our full‑stack aerial and ground robotic solution, DroneCore, provides capabilities that were previously unavailable to security organizations.

Position

Asylon is hiring a Site Reliability Engineer to join our Philadelphia team. You will be responsible for the reliability, availability, and performance of systems across cloud infrastructure, on‑prem servers in air‑gapped customer environments, and Kubernetes clusters on edge devices deployed with our robots. You will define and maintain SLOs, build observability into every layer of the stack, lead incident response, and drive automation that keeps our systems running without manual intervention. This role sits at the intersection of infrastructure engineering and operations—you should be comfortable writing code to eliminate toil as you are triaging an outage on a remote edge device.

Due to the nature of the projects, applicants must be a U.S. Person as defined by 22 C.F.R. §120.62 (U.S. Citizens, lawful permanent residents, refugees, or asylees).

Primary Duties
  • Own the reliability of production systems across cloud, on‑prem, and edge environments—define SLOs, track error budgets, and drive improvements.
  • Build and maintain observability infrastructure—monitoring, alerting, logging, and dashboards—to provide visibility into system health at every layer.
  • Lead incident response, conduct blameless post‑mortems, and implement remediation to prevent recurrence.
  • Develop automation to reduce toil, improve deployment reliability, and enable self‑healing infrastructure.
  • Build and maintain CI/CD pipelines for service deployment, testing, and infrastructure provisioning.
  • Manage Kubernetes clusters (K3s on edge, on‑prem, and managed cloud clusters)—deployments, upgrades, and troubleshooting.
  • Manage infrastructure‑as‑code for reproducible provisioning across cloud and air‑gapped on‑prem environments.
  • Collaborate with software and robotics engineers to build reliability into systems from the design phase.
Required Skills and Experience
  • 3+ years of professional experience in SRE, DevOps, or infrastructure engineering.
  • Strong in a high‑level language such as Python, Go, or Bash for building automation and tooling.
  • Proficient with Kubernetes—deploying, operating, debugging, and scaling containerized workloads.
  • Experience building and operating observability stacks—Prometheus, Grafana, Loki, or similar tools.
  • Background in CI/CD pipelines for automated testing, building, and deploying services.
  • Proficient with Linux systems administration and troubleshooting.
  • Experience with infrastructure‑as‑code tools such as OpenTofu, Terraform, or Ansible.
  • Comfortable with networking fundamentals—DNS, firewalls, VPNs, and debugging connectivity issues across distributed environments.
Bonus Points
  • Experience with K3s or lightweight Kubernetes on edge—running services on resource‑constrained hardware in the field.
  • Experience working in air‑gapped or disconnected environments where systems must operate without cloud dependencies.
  • Experience with on‑call rotations and structured incident management processes.
  • Familiarity with message brokers and streaming (MQTT, NATS, Kafka, or similar) for real‑time data pipelines.
  • Experience with robotics or IoT systems, particularly managing fleets of remote devices.
  • Experience with video streaming or processing pipelines in a production environment.
  • Comfortable getting hands‑on with hardware—building robots, tinkering with a Raspberry Pi, or debugging a device on a bench.
  • Experience with Bazel or similar build systems for managing complex, multi‑language codebases.
  • Experience with capacity planning and performance engineering.
Benefits
  • Competitive salary and equity packages.
  • 401(k) and 401(k) matching.
  • Medical, dental, and vision insurance.
  • Life insurance.
  • Paid time off.
  • Relocation assistance.
  • Flexible vacation/sick time.
Equal Opportunity Employer

Asylon is an Equal Opportunity Employer. We recruit, hire, employ, train, promote, and compensate individuals based on job‑related qualifications and abilities. We provide a work environment that respects the dignity and worth of each individual and is free from all forms of discrimination, including harassment, because of race, color, sex, gender, age, religion, national origin, citizenship, marital status, sexual orientation, gender identity, genetic information, disability, military status, or any other characteristic protected by law.

Pay: $118,000 – $150,000 per year

Location: Philadelphia, PA (Norristown, PA 19401)

Application Question: Are you a “U.S. Person” as defined in 22 C.F.R. §120.62?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer
Software Engineer

Decisive Point • Pennsylvania

On-site
USD 105,000 - 150,000
Parental leave
401(k)
Health insurance
+4
Full Stack Software Engineer
Full Stack Software Engineer

Decisive Point • Pennsylvania

On-site
USD 115,000 - 160,000
Competitive salary and equity packages
401(k) and matching
Flexible vacation/sick time
+4
Frontend Engineer/UX Engineer
Frontend Engineer/UX Engineer

Decisive Point • Pennsylvania

On-site
USD 120,000 - 150,000
Medical, dental, and vision plans
401(k) retirement savings plans
Life and disability insurance
+1
Security Robotics Operator
Security Robotics Operator

Decisive Point • Norristown

On-site
401(k)
Health insurance
Paid time off
+5
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Socket.dev • San Mateo (CA)

On-site
USD 240,000 - 300,000
Stock options
Health insurance
401K savings plan
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Skydio • United States

On-site
USD 240,000 - 300,000
Stock options
Health insurance
Paid vacation
+4
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Skydio • San Mateo (CA)

On-site
USD 240,000 - 300,000
Equity
Health insurance
Paid time off
+1
Site Reliability Engineer
Site Reliability Engineer

Skydio • United States

On-site
USD 180,000 - 240,000
Stock options
Health insurance
Paid vacation
+3
Site Reliability Engineer, Robotics
Site Reliability Engineer, Robotics

Hadrian Automation • Los Angeles (CA)

On-site
USD 164,000 - 270,000
Medical, dental, vision, and life insurance
401k
Flexible vacation policy
+1
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Figureai • San Jose (CA)

On-site
USD 175,000 - 250,000