Contract Azure Site Reliability Engineer

EVONA

El Segundo (CA)

On-site

USD 110,000 - 165,000

Full time

48 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

EVONA is seeking a Senior Site Reliability Engineer to own infrastructure behind ground systems and embedded spacecraft software on a 6-month engagement. The Kubernetes-based platform is defined in Terraform and treated as mission-critical.

You will operate live systems, improve observability, automation, and resilience while collaborating with software and hardware teams. You will build CI/CD pipelines, manage on-call rotations, and drive reliability through incident response and postmortems in

Qualifications

  • Experience with Infrastructure as Code and Terraform for provisioning and configuration management.
  • Proficient in observability tools (Prometheus, Grafana, InfluxDB) and alerting.
  • Hands-on experience with GitOps using ArgoCD; automation with Ansible and Salt.
  • Experience with Docker and containerized workloads.
  • Exposure to HPC and GPU workloads; Slurm scheduler design and resource management.
  • Ability to work in hybrid/cloud/on-prem environments and debug distributed systems at scale.
  • Previous contract or independent consulting work; ability to deliver value quickly.

Responsibilities

  • Deploy, maintain and operate mission-critical applications and infrastructure for spacecraft and company systems.
  • Build and evolve Terraform IaC frameworks.
  • Implement observability with metrics, logs, tracing, and actionable alerts.
  • Build and maintain CI/CD pipelines for safe and rapid deployments.
  • Provide tooling support to software and hardware engineers for rapid iteration.
  • Identify bottlenecks and reliability risks; improve performance and stability.
  • Respond to production incidents; perform root cause analysis and blameless postmortems.
  • Take turns on the on-call rotation.

Skills

Terraform/IaC
Observability
CI/CD
Kubernetes
On-call
Root cause analysis
Incident response

Tools

Terraform
Prometheus
Grafana
InfluxDB
ArgoCD
Ansible
Salt
Docker
Slurm

Job description

Open hourly rate

6-month contract - 40 hours per week

A space & defense company is hiring a Senior Site Reliability Engineer to own the infrastructure behind its ground systems and the embedded software running on its spacecraft through a 6-month engagement. The platform is Kubernetes-based, defined in Terraform, and treated as mission-critical.

The systems are live, the vehicles are flying. What's needed is an operator and builder who has run Kubernetes in production before, can pick up an existing IaC estate without a long runway, and can carry it from where it is now to something more observable, more automated and more resilient on a fixed clock.

What you'll be doing
  • Deploy, maintain and operate mission-critical applications and infrastructure supporting spacecraft and company-wide systems
  • Build and evolve Infrastructure as Code frameworks in Terraform
  • Implement and run observability (metrics, logging, tracing) with alerting that people act on
  • Build and maintain CI/CD pipelines for safe, repeatable, rapid deployments
  • Partner with software and hardware engineers so they have the tooling to iterate quickly
  • Find and fix bottlenecks and reliability risks; tune performance and land long-term stability improvements
  • Respond to production incidents, run root cause analysis and drive corrective actions through blameless postmortems
  • Take your turn on the on-call rotation
What you'll need
  • Infrastructure as Code with Terraform (or similar) for provisioning and configuration management
  • Prometheus, Grafana, InfluxDB or similar
  • GitOps - ArgoCD, Ansible, Salt
  • Docker
  • HPC and GPU workloads — Slurm queue/partition design, fair-share scheduling, cluster resource management
  • Hybrid cloud + on-prem/edge environments, and debugging distributed systems at scale
  • Prior contract or independent consulting work, you land and add value fast
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Reliability Engineer
Reliability Engineer

Chenega Agile Real Time Solutions, LLC • Atlanta (GA)

On-site
USD 95,000 - 105,000
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine • Chicago (IL)

Hybrid
USD 120,000 - 150,000
Professional growth
Competitive compensation
Exciting projects
+1
Reliability Engineer
Reliability Engineer

Chenega Agile Real Time Solutions, LLC • Washington

On-site
USD 95,000 - 105,000
Reliability Engineer
Reliability Engineer

Chenega Corporation • Washington

Hybrid
USD 95,000 - 105,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

On-site
USD 100,000 - 135,000
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine • Atlanta (GA)

Hybrid
USD 100,000 - 130,000
Professional growth
Competitive compensation
Exciting projects
+1
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mikealbert • Cincinnati (OH)

Hybrid
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

On-site
USD 120,000 - 155,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Socket.dev • Kentucky

Hybrid
USD 120,000 - 170,000