Site Reliability Engineer

OXIO Corporation

United States

Remote

USD 120,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

OXIO Corporation is seeking a Site Reliability Engineer to design and run scalable, reliable cloud platforms for our NeoTelco ecosystem. You will automate deployments, scaling, and recovery, while ensuring 24/7 uptime for critical services.

You will work with Linux systems, Docker/Kubernetes, and modern CI/CD pipelines, contributing to incident management and blameless postmortems. This role requires collaboration across engineering, telecom, and data teams in a fast-paced environment.

Qualifications

  • Linux/Unix administration and networking fundamentals are required.
  • Proficiency in at least one programming language (Python, Go, or Ruby) and strong scripting (Bash, Perl).
  • Experience with infrastructure provisioning tools (Terraform, CloudFormation, Ansible).
  • Familiarity with containerization (Docker) and orchestration (Kubernetes).
  • Experience with monitoring/observability tools (Prometheus, Grafana, Datadog).
  • Knowledge of incident management, runbooks, and blameless postmortems.
  • Experience building CI/CD pipelines (Jenkins, GitLab CI, CircleCI).
  • Hands-on with cloud providers (AWS, Google Cloud, Azure) and cloud-native architectures.

Responsibilities

  • Design and implement cloud-based platform to support backend services.
  • Automate deployments, scaling, and recovery processes.
  • Monitor production infrastructure to maximize uptime and reliability.
  • Participate in on-call rotations and postmortem-driven improvements.
  • Provide tooling and support to engineering teams for service operations.

Skills

Linux/Unix
Programming (Python/Go/Ruby)
Scripting (Bash/Perl)
IaC (Terraform/CloudFormation/Ansible)
Docker
Kubernetes
Monitoring (Prometheus/Grafana/Datadog

Tools

Terraform
CloudFormation
Ansible
Docker
Kubernetes
Prometheus
Grafana
Datadog
Jenkins
GitLab CI
CircleCI
AWS
GCP
Azure

Job description

Site Reliability Engineer

OXIO is the first NeoTelco. We are building the world’s largest, most accessible, and insightful Telecom network. Our platform empowers anyone to spin up their own carrier from a browser, scaling and supporting you as you scale your network to millions of users.

We ensure that users and devices are connected, and stay connected wherever they go: Cross- country, carrier, or cellular technology. We help them pay less for mobile data. This technology is provided through our Carrier-as-a-Service platform: BrandVNO, a fully customizable telecom service. In addition, we enable clients of our service to extract the value from telecom data - enriching their customer experience, business intelligence, and product understanding in the many markets in which we operate.

Come join us in creating a modern technology platform with a group of engineers dedicated to advancing our vision. Our team is passionate about what we build, open to new ideas and challenges, and has our sights set on the future of connectivity.

Responsibilities

  • Design and implement platform on the cloud to support OXIO backend services

  • Automate technical operations: deployments, scaling, recovery, etc.

  • Monitor and maintain mission-critical production infrastructure to ensure maximum uptime

  • Participate in an on-call rotation and culture of continuous improvement through blameless postmortems

  • Enable the Engineering/Telecom/Data Engineering teams by providing them the tools to operate the service they build

Essentials
  • Understanding of Linux/Unix systems (most systems are Linux-based).

  • Familiarity with Linux/Unix system internals like process management, filesystems, memory management, and networking.

  • Proficiency in at least one programming language (Python, Go, or Ruby) and strong skills in scripting (Bash, Perl).

  • Experience with infrastructure provisioning tools such as Terraform, CloudFormation, or Ansible.

  • Familiarity with containerization (Docker) and orchestration tools (Kubernetes).

  • Familiarity with monitoring tools like Prometheus, Grafana, or Datadog.

  • Knowledge of setting up alerts, analyzing logs, and creating dashboards for observability.

  • Familiarity with incident management practices (e.g., runbooks, postmortems).

  • Experience in being part of an on-call rotation and handling incidents.

  • Experience in setting up and maintaining Continuous Integration/Continuous Delivery pipelines (Jenkins, GitLab CI, CircleCI, etc.).

  • Hands-on experience with cloud providers (AWS, Google Cloud, Azure).

  • Knowledge of virtualization technologies (VMware, KVM) and cloud-native architecture.

  • Understanding of TCP/IP, DNS, HTTP/HTTPS, load balancing, and firewalls.

Nice to have
  • Strong understanding of deployment strategies (canary releases, blue-green deployments, etc.).

  • Familiarity with high availability and understanding failover mechanisms.

  • Familiarity with IAM (Identity and Access Management) and zero trust principles.

  • Experience working with distributed systems (e.g., Kafka, Cassandra, Elasticsearch).

  • Building custom monitoring tools or writing complex automation scripts.

  • Functional knowledge of database management (SQL and NoSQL).

  • Familiarity with distributed tracing (Jaeger, OpenTelemetry) and advanced log aggregation strategies (ELK stack, Splunk).

  • Familiarity with performance profiling tools and optimizing application performance under heavy load.

  • Familiarity in load testing and identifying bottlenecks.

  • Familiarity with Configuration Managment using SaltStack for maintaining server configurations.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

oxio • United States

Remote
USD 120,000 - 160,000
Site Reliability Engineer - Cloud Platform & Automation
Site Reliability Engineer - Cloud Platform & Automation

oxio • United States

Remote
USD 120,000 - 160,000
Staff Data Platform Engineer
Staff Data Platform Engineer

OXIO • New York (NY)

On-site
USD 150,000 - 210,000
Stock options
Healthcare
Flexible work
+3
Technical Lead / Exec IT Support
Technical Lead / Exec IT Support

OXIO Inc. • New York (NY)

On-site
USD 120,000 - 180,000
Equity
Health benefits
Flexible work
+4
Mobile Telecom Support Engineer
Mobile Telecom Support Engineer

oxio • United States

Remote
USD 120,000 - 180,000
Cloud Platform SRE — Automation & Uptime
Cloud Platform SRE — Automation & Uptime

OXIO Corporation • United States

Remote
USD 120,000 - 190,000
Staff Data Engineer
Staff Data Engineer

oxio • United States

Hybrid
USD 150,000 - 190,000
Stock options
Healthcare
Flexible work
+4
Junior Mobile Core Support Engineer
Junior Mobile Core Support Engineer

King River Capital Group • United States

On-site
USD 90,000 - 120,000
Product Manager
Product Manager

Multicoin • United States

On-site
USD 90,000 - 120,000
Competitive compensation
Equity participation
Comprehensive healthcare coverage
+2
Staff Data Engineer
Staff Data Engineer

OXIO Inc. • United States

On-site
USD 120,000 - 150,000
Competitive salary and stock options
Company paid healthcare
Flexible work arrangements
+2