Senior Site Reliability Engineer

Veloc Inc

Coppell (TX)

On-site

USD 140,000 - 210,000

Full time

17 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Veloc Inc is seeking a Senior Site Reliability Engineer to own the reliability, availability, and performance of our cloud infrastructure and SaaS platforms. You will work hands-on in both operations and platform engineering to drive reliability initiatives from conception through production.

You will lead automation efforts, IaC, CI/CD improvements, and incident response while collaborating across teams to improve deployment safety, capacity planning, and operational scalability.

Qualifications

  • 7+ years in Site Reliability Engineering, DevOps, Cloud Infrastructure, or Production Operations.
  • Strong experience operating workloads in Microsoft Azure, AWS, or Google Cloud.
  • Hands-on with Kubernetes, Docker, CI/CD pipelines, and IaC tools.
  • Scripting and automation in Python, Bash, PowerShell, or Go.
  • Experience with observability platforms (Datadog, Grafana, Prometheus, Splunk).
  • Solid networking, Linux/Windows admin, and cloud-native architectures.
  • Incident response, production troubleshooting, and governance expertise.

Responsibilities

  • Own day-to-day monitoring, alerting, on-call support for SaaS platforms and cloud infra.
  • Lead major incident response, root-cause analysis, and postmortems.
  • Design and maintain high-availability, backup, and DR procedures.
  • Investigate and resolve production incidents across infra, platform, and apps.
  • Design, implement, and maintain IaC, deployment automation, and CI/CD improvements.
  • Develop tooling to reduce toil and boost engineering productivity.
  • Partner with development teams to improve deployment safety and scalability.
  • Drive standardization of cloud infra and deployment governance.

Skills

SRE/DevOps expertise
Cloud infrastructure operations
Kubernetes & Docker
CI/CD automation
Scripting (Python, Bash, PowerShell,Go
Observability tooling
Incident response
Strong communication

Tools

Kubernetes
Docker
CI/CD pipelines
IaC tools
Datadog/Grafana/Prometheus/Splunk
Terraform/Bicep/ARM/Ansible

Job description

Senior Site Reliability Engineer — combination of deep operational expertise and hands-on engineering ability. The majority of your time (~70%) will be focused on owning the reliability, availability, scalability, and operational excellence of the cloud infrastructure and SaaS platforms powering our business. The remaining ~30% puts you directly in the platform engineering flow: building automation, improving deployment pipelines, and driving reliability initiatives from conception through production.

Key Responsibilities
Reliability Engineering & Operations (~40% of role)
  • Own day-to-day monitoring, alerting, operational health, and on-call support for mission-critical SaaS platforms and cloud infrastructure.
  • Lead major incident response activities including escalation coordination, root cause analysis, and postmortem reviews.
  • Design and maintain high-availability, failover, backup, and disaster recovery procedures; validate RTO/RPO targets regularly.
  • Investigate and resolve production incidents end-to-end across infrastructure, platform, and application layers.
Automation & Platform Engineering (~30% of role)
  • Design, implement, and maintain Infrastructure as Code (IaC), deployment automation, and CI/CD pipeline improvements.
  • Develop tooling and automation to reduce operational toil and improve engineering productivity.
  • Partner with development teams to improve deployment safety, release reliability, and operational scalability.
  • Drive standardization of cloud infrastructure, operational engineering practices, and deployment governance.
Observability & Performance Optimization (~15% of role)
  • Build and maintain monitoring, logging, tracing, and alerting capabilities across distributed systems.
  • Establish service-level objectives (SLOs), SLIs, and error budget policies.
  • Identify and remediate performance bottlenecks, scaling issues, and infrastructure inefficiencies.
  • Analyze operational telemetry and trends to improve reliability and capacity planning.
Security, Compliance & Architecture (~15% of role)
  • Implement operational security best practices including RBAC, least privilege access, and infrastructure hardening.
  • Ensure compliance with SOC 2, HIPAA, GDPR, and organizational security standards.
  • Participate in architecture reviews and operational readiness assessments for new services and platforms.
  • Mentor junior engineers on reliability engineering, cloud operations, automation, and incident management best practices.
Required Qualifications
  • 7+ years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, or Production Operations roles.
  • Strong experience operating workloads in cloud environments such as Microsoft Azure, AWS, or Google Cloud.
  • Hands-on experience with Kubernetes, Docker, CI/CD pipelines, and Infrastructure as Code tools.
  • Strong scripting and automation skills using Python, Bash, PowerShell, Go, or similar languages.
  • Experience with observability and monitoring platforms such as Datadog, Grafana, Prometheus, or Splunk.
  • Strong understanding of networking, Linux/Windows administration, distributed systems, and cloud-native architectures.
  • Experience with incident response, production troubleshooting, and operational governance.
  • Strong communication skills and ability to collaborate across engineering and business teams.
Preferred Qualifications
  • Experience supporting multi-tenant SaaS environments.
  • Experience with Terraform, Bicep, ARM templates, or Ansible.
  • Familiarity with GitOps and modern deployment strategies such as canary or blue/green deployments.
  • Experience working within regulated or compliance-driven environments.
  • Relevant cloud or Kubernetes certifications.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Denver (CO)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Seattle (WA)

On-site
USD 140,000 - 190,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Jobtailor • Arizona

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

On-site
USD 100,000 - 135,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000