DevOps & Site Reliability Engineer

VoltaGrid, LLC

Houston (TX)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

VoltaGrid, LLC is seeking a DevOps & SRE Engineer in Houston to design, build, and evolve scalable infrastructure. You will implement robust deployment pipelines and drive reliability across production systems, collaborating closely with software teams to ensure operational excellence.

The role emphasizes Kubernetes management, Terraform-based IaC, and strong Linux administration. You will participate in on-call rotations and help improve monitoring and incident response processes.

Qualifications

  • 4+ years of experience in DevOps, SRE, or infrastructure engineering roles.
  • Strong hands-on experience with Kubernetes and Docker in production.
  • Proficiency with infrastructure-as-code tooling, particularly Terraform.
  • Experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins).
  • Solid understanding of monitoring and observability (metrics, logs, traces).
  • Strong Linux systems administration skills (Ubuntu, RHEL/CentOS).
  • Experience with virtualization platforms and capacity planning.

Responsibilities

  • Design, build, and maintain cloud infrastructure.
  • Manage and optimize Kubernetes clusters and containerized workloads in production.
  • Develop and maintain infrastructure-as-code using Terraform (or equivalent).
  • Build and improve CI/CD pipelines for fast, safe deployments.
  • Implement and maintain monitoring, alerting, and observability systems (Prometheus, Grafana, Datadog).
  • Define and track SLIs/SLOs; participate in incident response and blameless postmortems.
  • Identify and eliminate toil through automation and self-service tooling.
  • Collaborate with development teams on system design and capacity planning.
  • Participate in on-call rotations to ensure production readiness of new services.

Tools

Terraform
GitHub Actions
GitLab CI
Jenkins
Prometheus

Job description

Position Title: DEVOPS & SRE ENGINEER
Location: HOUSTON, TX
FLSA Class: EXEMPT
Responsible to: Directo of Software Engineering


Position Summary: DevOps / Site Reliability Engineer to implement and evolve the infrastructure, deployment pipelines, and reliability posture of our systems. You'll work closely with engineering teams to build scalable, observable, and resilient infrastructure while driving a culture of operational excellence.


Essential Duties and Responsibilities:



  • Design, build, and maintain cloud infrastructure

  • Manage and optimize Kubernetes clusters and containerized workloads in production

  • Develop and maintain infrastructureascode using Terraform (or equivalent tooling)

  • Build and improve CI/CD pipelines to enable fast, safe, and reliable deployments

  • Implement and maintain monitoring, alerting, and observability systems (Prometheus, Grafana, Datadog, or similar)

  • Define and track SLIs/SLOs, participate in incident response, root cause analysis, and blameless postmortems

  • Identify and eliminate toil through automation and selfservice tooling

  • Configure and maintain onprem baremetal servers and Linuxbased infrastructure

  • Configure, maintain, and optimize virtualized assets

  • Collaborate with development teams on system design, capacity planning, and performance optimization

  • Participate in oncall rotations and ensure production readiness of new services


Other Requirements:



  • 4+ years of experience in DevOps, SRE, or infrastructure engineering roles

  • Strong experience with at least one major cloud provider (AWS, GCP, or Azure AWS preferred)

  • Deep hands-on experience with Kubernetes and Docker in production environments

  • Proficiency with infrastructureascode tools, particularly Terraform

  • Experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, or similar)

  • Solid understanding of monitoring and observability (metrics, logs, traces)

  • Strong scripting skills (Bash, Python, or Go)

  • Experience with incident management, SLObased reliability practices, and capacity planning

  • Strong Linux systems administration skills (Ubuntu, RHEL/CentOS, or similar)

  • Experience with virtualization platforms including VM provisioning, storage, networking, and cluster management

  • Solid understanding of networking, DNS, load balancing, and security fundamentals


Nice to Have:



  • Contributions to internal developer platforms or platform engineering initiatives

  • Proxmox VE experience

  • Certifications in cloud platforms (AWS SA, CKA, etc.)


The above statements are intended to describe the general nature and level of work being performed by employees assigned to this classification. All personnel may be required to perform duties outside of their normal responsibilities from time to time, as needed.


VoltaGrid is an Equal Opportunity Employer that does not discriminate on the basis of actual or perceived race, creed, color, religion, alienage or national origin, ancestry, citizenship status, age, disability or handicap, sex, marital status, veteran status, sexual orientation, genetic information, arrest record, or any other characteristic protected by applicable federal, state or local laws.


Our management team is dedicated to this policy with respect to recruitment, hiring, placement, promotion, transfer, training, compensation, benefits, employee activities, and general treatment during employment. #LI-KM1


Qualifications

Skills

Behaviors

:


Motivations

:


Education

Experience

Licenses & Certifications

Equal Opportunity Employer
This employer is required to notify all applicants of their rights pursuant to federal employment laws.For further information, please review the Know Your Rights notice from the Department of Labor.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

DevOps & Site Reliability Engineer
DevOps & Site Reliability Engineer

VoltaGrid LLC. • Houston (TX), Northern (KY)

On-site
USD 120,000 - 150,000
DevOps & Site Reliability Engineer
DevOps & Site Reliability Engineer

VoltaGrid • Houston (TX)

On-site
USD 120,000 - 180,000
Senior Software Engineer
Senior Software Engineer

VoltaGrid LLC • Houston (TX), Northern (KY)

On-site
USD 120,000 - 150,000
Sr Software Engineer, Fullstack
Sr Software Engineer, Fullstack

VoltaGrid • Houston (TX)

On-site
USD 120,000 - 170,000
DevOps Engineer
DevOps Engineer

REGION 4 ESC • Houston (TX)

On-site
USD 87,000 - 104,000
Infrastructure/DevOps Engineer
Infrastructure/DevOps Engineer

Veritas Automata • Smithfield (RI)

On-site
USD 95,000 - 120,000
Flexible working hours
Continuous education opportunities
Health and wellness programs
Infrastructure/DevOps Engineer
Infrastructure/DevOps Engineer

Veritas Automata • Merrimack (NH)

On-site
USD 110,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Talentify • Greenville (SC)

On-site
USD 120,000 - 170,000
Competitive Wages
Health/Life Benefits
401(k) with company match
+8
Infrastructure/Cloud DevOps - SRE
Infrastructure/Cloud DevOps - SRE

Bayside Solutions • Cupertino (CA)

On-site
USD 150,000 - 230,000
Sr Software Engineer, Fullstack
Sr Software Engineer, Fullstack

Worky • Town of Texas (WI)

On-site
USD 140,000 - 175,000