Senior DevOps Engineer

HCSS

Sugar Land (TX)

Hybrid

USD 140,000 - 170,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Remote work flexibility
Medical coverage
Dental coverage
Vision coverage
Paid holidays
401(k) with 5% match
Professional development
ERGs
On-site amenities
Dog-friendly campus

Job summary

HCSS is seeking a Senior DevOps Engineer with a strong SRE focus to drive resilience and automation across our cloud environments in Sugar Land, TX. You will own Azure SQL Elastic Pool design, tuning, and automation, and lead high-availability initiatives with IaC practices using Terraform, Bicep, or ARM templates.

Qualifications include 8+ years in DevOps/SRE, 3+ years with Azure Elastic Pools, and 5+ years in Azure services.

Qualifications

  • 8+ years in DevOps or SRE focusing on cloud reliability.
  • 3+ years managing Azure SQL Elastic Pools, including tuning and automation.
  • 5+ years in Azure cloud services: networking, compute, databases, and identity.
  • 3+ years applying SRE: SLIs, SLOs, incident management.

Responsibilities

  • Azure Elastic Pool Management: design, scaling, and optimization of Elastic Pools.
  • Monitor performance and establish observability and alerting for SQL resources.
  • Automate provisioning, scaling, and failover using IaC and scripting tools.
  • Collaborate with database and application teams to optimize resource usage.
  • High Availability and Disaster Recovery: design and implement highly available and fault-tolerant systems.
  • Develop and maintain disaster recovery strategies and runbooks.
  • Monitoring, Observability, and Automation: build dashboards in Grafana and automate tasks.
  • Incident Management and Operational Excellence: lead RCAs and post-incident reviews.
  • Cloud Infrastructure Management: manage Azure or AWS infrastructure and optimize costs.
  • Infrastructure as Code: define IaC with Terraform and maintain modular code.
  • Collaboration and Leadership: mentor juniors and partner with cross-functional teams.

Skills

DevOps experience
Azure SQL Elastic Pools
Azure networking
Azure compute
Azure databases
Azure identity
SLIs SLOs
Terraform
Bicep
ARM templates
CI/CD pipelines
Azure DevOps
GitHub Actions
Azure CLI
PowerShell
Grafana
Observability tools

Tools

Terraform
Bicep
ARM templates

Job description

We are HCSS. For the last 40 years, we have been developing software to help construction companies streamline their operations. Based in Sugar Land, TX, our mission is helping customers achieve excellence through our proven customer-centric, end-to-end solutions and exceptionally helpful service, while providing a great life for our employees. With this mission at the core of everything we do, HCSS is a pioneer and leader in the construction software space and a consistently recognized employer. We have earned Best Companies to Work for in Texas honors for 18 consecutive years and have been named a USA Today Top Workplace. HCSS has also been recognized by Built In as a Best Place to Work in Greater Houston and by Construction Executive for our technology innovation, reflecting our strong culture, industry leadership, and commitment to excellence.

WHO WE NEED:

As a Senior DevOps Engineer with a focus on Site Reliability Engineering (SRE), you will play a key role in driving infrastructure resilience, availability, and operational excellence across our cloud environments. A core responsibility of this role is managing and optimizing Azure SQL Elastic Pools, ensuring performance, cost-efficiency, observability, and automation are aligned with business objectives. You will lead initiatives around high availability, disaster recovery, incident response, and reliability automation. Your expertise in Azure or AWS, observability tools such as Grafana, and scalable infrastructure will be essential to ensuring our systems are robust, performant, and recoverable.

Qualifications:
  • 8+ years of experience in DevOps or SRE roles with a strong focus on cloud infrastructure and systems reliability
  • 3+ years of hands-on experience with managing Azure SQL Elastic Pools, including performance tuning, scaling, and automation
  • 5+ years of expertise in Azure cloud services including networking, compute, databases, and identity
  • 3+ years of experience applying SRE principles including SLIs, SLOs, and incident management best practices
  • Extensive experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM templates
  • Extensive experience building and managing CI/CD pipelines with tools like Azure DevOps or GitHub Actions
  • Strong scripting skills using Azure CLI and PowerShell for automation and operational tasks
  • Experience with monitoring and observability platforms, ideally Grafana, or a strong foundation in similar tools
Soft Skills:
  • Strong troubleshooting and problem-solving abilities.
  • Excellent communication skills and a collaborative mindset to work with cross-functional teams.
  • Ability to work independently, manage multiple tasks, and prioritize efficiently.
  • A proactive attitude toward continuous improvement and learning.
Preferred Qualifications:
  • Managed 10+ Elastic pools and 100+ databases in Azure
  • Advanced level certifications on cloud infrastructure like Az-400 or equivalent
Role Responsibilities:
Azure Elastic Pool Management:
  • Take ownership of the design, scaling, and optimization of Azure SQL Elastic Pools
  • Monitor and tune pool performance to ensure efficiency and SLA compliance
  • Establish observability and alerting for SQL resource consumption, errors, and performance anomalies
  • Automate provisioning, scaling, and failover using infrastructure and scripting tools
  • Collaborate with database and application teams to align on resource usage strategies
High Availability and Disaster Recovery:
  • Design and implement highly available and fault tolerant systems
  • Develop and maintain disaster recovery strategies across critical services
  • Perform regular failover testing, documentation, and validation of recovery procedures
  • Work closely with infrastructure and development teams to ensure business continuity objectives are met
Monitoring, Observability, and Automation:
  • Implement and manage observability stacks with logs, metrics, traces, and alerting
  • Create dashboards and alerts in Grafana or similar platforms to track key system indicators
  • Develop automated solutions for provisioning, monitoring, and maintenance tasks
  • Continuously improve system visibility and reduce time to detect and resolve issues
  • Assist in the automation of performance testing to proactively identify bottlenecks, validate scalability and ensure reliable system behavior
Incident Management and Operational Excellence:
  • Establish and refine incident response processes, escalation workflows, and resolution protocols
  • Lead root cause analysis, post-incident reviews, and continuous improvement efforts, including following up to ensure identified improvements to the application or process are implemented.
  • Develop and maintain runbooks, diagnostic tools, and automated remediation solutions
  • Champion a blameless culture of reliability and operational readiness across engineering teams
Cloud Infrastructure Management:
  • Architect and manage scalable and secure cloud infrastructure in Azure or AWS
  • Provision and manage services including compute, networking, storage, and containerized workloads
  • Continuously monitor performance, latency, and uptime to ensure system health
  • Apply cost optimization practices while aligning infrastructure with business goals
Infrastructure as Code (IaC):
  • Define and implement infrastructure using tools such as Terraform
  • Maintain modular, version-controlled infrastructure code that supports environment consistency
  • Apply automation and policy enforcement to reduce drift and improve auditability
  • Ensure IaC best practices are embedded in the development lifecycle
Collaboration and Leadership:
  • Mentor and support junior DevOps engineers through code reviews, knowledge sharing, and technical guidance
  • Partner with cross-functional teams including development, security, and database operations to drive initiatives
  • Act as a subject matter expert in SRE practices and reliability-driven engineering
  • Lead continuous improvement efforts across infrastructure and operations processes
Travel Requirements:
  • Remote Requirements
    • Employees will be expected to come into the office on a periodic basis.
    • Baseline expectations for roles are as follows but may fluctuate based on manager’s discretion:
      • Individual Contributor - Up to 2x per year
  • Employees will be expected to attend HCSS sponsored events per manager discretion (ex. UGM)
BENEFITS & PERKS:

Part of our mission is to provide a great life for our employees. We believe that when our people are happy, they do their best work. Some of the benefits and perks we offer include:

  • Flexibility to work Remotely
  • Medical, dental, and vision coverage with company-paid and employee-paid options
  • Paid holidays, sick days, and personal time off
  • Employee Resource Groups (ERGs) that foster connection and inclusion
  • On-site amenities including a covered basketball court, soccer field, track, pickleball/tennis courts, gym, etc.
  • Dog-friendly campus and WiFi-accessible courtyards
  • 401(k) with a 5% company match
  • Coverage for employee professional development and wellness
  • And more!
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DevSecOps Engineer
Senior DevSecOps Engineer

HCSS • Sugar Land (TX)

On-site
USD 120,000 - 180,000
Remote work
Medical insurance
Dental insurance
+5
Design Systems Engineer
Design Systems Engineer

HCSS • Sugar Land (TX)

On-site
USD 110,000 - 170,000
Flexible remote options
Medical, dental, and vision coverage
401(k) with 5% company match
+1
Design Systems Engineer II
Design Systems Engineer II

HCSS • Sugar Land (TX)

On-site
USD 120,000 - 170,000
Remotely flexible
Medical, dental, vision coverage
401(k) with match
+2
Senior Business Analyst
Senior Business Analyst

Talentify • Sugar Land (TX)

Hybrid
USD 90,000 - 120,000
Flexibility to work Remotely
Medical, dental, and vision coverage (
company-paid options
+8
Product Operations Analyst I
Product Operations Analyst I

HCSS • Sugar Land (TX)

On-site
USD 60,000 - 76,000
Flexible hybrid schedule
Medical, dental, and vision coverage
401(k) with company match
+4
Enterprise Account Executive III
Enterprise Account Executive III

HCSS • Sugar Land (TX)

On-site
USD 90,000 - 150,000
Flexibility to work Remotely
Medical, dental, and vision coverage
Paid holidays, sick days, PTO
+4
Managing Legal & Compliance Counsel
Managing Legal & Compliance Counsel

HCSS • Houston (TX)

On-site
USD 150,000 - 190,000
Flexible hybrid schedule
Medical, dental, and vision coverage
Paid holidays & time off
+5
Managing Legal & Compliance Counsel
Managing Legal & Compliance Counsel

HCSS • Sugar Land (TX)

On-site
USD 180,000 - 260,000
Flexible hybrid schedule
Medical, dental, and vision coverage
Paid holidays, sick days, and personal
+2
Senior Analyst, Corporate FP&A
Senior Analyst, Corporate FP&A

HCSS • Sugar Land (TX)

On-site
USD 110,000 - 160,000
Remote work flexibility
Medical, dental, vision coverage
Paid holidays & PTO
+4
Managing Legal & Compliance Counsel
Managing Legal & Compliance Counsel

Heavy Construction Systems Specialists • Sugar Land (TX)

On-site
USD 180,000 - 240,000
Flexible hybrid schedule
Medical, dental, and vision coverage
Paid holidays, sick days, and personal
+2