Sr Software Engineer (Site Reliability) Austin or Dallas, TX

H-E-B

San Antonio (TX)

Hybrid

USD 140,000 - 180,000

Full time

37 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

H-E-B seeks a Senior Software Engineer-Site Reliability on the Digital Fulfillment team in a hybrid environment. You will support build/deploy pipelines, diagnose production issues, and influence system design for reliability and performance.

You will work on Kubernetes, cloud infra (AWS/GCP), and CI/CD pipelines, collaborating across engineering teams to ensure scalable services.

Qualifications

  • 5+ years designing, analyzing, developing, or troubleshooting distributed systems.
  • 3+ years SRE experience with cloud environments (GKE, AWS).
  • 2+ years Java (Spring) programming experience preferred.
  • 3+ years Terraform for cloud infra.
  • 3+ years CI/CD pipelines (GitLab/GitHub).
  • Experience with GitLab, JIRA, Slack, Confluence, IntelliJ preferred.
  • Experience with microservices architectures.
  • Experience with PostgreSQL, Kubernetes, Docker, Linux, GCP, Terraform, REST/GraphQL.
  • Monitoring/visualization tools: Datadog, Grafana, New Relic.
  • Strong scripting: Python, Ruby, Groovy, Bash.
  • Proven HA/Scalability principles application.

Responsibilities

  • Develops and maintains tooling for environment monitoring and task automation.
  • Contributes to lifecycle of services from inception to deployment and operation.
  • Optimize configurations for software, servers, DB connections/indexes and drivers.
  • Design service architectures, capacity planning, release plans and launches.
  • Monitor SLOs/SLAs and identify gaps; work with teams to resolve.
  • Serve as SME for cross-functional engineering teams and troubleshoot systems.

Skills

Distributed systems
SRE principles
On-call support
Python/Ruby/Groovy/Bash
System monitoring
Communication
Agile / cross-team

Education

Computer Science degree

Tools

Kubernetes
Google Kubernetes Engine
AWS
Terraform
CI/CD (GitLab/GitHub Actions)
Docker
PostgreSQL
Grafana/Datadog/New Relic
Linux
REST/GraphQL APIs

Job description

Responsibilities

Job Summary: As a Senior Software Engineer-Site Reliability on the Digital Fullfilment team, you'll deliver complex code solutions. You'll support the build and deployment pipeline and when necessary, diagnose / solve production support or on-call issues. You'll contribute to overall system design, architecture, security, scalability, reliability, application performance and provide end-to-end support.

Location

Austin (preferred), open to Dallas, TX (Hybrid)

Key Responsibilities & Essential Functions
  • Develops and maintains tooling used for environment monitoring and task automation
  • Engages in and improves whole lifecycle of services, including inception and design, deployment, operation, and refinement
  • Analyses and establishes efficient configurations for software and servers, DB connections / indexes, drivers, etc.
  • Collaborates with development teams to design service architectures, software platforms and frameworks, capacity planning, release plans and launch reviews
  • Monitors internal and vendor service level objectives (SLOs) and agreements (SLAs); identifies / resolves SLO / SLA gaps
  • Serves as technical subject matter expert (SME) for cross-functional engineering Teams; assists with / troubleshoots systems-related issues and maintenance

The responsibilities and essential functions outlined above describe the general nature and level of work assigned to this position. This is not an exhaustive list of all duties, responsibilities, and skills required. Duties and responsibilities may be modified at any time based on business needs. Employees may be required to perform other job-related tasks as requested by their supervisor, subject to reasonable accommodations.

Work Experience

Qualifications & Key Requirements:

  • 5+ years experience designing, analyzing, developing, or troubleshooting distributed systems
  • 3+ years of SRE experience managing Google Kubernetes Engine (preferred), K8s, or AWS environments
  • 2+ years of Java (Spring) programming experience preferred
  • 3+ years of using Terraform to maintain cloud infrastructure
  • 3+ years of CI Pipeline experience with either Gitlab Pipelines, or GitHub Actions
  • Experience with tools such as Gitlab, JIRA, Slack, Confluence and Intellij is preferred
  • Experience with microservices architecture patterns
  • Experience working with PostgreSQL, Kubernetes, Docker, Linux, GCP, Terraform, and APIs using REST and GraphQL
  • Experience working with monitoring and visualization tools such as Datadog, Grafana, or New Relic
  • Strong proficiency with scripting languages such as Python, Ruby, Groovy, Bash
  • Proven track record of researching, understanding, and effectively applying Scalability and High Availability principles
Knowledge/Skills/Abilities
  • Advanced knowledge in system and data architecture, data modeling, and design and capable of architecting and designing at the application or service level using well-accepted design patterns -
  • Able to review platform designs for strength of engineering solutions, namely performance, sustainability, and iterative development potential. -
  • Comprehensive knowledge of Computer Science fundamentals: data structures, algorithms, design patterns, system architecture and design patterns -
  • Advanced understanding of development methodologies and processes -
  • High degree of personal accountability to self and team for continued growth -
  • Adjust - Leverages Agile metrics to improve team performance and deliverables. Evaluates and adjusts resources, self, and team as necessary. -
  • Collaborate - Ability to work on tasks which span multiple domains, requiring cross-team collaboration, which have a high impact on your project. -
  • Agility - Embraces risk, change, and helps team manage ambiguity within the team's scope of work. -
  • Able to drive progress without having a complete picture and can articulate potential tradeoffs and prioritize when faced with ambiguity. -
  • Connect - Delivers clear, concise, effective messages across different levels; can tailor communication based on intended audience. -
  • Growth Mindset - Fosters a culture of mentoring and coaching across multiple technical teams and other stakeholders. -
  • Relate - Fosters a culture within their team where people are encouraged to share their opinions and contribute to discussions in a respectful manner, approach disagreement non-defensively with inquisitiveness, and use contradictory opinions as a basis for constructive, productive conversations. -
Education
  • A Computer Science degree or comparable formal training, certification, or work experience -involving software / systems engineering
Physical Demands & Working Conditions
  • Travel by car or plane with overnight stays
  • Work extended hours; sit for extended periods
  • Work rotating and on-call schedules, as needed

The work environment characteristics described here are representative of those a Partner encounters while performing the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.

JDENGINEERING

DEV3232

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr Staff Software Engineer
Sr Staff Software Engineer

H.E.B. • San Antonio (TX)

On-site
USD 150,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

On-site
USD 120,000 - 155,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

Hybrid
USD 120,000 - 150,000
Systems Analyst 3 529601671
Systems Analyst 3 529601671

LMG Technology Services LLC • Austin (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
SRE - Site Reliability Engineer - Senior
SRE - Site Reliability Engineer - Senior

ManpowerGroup Global, Inc. • Austin (TX)

On-site
USD 66,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
Software Engineer II - Reliability Engineering Tooling (Remote)
Software Engineer II - Reliability Engineering Tooling (Remote)

The Home Depot • Atlanta (GA)

On-site
USD 100,000 - 140,000