Principal Reliability Engineer

GoTo Foods, LLC

Blythe (GA)

On-site

USD 180,000 - 240,000

Full time

28 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

GoTo Foods, LLC is building an industry-leading Digital Platform powering seven brands and enabling future growth. The Principal Site Reliability Engineer is an IC role owning enterprise-wide reliability, scalability, and performance of critical production services, shaping architectural standards and guiding infrastructure changes.

The role requires strong communication to explain technical solutions to diverse partners and close collaboration with software engineers, analysts, and architects

Qualifications

  • Bachelor’s degree in Information Systems, Computer Science, or a related field.

Responsibilities

  • Lead resiliency and capacity planning for a high-performance digital platform.
  • Drive reliability, scalability, and operational excellence for critical user-facing systems.
  • Identify systemic risks and bottlenecks across services, dependencies, deployments, and infrastructure.
  • Automate repetitive operational tasks and improve deployment safety and recovery workflows.
  • Establish best practices around observability, incident response, and postmortems.
  • Collaborate with engineers and architects to plan, design, and maintain enterprise platform services.

Skills

Distributed systems
Linux
Cloud native
Typescript
Python
Observability
Terraform
Kubernetes
Azure
AWS
GCP
CI/CD
Git
Jira
Confluence

Education

Bachelor’s degree in Information Systems/CS
Master’s degree (preferred)

Tools

Terraform
Kubernetes
Azure
AWS
GCP
Datadog/Prometheus/Grafana" observability tools
Git

Job description

Job Summary

GoTo Foods is on a journey to build out an industry leading Digital Platform which will power its 7 existing brands and enable smooth integration of future brands. The Principal Site Reliability Engineer is an Individual Contributor role owning the enterprise-wide reliability, scalability, and performance of our critical production services. As a foundational pillar of our engineering organization, this role drives architectural standards for the full service lifecycle and will help us scale our backend, raise our reliability and performance bar by providing sharp technical judgment on infrastructure changes. To be successful, the candidate will require excellent communication skills, and the ability to explain technology solutions with diverse partners.

Job Summary

GoTo Foods is on a journey to build out an industry leading Digital Platform which will power its 7 existing brands and enable smooth integration of future brands. The Principal Site Reliability Engineer is an Individual Contributor role owning the enterprise-wide reliability, scalability, and performance of our critical production services. As a foundational pillar of our engineering organization, this role drives architectural standards for the full service lifecycle and will help us scale our backend, raise our reliability and performance bar by providing sharp technical judgment on infrastructure changes. To be successful, the candidate will require excellent communication skills, and the ability to explain technology solutions with diverse partners.

Essential Functions
  • Lead the resiliency and capacity planning of high-performance Digital Platform for online ordering across 7 brands processing 50,000+ online orders daily.
  • Drive reliability, scalability, and operational excellence for critical user facing systems and services. Improve performance and resiliency across APIs, content delivery, and real-time experience.
  • Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure.
  • Eliminate repetitive operational work through automation and tooling. Build systems that improve deployment safety, remediation workflows, and reliability guardrails.
  • Maintain high standards of software quality within the team by establishing best practices around observability and incident response.
  • Participate in on-call duties and lead complex incident response efforts across engineering teams. Drive blameless postmortems, identify root causes, and ensure sustainable long-term fixes are implemented.
  • Define and champion best practices around reliability engineering, SLIs/SLOs, capacity management, release engineering, and operational maturity across the company.
  • Foster CI/CD practices, automated testing, and deployment pipelines to ensure a smooth development and release process.
  • Collaborate with other software developers, business analysts and software architects to plan, design, develop, test, and maintain enterprise platform services.
  • Provide guidance to engineering teams in support of cloud infrastructure.
  • Prepare reports, manuals and other documentation on the status, operation and maintenance of software and infrastructure.
  • Develop, refine, and tune integrations between applications.
  • Analyze and resolve technical and application problems.
  • Assess opportunities for application and process improvement and prepare documentation of rationale to share with team members and other affected parties.
  • Adhere to high-quality development principles while delivering solutions on-time.
  • Research and evaluate a variety of software products.
Education
  • Bachelor’s degree in Information Systems, Computer Science, or a related field, required.
  • Master’s degree, preferred.
Work Experience
  • 7+ years’ experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems.
  • 5+ years’ experience with cloud platforms like Azure, GCP, or AWS.
  • 5+ years’ experience with Kubernetes.
  • 3+ years’ experience with Splunk, Dynatrace, Datadog, New Relic, Prometheus, Grafana or other observability tool with in-depth understanding of otel.
  • Experience with Terraform or other cloud SDKs.
  • Experience with CDN such as Cloudflare and Azure Front Door.
  • Prior experience in technical leadership, team development, and supporting delivery.
  • In-depth knowledge and experience with developing web applications with service‑oriented framework, microservices and Rest APIs.
  • Extensive experience designing and developing enterprise grade software.
  • Experience with source control management systems like Git and continuous integration/deployment environments.
  • Experience with agile development methodologies including Kanban and Scrum.
  • Experience with debugging, performance profiling, and optimization.
  • Experience improving reliability through SLOs, automation, incident management, and performance optimization.
Skills
  • Deep understanding of one or more: distributed systems, networking, Linux systems, cloud native architectures.
  • Strong programming skills in one or more languages such as Typescript, Python, or similar.
  • Strong understanding of observability systems including metrics, logging, tracing, and alerting.
  • Demonstrated ability to troubleshoot complex issues across applications, infrastructure, networking, and services.
  • Expert level knowledge with Terraform, Kubernetes, Azure.
  • Familiar with AWS, GCP, Node.
  • Deep understanding of authentication strategies.
  • Advanced knowledge of CI/CD practices and ability to understand pipelines.
  • Advanced user of Git, Jira, Confluence, and other supporting tools.
  • Internally motivated, able to work proficiently both independently and in a team environment.
  • Strong communication skills with both internal team members and external business stakeholders.
  • Strong initiative to find ways to improve solutions, systems, and processes.
  • High level working knowledge of SSO, application implementation, administration, SAML Authentication, etc.
  • Ability to communicate complex, technical concepts to business leaders and technical resources in clear concise language; to convey clear, concise information in verbal, written, electronic, and other communication formats; to demonstrate active listening while engaging others; and to articulate ideas and present information to all levels of the organization and varying sizes of audiences.
  • Ability to develop and maintain positive business relationships and foster an environment of mutual respect, understanding, trust, and support.
  • Ability to adapt and adjust planned work through analyzing work demands, competing priorities, and tight deadlines; and to understand the most effective and efficient means to accomplish tasks within the parameters of the organizational structure, processes, systems, and policies.
  • Ability to exercise judgment and discretion in dealing with matters of significance; and to conduct research, analyze data, and arrive at valid conclusions.
  • Ability to conduct research, perform analysis, and communicate results effectively.
  • Ability to anticipate and respond to the needs of stakeholders (e.g., internal, and external customers, etc.) in a timely manner.
  • Comfortable using AI coding tools.
CertificationsTravel Requirement
  • None
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Reliability Engineer
Principal Reliability Engineer

GoToFoods • Augusta (GA)

Hybrid
USD 150,000 - 200,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Prosum • Scottsdale (AZ)

On-site
USD 150,000 - 190,000
Software Engineer Manager - Supply Chain RE (Remote)
Software Engineer Manager - Supply Chain RE (Remote)

The Home Depot • Atlanta (GA)

Remote
USD 180,000 - 240,000
Senior Software Engineer - IP&R Reliability Engineering (Remote)
Senior Software Engineer - IP&R Reliability Engineering (Remote)

The Home Depot • Atlanta (GA)

Remote
USD 120,000 - 180,000
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Staffworxs • Louisville (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Vertafore • Denver (CO)

On-site
USD 160,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package