Site Reliability Engineer

Thales

Austin (TX)

Hybrid

USD 110,000 - 183,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, Dental, Vision coverage
Retirement Savings Plan
Paid Time Off
Life Insurance & AD&D
Well-being Program

Job summary

Thales seeks a Site Reliability Engineer in Austin to ensure high service levels for a cloud-based telecom solution. You will own availability, design scalable infra with Terraform/Ansible/Kubernetes, and run 24/7 on-call incident response in a hybrid setting.

You will define SLOs/SLIs, manage error budgets, and collaborate with security teams on best practices while improving monitoring with Datadog and postmortems for continuous improvement.

Qualifications

  • Engineer or equivalent.
  • At least 5 years of experience.
  • Experience with public cloud (GCP/AWS) and microservices architecture.

Responsibilities

  • Design, build, and maintain scalable infrastructure using Terraform, Ansible, and Kubernetes.
  • Develop automated CI/CD pipelines via GitLab to reduce manual toil.
  • Define and monitor SLOs/SLIs and manage error budgets.
  • Participate in 24/7 on-call rotations and perform deep-dive incident troubleshooting.
  • Perform performance analysis and capacity planning for growth.
  • Implement observability with Datadog and refine monitoring strategies.
  • Lead blameless postmortems and drive long-term fixes.
  • Collaborate with Cloud Security teams on best practices and access controls.
  • Interface with stakeholders to define solution improvement plans.
  • You will own solution service availability.

Skills

Java development
Public Cloud familiarity
Containers & microservices
CI/CD & automation
NoSQL databases

Education

Engineer or equivalent

Tools

Docker
Kubernetes
Jenkins
GitLab
Helm
NoSQL database

Job description

Location: Austin, United States of America

Thales people architect identity management and data protection solutions at the heart of digital security. Business and governments rely on us to bring trust to the billions of digital interactions they have with people. Our technologies and services help banks exchange funds, people cross borders, energy become smarter and much more. More than 30,000 organizations already rely on us to verify the identities of people and things, grant access to digital services, analyze vast quantities of information and encrypt data to make the connected world more secure.

Austin, TX - Hybrid (3 days a week)

Position Summary

We are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication solution (high availability, strong performance constraints) deployed in the public cloud. This product requires the establishment of a product specific SRE team.

Essential Functions
  • Automation & Infrastructure as Code: Design, build, and maintain scalable infrastructure using tools such as Terraform, Ansible, and Kubernetes. Develop automated CI/CD pipelines via GitLab to reduce manual toil.
  • Availability & Reliability Engineering: Define and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Manage "Error Budgets" to balance the velocity of new features with the stability of the platform.
  • Incident Management & On-Call Support: Participate in 24/7 on-call rotations to provide emergency response and perform deep-dive troubleshooting for production issues.
  • Performance & Capacity Planning: Conduct system performance analysis, identify bottlenecks, and perform capacity planning to ensure the infrastructure can handle growth and peak loads.
  • Observability & Monitoring: Implement and refine symptom-based alerting and comprehensive monitoring strategies using platforms like Datadog to ensure high visibility into system health.
  • Continuous Improvement & Postmortems: Lead blameless postmortems after incidents to identify root causes and implement long-term technical fixes to prevent recurrence.
  • Security & Compliance Collaboration: Partner with Cloud Security teams to implement security best practices, manage access controls, and respond to security breaches or vulnerabilities.
  • Support customer relationship
  • Interface with other stakeholders to define solution improvement plan
  • You will have the ownership of solution service availability.
Minimum Requirements
Education
  • Engineer or equivalent
Experience
  • at least 5 years of experience
Skills and Abilities
  • Java development skill is required.
  • You are familiar with Public Cloud (GCP, AWS), containers and microservices (Docker, Kubernetes, Java), CI/CD and automation (Jenkins, Gitlab, Helm), NoSQL database.
Certification
  • GCP cloud architect certification is a plus
Preferred Qualifications
  • You have already set up product monitoring and the underlying infrastructure
  • You have development experience in a distributed systems and/or high availability context
  • You are familiar with microservices development
  • You participated in the definition of architectures, data structures, algorithms with performance, security, reliability constraints, etc.
  • Public cloud architect certification
  • You are interested in aspects of Site Reliability Engineer: CI/CD, automation, monitoring and observability, and continuous improvement.
  • You are an accomplished, versatile and multi-tasking developer engineer.
Must have U.S. or Dual Citizenship and be able to obtain post-hire clearance from the Committee on Foreign Investments in the U.S. (CFIUS) and Department of Treasury

This position will require successfully completing a post-offer background check. Qualified candidates with [a] criminal history will be considered and are not automatically disqualified, consistent with federal law, state law, and local ordinances.

Thales champions inclusion and we believe diversity strengthens the fabric of our culture. Thales is an Equal Opportunity Employer, including disability/veterans.

If you need an accommodation or assistance in order to apply for a position with Thales, please contact us at talentacquisition@us.thalesgroup.com.

The reference Total Target Compensation (TTC) market range for this position, inclusive of annual base salary and the variable compensation target, is between

Total Target Cash (TTC) 109,653.00 - 182,755.00 USD Annual

This reflects how companies in a similar industry and geographic region generally pay for similar jobs. This range helps the Company make pay decisions as one data point among many. Where a position falls within this range is also dependent on other factors including – but not limited to – the employee’s career path history, competencies, skills and performance, as well as the company’s annual salary budget, the customer’s program requirements, and the company’s internal equity. Thales may offer additional benefits and other compensation, depending on circumstances not related to an applicant’s status protected by local, state, or federal law.

Thales provides an extensive benefits program for all full-time employees working 30 or more hours per week and their eligible dependents, including the following:

  • Elective Health, Dental, Vision, FSA/HSA, Voluntary Life and AD&D, Whole Group Life w/LTC, Critical Illness, Hospital Indemnity, Accident Insurance, Legal Plan, Identity Theft, and Pet Insurance
  • Retirement Savings Plan after 30 days of employment with a company contribution and a match, and with no vesting period
  • Company paid holidays and Paid Time Off
  • Company provided Life Insurance, AD&D, Disability, Employee Assistance Plan, and Well-being Program
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Relha LLC • Austin (TX)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Thales Group • Austin (TX)

Hybrid
USD 120,000 - 180,000
Hybrid work model
PTO & holidays
Retirement plan
DevOps System Engineer
DevOps System Engineer

Thales • United States

On-site
USD 109,000 - 183,000
Health, Dental, Vision Insurance
Retirement Savings Plan with company match
Paid Time Off and Holidays
Senior Service Delivery Manager
Senior Service Delivery Manager

Thales • United States

On-site
USD 124,743 - 208,378
Senior Software Engineer (Hands on, Tech Lead)
Senior Software Engineer (Hands on, Tech Lead)

Thales • Austin (TX)

On-site
USD 139,000 - 234,000
Health/Dental/Vision benefits
Retirement Savings Plan
Paid Time Off
Cloud Security Engineer
Cloud Security Engineer

Thales • Austin (TX)

On-site
USD 105,000 - 184,000
Health, Dental, and Vision Insurance
Retirement Savings Plan
Paid Time Off
Sr Engineer, Advanced Security Response Team (ASRT)
Sr Engineer, Advanced Security Response Team (ASRT)

Thales • Chandler (AZ)

On-site
USD 140,000 - 190,000
Comprehensive health benefits
Retirement savings plan with company  
Paid time off
Regional Sales Manager
Regional Sales Manager

Thales • Austin (TX)

Hybrid
USD 148,000 - 290,000
Health, Dental, Vision benefits
Retirement Savings Plan with match
Paid Time Off
Administrative Assistant Project Coordinator
Administrative Assistant Project Coordinator

Thales • Austin (TX)

Hybrid
USD 59,000 - 103,000
Health insurance
Retirement plan
Paid time off
+1
Test Technician
Test Technician

Thales • Salt Lake City (UT)

On-site
USD 49,000 - 91,000
Health Insurance
Retirement Savings Plan
Paid Time Off