Site Reliability Engineer - Service Assurance Systems

Viasat

Batam

On-site

IDR 300,000,000 - 700,000,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Viasat is seeking a Site Reliability Engineer to bridge software development and operational reliability in the SAS group. You will manage deployments, investigate incidents, and drive DevOps practices across cloud and on-premise services.

You will collaborate with developers and operations teams to improve resilience, automate tasks, and maintain clear operational documentation and runbooks.

Qualifications

  • Familiarity with Windows and Linux operating systems and command-line tools.
  • Experience with SQL or data pipelines and ETL concepts.
  • Hands-on experience with CI/CD, monitoring and cloud services.
  • Proficiency in Python for automation and scripting.
  • Ability to communicate clearly with technical and non-technical audiences.
  • Experience using ITSM/ServiceNow for incident and change management.
  • Understanding of security best practices in operations.

Responsibilities

  • Investigate and resolve incidents and service requests within SLAs.
  • Manage deployments across development, staging, and production environments.
  • Monitor health using Prometheus, CloudWatch; detect anomalies and respond promptly.
  • Own CI/CD pipelines and improve automation for builds and deployments.
  • Promote DevOps practices and infrastructure-as-code within the SAS group.
  • Manage containerised workloads with Docker and AWS services.
  • Write operational scripts (Python) to automate tasks and incidents.
  • Collaborate with developers to address recurring operational issues.
  • Maintain runbooks, deployment guides and post-mortems.

Skills

Windows & Linux familiarity
SQL or data pipelines
Airflow/Flink/Dataflow
CI/CD experience
AWS services
Docker containers
Python scripting
Communication skills
ITSM/ServiceNow
Proactive problem-solving
Security best practices

Tools

Jenkins
GitLab CI
GitHub Actions
Terraform
Ansible
Prometheus
AWS CloudWatch
Docker
ServiceNow

Job description

About us

One team. Global challenges. Infinite opportunities. At Viasat, we’re on a mission to deliver connections with the capacity to change the world. For more than 35 years, Viasat has helped shape how consumers, businesses, governments and militaries around the globe communicate. We’re looking for people who think big, act fearlessly, and create an inclusive environment that drives positive impact to join our team.

What you'll do

The Service Assurance Systems (SAS) Group, part of Global Operations, develop and maintain many software systems and applications which support the operation of Viasat services.

The team collects assurance and operational data from a wide range of systems across the Viasat estate. This data is shared between internal platforms and distributed through Google Cloud Platform (GCP), where it underpins observability and monitoring capabilities that give operations teams real-time insight into service health and performance. To ensure the data is accurate and fit for purpose, the team builds and maintains data pipelines that cleanse, transform and enrich raw data, making it readily consumable by our stakeholders across operations, engineering and management.

As a Site Reliability Engineer (SRE), you will play a key role in bridging the gap between software development and operational reliability. You will be responsible for supporting the applications built and maintained by the SAS group, managing deployments, investigating operational issues, and driving the team’s transition toward a modern DevOps culture. You will act as a first point of contact for service issues raised through the ServiceNow ticketing system, working to diagnose, triage and resolve incidents efficiently while collaborating closely with developers and operations teams.

The day-to-day

You will be working as part of a small team of developers and operations engineers, supporting the evolution of Viasat’s network and service monitoring capabilities, ensuring it remains world class in support of existing and future services.

Day-to-day the role will involve:

  • Investigate and resolve incidents, service requests and problems raised through the ServiceNow ticketing system, ensuring timely and thorough resolution within agreed SLAs.
  • Manage and oversee application deployments across development, staging and production environments, ensuring smooth and reliable release processes.
  • Monitor application and infrastructure health using observability tools such as Prometheus and AWS CloudWatch; proactively identify and respond to anomalies and performance degradation.
  • Own and maintain CI/CD pipelines, working to improve build, test and deployment automation to reduce manual effort and increase release confidence.
  • Champion and drive the adoption of DevOps practices and culture within the SAS group, working towards greater automation, infrastructure-as-code, and operational maturity.
  • Manage and maintain containerised workloads using Docker, and support the operation of services hosted on AWS cloud infrastructure.
  • Write and maintain operational scripts (primarily Python) to automate routine tasks, support incident investigations and improve team efficiency.
  • Collaborate with software developers to identify recurring operational issues and feed findings back into the development process to improve application resilience.
  • Maintain clear and up-to-date operational documentation including runbooks, deployment guides and incident post-mortems.
  • Provide written and verbal progress updates on open incidents, deployments and operational improvements to the SAS group and wider stakeholders.
  • Support on-call and out-of-hours incident response as required.
  • Liaise with engineering and infrastructure teams to ensure system changes are communicated and operationally risk-assessed before deployment.
What you’ll need
  • Familiarity with both Windows and Linux operating systems, including command-line administration and troubleshooting.
  • Demonstrable experience with SQL or data pipeline technologies.
  • Experience with data processing frameworks such as Apache Airflow, Apache Flink or Google Dataflow, with a practical understanding of ETL pipeline design and the ability to diagnose issues across data transformation and scheduling workflows.
  • Demonstrable CI/CD experience - building, maintaining and improving pipelines using tools such as Jenkins, GitLab CI, GitHub Actions or equivalent.
  • Hands-on experience with AWS services (e.g. EC2, S3, ECS, Lambda, CloudWatch) and containerisation using Docker.
  • Experience with monitoring and observability tooling, including Prometheus and AWS CloudWatch, for metrics collection, alerting and dashboarding.
  • Proficiency in scripting with Python for automation, tooling and operational support tasks.
  • Strong communication skills — able to clearly articulate technical issues and their impact to both technical and non-technical audiences, and to provide timely updates on incident progress.
  • Experience using an ITSM/ticketing system (e.g. ServiceNow) for incident management, request fulfilment and problem tracking.
  • A proactive, solution-oriented approach with strong attention to detail and a commitment to operational excellence.
  • A reasonable understanding and appreciation of IT and security standard methodologies in an operational environment.
What will help you on the job
  • Understanding of TCP/IP networking principles, including the ability to diagnose connectivity issues and interpret network traffic.
  • Familiarity with event streaming platforms such as Apache Kafka, including an understanding of producer/consumer patterns and how message queues are used to move data reliably between systems at scale.
  • Familiarity with infrastructure-as-code tools such as Terraform or Ansible.
  • Experience with log aggregation and analysis platforms such as the OTEL Stack or AWS CloudWatch Logs Insights.
  • Exposure to Kubernetes or other container orchestration platforms.
  • Experience working in an Agile or DevOps team environment.
EEO Statement

Viasat is proud to be an equal opportunity employer, seeking to create a welcoming and diverse environment. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, ancestry, physical or mental disability, medical condition, marital status, genetics, age, or veteran status or any other applicable legally protected status or characteristic. If you would like to request an accommodation on the basis of disability for completing this on-line application.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud SRE & DevOps Engineer
Cloud SRE & DevOps Engineer

Viasat • Batam

On-site
IDR 300,000,000 - 700,000,000
NOC Controller Intern
NOC Controller Intern

Viasat • Batam

On-site
Housing assistance
Relocation assistance
Benefits overview
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
Site Reliability Engineer
Site Reliability Engineer

PARTECH PARTNERS • Daerah Khusus Ibukota Jakarta

On-site
IDR 272,380,000 - 453,968,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
Senior Site Reliability Engineer
Senior Site Reliability Engineer

StraitsX • Indonesia

On-site
IDR 300,000,000 - 540,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

StraitsX • Kebayoran Baru

On-site
IDR 600,000,000 - 1,000,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AccelByte • Kota Yogyakarta

On-site
IDR 300,000,000 - 600,000,000
Site Reliability Engineer
Site Reliability Engineer

Macquarie Group • Indonesia

On-site
IDR 400,000,000 - 650,000,000
Senior Site Reliability Engineer Associate
Senior Site Reliability Engineer Associate

PT Alto Network • Jawa Tengah

On-site
IDR 180,000,000 - 300,000,000