Site Reliability Engineer (Splunk, Python, OCI, Dynatrace, RCA, Terraform, Ansible )

NEPTUNEZ SINGAPORE PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NEPTUNEZ SINGAPORE PTE. LTD. is seeking an experienced Site Reliability Engineer to manage scalable production environments, lead incident management, and build automation using Python and Shell scripting.

The role requires strong observability expertise with Splunk and experience administering Linux infrastructure. You will collaborate with development and infrastructure teams to improve reliability and deployment processes, while driving automation to reduce manual effort and improve SLA

Qualifications

  • Bachelor's degree in CS/IT or related field.
  • 8+ years in SRE/Production Support/Infra Ops.
  • Strong Python and Shell scripting experience.
  • Experience with Dynatrace and AppDynamics.
  • Splunk experience for monitoring and observability.
  • Infra admin and troubleshooting Linux environments.
  • IaC with Terraform and configuration management with Ansible.
  • SQL and app performance tuning preferred.

Responsibilities

  • Manage highly available, scalable production environments.
  • Lead incident management, RCA, and problem management.
  • Develop automation with Python and Shell scripting.
  • Implement monitoring and observability with Splunk and others.
  • Administer and troubleshoot Linux servers and production infra.
  • Provide L2/L3 production support within SLA.
  • Automate provisioning and config management using IaC tools.
  • Collaborate with development and infra teams to improve reliability.
  • Identify process improvements and automate repetitive tasks.
  • Create and maintain runbooks and production best practices.

Skills

Python automation
Shell scripting
Analytical skills
Troubleshooting
Communication
Collaboration

Education

Bachelor's degree in Computer Science or IT

Tools

Dynatrace
AppDynamics
Splunk
Terraform
Ansible

Job description

Responsibilities


  • Manage and maintain highly available, scalable, and reliable production environments.

  • Lead incident management, root cause analysis (RCA), and problem management activities to ensure service stability.

  • Develop and maintain automation solutions using Python and Shell scripting to improve operational efficiency.

  • Implement and enhance monitoring, logging, and observability using Splunk and other enterprise monitoring tools.

  • Administer, troubleshoot, and optimize Linux servers and production infrastructure.

  • Provide L2/L3 production support and resolve critical application and infrastructure issues within SLA.

  • Automate infrastructure provisioning and configuration management using Infrastructure-as-Code tools.

  • Collaborate with development and infrastructure teams to improve system reliability, performance, and deployment processes.

  • Identify opportunities for process improvement and implement automation to reduce manual effort.

  • Create and maintain operational documentation, runbooks, and best practices for production support.


Requirements


  • Bachelor's degree in Computer Science, Information Technology, or a related field.

  • 8+ years of experience in Site Reliability Engineering (SRE), Production Support, or Infrastructure Operations.

  • Strong hands-on experience with Python automation and Shell scripting.

  • Strong experience in Dynatrace, AppDynamics

  • Hands on experience in Splunk for monitoring, log analysis, troubleshooting, and observability.

  • Strong experience in administering and troubleshooting Linux environments.

  • Experience in incident management, RCA, problem management, and production support for mission-critical systems.

  • Experience with cloud platforms such as AWS and/or Oracle Cloud Infrastructure (OCI).

  • Hands-on experience with Infrastructure-as-Code tools such as Terraform and configuration management tools like Ansible.

  • Experience in SQL, and application performance tuning is preferred.

  • Excellent analytical, troubleshooting, communication, and collaboration skills.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer(Senior SRE)
Site Reliability Engineer(Senior SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 200,000
Site Reliability Engineer - Data Availability
Site Reliability Engineer - Data Availability

SIX • Singapore

Hybrid
SGD 100,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

IDEMIA Public Security • Singapore

On-site
SGD 120,000 - 180,000
Sr. SRE
Sr. SRE

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 180,000
On-site in Singapore (3 days/wk)
Senior SRE: Dynatrace, Splunk, Python Automation & OCI
Senior SRE: Dynatrace, Splunk, Python Automation & OCI

NEPTUNEZ SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

VANGUARD SOFTWARE PTE. LTD. • Singapore

On-site
SGD 100,000 - 150,000
Technical Leadership
Career Growth
High-Performance Team
+1
Software Engineer/ Site Reliability Engineer
Software Engineer/ Site Reliability Engineer

United States Digital Space LLC • Singapore

On-site
SGD 90,000 - 150,000
Senior SRE Engineer
Senior SRE Engineer

JONDAVIDSON PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

United States Digital Space LLC • Singapore

On-site
SGD 100,000 - 130,000