Site Reliability engineering (SRE)

TechDigital Group

San Leandro (CA)

On-site

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

An established industry player is seeking a skilled Site Reliability Engineer with a robust Java development background. In this pivotal role, you will be part of a dedicated SRE team, utilizing cutting-edge tools to enhance the performance and availability of digital platforms. Your expertise in building dashboards, automating processes, and supporting web/API platforms will be crucial in driving operational excellence. This role offers the opportunity to work with a variety of technologies, influence DevOps practices, and contribute to a culture of continuous improvement. If you are passionate about technology and thrive in a dynamic environment, this position is perfect for you.

Qualifications

  • 10+ years in Software Engineering with a focus on SRE and platform health.
  • Hands-on experience with monitoring tools and cloud infrastructure.

Responsibilities

  • Support and maintain the resiliency and performance of digital platforms.
  • Collaborate with engineering teams to resolve production outages.

Skills

Java Development
Site Reliability Engineering
Agile Practices
Automated Testing
Process Automation
Shell Scripting
DevOps Tools
Cloud Infrastructure
Monitoring Tools
API Development

Education

Bachelor's Degree in Computer Science or related field
Equivalent experience in Software Engineering

Tools

Splunk
Grafana
GCL
ELK
Prometheus
Jenkins
Git
Kubernetes
AWS
Azure

Job description

Need SRE candidate with good Java Dev background interested in this role with strong hands-on experience in building dashboards and setting up alerts using Splunk, Grafana and GCL.


Required Qualifications:

  • 10+ years of Software Engineering experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
  • 10+ years of experience in Production support/Site Reliability Engineering teams with continued focus on improving Platform health
  • Familiar with Agile or other rapid application development practices
  • Hands-on expertise with Automated testing, Process Automation & building dashboards using APM tools.
  • Experience with distributed (multi-tiered) systems, algorithms, relational databases, and NoSQL databases.
  • Knowledge & Exposure caching tools (Redis, memcache) or messaging tools such as MQ, Kafka.
  • Must have working knowledge of APM tools such as splunk, GCL, ELK, Grafana, Prometheus etc.
  • Able to create Dashboards using GCL/Splunk/ELK and setup alerts.
  • Working knowledge of CICD is a plus – Source control like Git, Continuous Integration – Jenkins / UCD Release etc.
  • Ability to work with Engineering teams across the ecosystem such as Security, Networking & Infrastructure challenges which can impact platform health & resiliency.
  • Shell Scripting / DevOps tools like Ansible with good knowledge of yaml file to write playbooks.
  • Experience with distributed storage technologies like NFS as well as dynamic resource management frameworks PCF, Kubernetes / OpenShift, AWS or Azure.
  • Tech Stack: Java/J2EE (Spring, Spring Boot, Python, Shell Scripting, Kafka, Oracle, MongoDB etc.).
  • Able to work on shift duty in a 12/7 support organization.

Job Expectations:

  • You will be a core member of a SRE support team, utilizing the latest technology tools to write code, test cases, working with API specs and automate to maintain the resiliency, performance and availability of Digital Sales & Marketing platforms.
  • Strong & relevant experience in supporting Web/API platforms built using Java/java script Stack (Spring/Spring boot, Javascript -Angular/react)
  • Proficiency in dealing with Legacy infrastructure along with cloud infrastructure (on prem & 3rd party) such as PCF or Azure.
  • Identifying opportunities to adopt to new technologies while improving efficiency by removing toil and continues to drive efficiency & optimization.
  • Proactive monitoring of app performance through Splunk, App dashboards, App dynamics & Dynatrace etc.
  • Represent Platform engineering teams during production outages and collaborate with engineering teams to resolve production outages. Collaborate with stakeholders across engineering functions to own/derive RCA & work towards permanent resolution.
  • Plan, support, execute and comply with governance programs/processes in support of a strong control environment in your functional area. Leverage process documentation to improve operational controls and identify and remediate process deficiencies.
  • Proactively identify, communicate, mitigate and escalate risk originating from non-compliance of processes, operational errors, and data integrity issues in all applicable processes.
  • Ability to influence SRE practices within and outside teams to enable a strong DevOps culture within the organization.
  • Able to work on shift duty in a 12/7 support organization.
  • Responsible for working with Engineering teams to maintain the SLAs & SLOs. Constantly looking out for opportunities to improve platform metrics & communicate the same to stakeholders.
  • Exposure and proficiency in different API styles such as SOAP, REST, Micro services etc.
  • Working knowledge of Unix, Linux and Postman.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead
SRE Lead

TechDigital Group • San Leandro (CA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

Matlen Silver • Charlotte (NC)

Hybrid
USD 90,000 - 94,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

State of Wisconsin Investment Board • Madison (WI)

On-site
USD 150,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000