Site Reliability Engineer

HCLTech Germany

Deutschland

Vor Ort

EUR 90.000 - 120.000

Vollzeit

Vor 13 Tagen
Bewerbungsgenerator

Eine maßgeschneiderte Bewerbung für diese Stelle — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

HCLTech Germany is seeking a Site Reliability Engineer (24x7 Operational Support) to strengthen platform operations with a focus on observability, secure logging, and automation.

You will manage Kubernetes, CI/CD pipelines, and Elastic Stack, participate in 24x7 on-call rotations, and help improve SOPs and security compliance across the stack.

Qualifikationen

  • Strong Linux and Kubernetes knowledge is expected.
  • Experience with observability platforms such as Prometheus and Grafana is required.
  • Proficiency in Python, Go, or Bash scripting is essential.
  • Experience with Elasticsearch and/or OpenSearch.
  • Familiarity with CI/CD tooling (Helm/Jenkins/ArgoCD).
  • Willingness to work 24x7 on-call shifts including weekends and holidays.

Aufgaben

  • Platform Engineering & DevOps: manage Kubernetes, Helm, CI/CD pipelines, and IaC.
  • Observability & Monitoring: maintain Prometheus configs, Grafana dashboards, and Thanos.
  • Elastic Stack Ops: configure Elasticsearch/Logstash/Kibana for secure log processing.
  • Incident Response: participate in 24x7 on-call rotations and MIM activities.
  • Secure Operations: ensure security/compliance and maintain SOPs.

Kenntnisse

Linux concepts
Kubernetes environments
Networking fundamentals
REST APIs
Python/Go/Bash
Git workflows
OpenSearch/Elasticsearch
On-call support

Tools

Helm
Jenkins
ArgoCD
Prometheus
Grafana
Thanos

Jobbeschreibung

We’re Hiring | Site Reliability Engineer (24x7 Operational Support)

We are HCLTech, one of the fastest-growing large tech companies in the world and home to 225,000+ people across 60 countries, supercharging progress through industry-leading capabilities centered around Digital, Engineering and Cloud. The driving force behind that work, our people, are diverse, creative, and passionate, raising the bar for excellence on a regular basis. We, in turn, work hard to bring out the best in them as we strive to help them find their spark and become the best version of themselves that they can be

Mandate Skills: Grafana and ELK stack

Candidates must be based in Germany or willing to relocate to Germany.

About the Role:

We are seeking a Site Reliability Engineer (SRE) with a strong background in observability, secure logging, and automation. The ideal candidate will have hands‑on experience with Elasticsearch and/or Prometheus platforms. This role encompasses critical responsibilities in platform operations, including incident management, execution of scheduled maintenance, and contributing to engineering tasks focused on enhancing system stability. The SRE will also be responsible for adhering to standard operating procedures (SOPs) and actively contributing to their continuous improvement by providing constructive feedback.

Key Responsibilities:
  • Platform Engineering & DevOps: Manage Kubernetes and container orchestration, including Helm chart configurations and CI/CD pipelines (Jenkins, ArgoCD). Develop automation scripts (Python, Bash, Go) and deploy Infrastructure-as-Code (IaC) solutions.
  • Observability, Monitoring & Visualisation: MaintainPrometheus solutions (scrape configurations, alert rules, PromQL queries), administer Thanos and Grafana.
  • Elastic Stack Operations & Log Management: Configure and optimise Elasticsearch clusters, Logstash pipelines, and Kibana dashboards for secure, scalable log processing.
  • Incident Response, Troubleshooting & Collaboration: Participate in 24x7 on‑call rotations for rapid incident response, troubleshoot platform, data and performance issues, and engage in Major Incident Management (MIM).
  • Secure Operations & Compliance: Ensure system operations meet security and data protection requirements, maintainsecure documentation, and manage access control policies.
Qualifications, Requirements, and Skills
  • Strong grasp of Linux concepts, preferably in Kubernetes environments.
  • Solid understanding of networking fundamentals and REST APIs.
  • Proficiency in Python, Go, or Bash.
  • Proficiency in Git-based configuration management workflows.
  • Familiarity with CI/CD tools like Helm, Jenkins, or ArgoCD.
  • Experience with Elasticsearch and/or OpenSearch.
  • Willingness to work shift‑based 24x7 on‑call support, including weekends and holidays.
Preferred Certifications:

Elastic Certified Engineer, LPIC Level 2, Kubernetes Administrator.

This position requires eligibility for and successful completion of a German Ü2 Security Clearance. Please ensure you are willing and able to complete the required background screening process as a condition of employment.

____________________________

We promote equal opportunities for all employees, regardless of their cultural and social background, gender, disability, age, religion, beliefs, and sexual identity. We give priority consideration to severely disabled applicants and those of equal status in the case of equal suitability.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Site Reliability Engineer
Site Reliability Engineer

Reysion Technologies • Dresden

Vor Ort
EUR 90.000 - 130.000
Site Reliability Engineer
Site Reliability Engineer

Helsing • München

Vor Ort
Confidential
Focus on outcomes, not time-tracking
Competitive compensation and stock options
Relocation support
+3
Site Reliability Engineer
Site Reliability Engineer

Helsing • Berlin

Vor Ort
EUR 70.000 - 90.000
Competitive compensation
Stock options
Relocation support
+2
Senior Site Reliability Engineer / Kubernetes
Senior Site Reliability Engineer / Kubernetes

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 120.000
Sovereign Cloud Engineer (m/w/d)
Sovereign Cloud Engineer (m/w/d)

United States Digital Space LLC • Walldorf

Vor Ort
EUR 90.000 - 130.000
Remote work
Hybrid work
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Berlin

Vor Ort
EUR 110.000 - 150.000
Hybrid work model
Flexible work policy
AI-driven environment
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FactFinder • Berlin

Vor Ort
Confidential
Hybrid work model
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Pforzheim

Vor Ort
EUR 90.000 - 125.000
Hybrid work model
Staff Site Reliability Engineer (f/m/d)
Staff Site Reliability Engineer (f/m/d)

IONOS • Berlin

Vor Ort
EUR 85.000 - 110.000
Canteen subsidy
Free drinks
Employee discounts
+3
Site Reliability Engineer
Site Reliability Engineer

Apprize Technology Solutions • Deutschland

Vor Ort
EUR 70.000 - 90.000