SRE Monitoring Platform Software Engineer, Entry Level

Jobtailor

Deutschland

Remote

EUR 42.000 - 64.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Mach aus dieser Rolle ein Vorstellungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor is seeking an entry-level Software Engineer to contribute to a multi-region GPU rental platform. You will help observe, protect, and operate the NeoCloud SRE stack, building collection agents, dashboards, and testing suites.

Under senior guidance, you will implement alerts, SLOs, and topology services for Kubernetes and related tooling, while gaining ownership of a sub-context within 12 months and growing your skills across distributed systems and observability.

Qualifikationen

  • 0–2 years of software engineering experience; new graduates with strong projects or internships welcome.
  • Solid fundamentals in Go, Python, Java, or Rust; Go preferred.
  • Ability to write clean, tested, readable code and explain design choices.
  • Data structures, algorithms, concurrency, TCP/HTTP networking, and operating-system fundamentals.
  • Understanding of distributed-systems concepts including idempotency, retries, back-pressure, caching, and eventual consistency.
  • Hands-on exposure to Prometheus, Grafana, Loki, or similar observability tools.
  • Ability to write basic PromQL queries and instrument services.
  • Familiarity with Linux, the shell, system logs, and standard debugging tools.
  • Kubernetes basics, including Pods, Services, and Deployments; experience running something on Kubernetes.
  • Git, branching, pull requests, and CI pipelines such as GitHub Actions or GitLab CI.
  • Unit and integration testing discipline.
  • Clear written and verbal English.
  • Curiosity and willingness to learn GPU/AI infrastructure, AIOps, distributed systems, and observability.
  • Nice-to-have: internship or project experience in monitoring, observability, telemetry pipelines, or platform/SRE tooling
  • Nice-to-have: exposure to GPU/AI infrastructure such as DCGM, InfiniBand/RoCE, Kubernetes GPU Operator, Slurm, or Ray
  • Nice-to-have: exposure to AIOps/ML-adjacent tooling
  • Nice-to-have: contributions to open-source observability or cloud-native projects

Aufgaben

  • Contribute to Bitdeer's NeoCloud SRE platform for observing, protecting, and operating a multi-region GPU rental fleet
  • Build collection agents, metrics/logs/traces/profiles stores, enrichment services, and collection monitors
  • Write ingestion, query, and storage-path code
  • Contribute to alerting, correlation, and SLO frameworks; implement and tune default alert rules
  • Contribute to topology services, cluster-health rollups, and OSS-SRE-tool collection plugins for Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay
  • Help build remediation actuators, orchestration/workflow components, inspection probes, and job schedulers
  • Instrument services with metrics, logs, and traces using OpenTelemetry
  • Build dashboards and write actionable on-call runbooks
  • Write unit, integration, and contract tests for shipped components
  • Participate in chaos and soak tests led by senior engineers
  • Develop components from design through production using GitOps and the CI/CD release pipeline
  • Meet declared SLOs and maintain drift-free systems
  • Operate what you build under senior-engineer guidance
  • Participate in on-call as a shadow before taking primary responsibility
  • Progress toward independently delivering components and owning a sub-context within 12 months

Kenntnisse

Go
Python
Java
Rust
Distributed Systems
Prometheus
Grafana
Loki
OpenTelemetry
Kubernetes
Git
CI/CD
Linux
PromQL
Unit Testing

Tools

Prometheus
Grafana
Loki
OpenTelemetry
Kubernetes
Git
GitHub Actions
GitLab CI
Docker
Linux

Jobbeschreibung

  • Contribute to Bitdeer's NeoCloud SRE platform for observing, protecting, and operating a multi-region GPU rental fleet
  • Build collection agents, metrics/logs/traces/profiles stores, enrichment services, and collection monitors
  • Write ingestion, query, and storage-path code
  • Contribute to alerting, correlation, and SLO frameworks; implement and tune default alert rules
  • Contribute to topology services, cluster-health rollups, and OSS-SRE-tool collection plugins for Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay
  • Help build remediation actuators, orchestration/workflow components, inspection probes, and job schedulers
  • Instrument services with metrics, logs, and traces using OpenTelemetry
  • Build dashboards and write actionable on-call runbooks
  • Write unit, integration, and contract tests for shipped components
  • Participate in chaos and soak tests led by senior engineers
  • Develop components from design through production using GitOps and the CI/CD release pipeline
  • Meet declared SLOs and maintain drift-free systems
  • Operate what you build under senior-engineer guidance
  • Participate in on-call as a shadow before taking primary responsibility
  • Progress toward independently delivering components and owning a sub-context within 12 months
Requirements
  • 0–2 years of software engineering experience; new graduates with strong projects or internships welcome
  • Solid fundamentals in Go, Python, Java, or Rust; Go preferred
  • Ability to write clean, tested, readable code and explain design choices
  • Data structures, algorithms, concurrency, TCP/HTTP networking, and operating-system fundamentals
  • Understanding of distributed-systems concepts including idempotency, retries, back-pressure, caching, and eventual consistency
  • Hands-on exposure to Prometheus, Grafana, Loki, or similar observability tools
  • Ability to write basic PromQL queries and instrument services
  • Familiarity with Linux, the shell, system logs, and standard debugging tools
  • Kubernetes basics, including Pods, Services, and Deployments; experience running something on Kubernetes
  • Git, branching, pull requests, and CI pipelines such as GitHub Actions or GitLab CI
  • Unit and integration testing discipline
  • Clear written and verbal English
  • Curiosity and willingness to learn GPU/AI infrastructure, AIOps, distributed systems, and observability
  • Nice-to-have: internship or project experience in monitoring, observability, telemetry pipelines, or platform/SRE tooling
  • Nice-to-have: exposure to GPU/AI infrastructure such as DCGM, InfiniBand/RoCE, Kubernetes GPU Operator, Slurm, or Ray
  • Nice-to-have: exposure to AIOps/ML-adjacent tooling
  • Nice-to-have: contributions to open-source observability or cloud-native projects
Core Competencies

Demonstrates strong software engineering fundamentals with proficiency in Go, Python, or Java, and a solid understanding of distributed systems and observability tools. Capable of writing clean, tested code and instrumenting services for performance monitoring and reliability.

Highest-signal resume keywords
  • Proficiency In Go, Python, Java, Or Rust
  • Experience With Kubernetes
  • Familiarity With Prometheus And Grafana
  • Understanding Of Distributed Systems Concepts
  • Unit And Integration Testing Discipline
ATS Optimization Keywords
Hard Skills
  • Software Engineering
  • Clean Code Writing
  • Data Structures
  • Algorithms
  • Concurrency
  • TCP/HTTP Networking
  • Operating System Fundamentals
  • PromQL Queries
  • Unit Testing
  • Integration Testing
Soft Skills
  • Clear Written And Verbal English
  • Curiosity
  • Willingness To Learn
Industry Keywords
  • Observability
  • Telemetry Pipelines
  • Platform/SRE Tooling
  • GPU/AI Infrastructure
  • AIOps
Tools & Technologies
  • Kubernetes
  • Prometheus
  • Grafana
  • Loki
  • Git
  • GitHub Actions
  • GitLab CI
  • OpenTelemetry
  • CI/CD Release Pipeline
  • Linux
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Mid-Level SRE Analyst
Mid-Level SRE Analyst

Jobtailor • Deutschland

Remote
EUR 90.000 - 130.000
Senior GPU Cloud, K8S Expert
Senior GPU Cloud, K8S Expert

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Senior DevOps Engineer
Senior DevOps Engineer

Jobtailor • Düsseldorf

Vor Ort
EUR 90.000 - 130.000
Staff Software Engineer – Platform Team
Staff Software Engineer – Platform Team

Jobtailor • Deutschland

Remote
EUR 90.000 - 120.000
Mid-Level SRE
Mid-Level SRE

Jobtailor • Deutschland

Hybrid
EUR 70.000 - 110.000
DevOps Engineer, Blockchain Infra
DevOps Engineer, Blockchain Infra

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Senior Systems Software Engineer, Observability and Telemetry Platform
Senior Systems Software Engineer, Observability and Telemetry Platform

Jobtailor • Deutschland

Remote
EUR 90.000 - 130.000
Senior System Software Engineer, Software Defined Networking
Senior System Software Engineer, Software Defined Networking

Jobtailor • Deutschland

Remote
EUR 90.000 - 140.000
Senior Principal Platform Engineer – Connected Aviation
Senior Principal Platform Engineer – Connected Aviation

Jobtailor • Deutschland

Hybrid
EUR 110.000 - 170.000
Senior AI Platform Engineer – Enterprise Systems
Senior AI Platform Engineer – Enterprise Systems

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000