Infrastructure & Platform Operations Engineer

Embedded Shishya

Deutschland

Vor Ort

EUR 90.000 - 120.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Competitive compensation package
Government‑mandated benefits plus HMO
Hybrid/Remote work arrangement

Zusammenfassung

DysrupIT is seeking an Infrastructure & Platform Operations Engineer to support, maintain and improve production platforms and cloud/on‑prem infrastructure. The role focuses on BAU production support, service stability and knowledge transfer, with opportunities to contribute to automation and platform projects.

You will work across Linux, databases, containers, monitoring and enterprise platforms, and collaborate with cross‑functional teams to deliver end‑to‑end platform outcomes.

Qualifikationen

  • Typically 5+ years of experience in infrastructure engineering, platform operations, systems engineering or production support.
  • Strong BAU and production support experience with live incidents and root-cause analysis.
  • Hands-on AWS experience in production environments is essential.
  • Very strong Linux administration and troubleshooting skills (Ubuntu/Red Hat).
  • Deep networking knowledge including TCP/IP, DNS, routing, firewalls and cloud networking.
  • Solid PostgreSQL and MongoDB operational experience.
  • Experience with Docker and Kubernetes in production environments.
  • Scripting ability in Python and PowerShell; Bash is valuable.
  • Knowledge of Git and GitHub for version control of scripts and configurations.

Aufgaben

  • Work as part of Service Operations to support enterprise infrastructure and production platforms.
  • Administer, support and troubleshoot customer environments including storage, networking and compute resources.
  • Support Linux environments, primarily Ubuntu and Red Hat; Windows support where needed.
  • Diagnose complex network issues across TCP/IP, DNS, routing, firewalls and VPNs.
  • Manage PostgreSQL and MongoDB production environments including backups and recovery.
  • Support containerised environments with Docker and Kubernetes and cloud services (AWS ECS/EKS).
  • Use scripting to automate operational tasks and improve repeatability.
  • Collaborate with DevOps, Data Services, Product and Delivery teams to deliver platform outcomes.
  • Attend CAB and manage Change records with appropriate rollback plans.

Kenntnisse

AWS
Linux administration
Network troubleshooting
Docker & Kubernetes
Python/PowerShell scripting
Git/GitHub
Incident & problem management
Customer-focused communication

Tools

Ubuntu
Red Hat
PostgreSQL
MongoDB
Docker
Kubernetes
AWS ECS/EKS

Jobbeschreibung

About Dysrupit

DysrupIT is a consulting-led technology firm. We help mid-market to enterprise businesses solve business problems through technology — whether that's consulting, execution, managed services, or staff augmentation — and we take accountability for the outcome. We are dedicated to making a positive impact in the communities we serve.

Company Culture

At DysrupIT, success isn't measured in headcount placed or hours billed — it's measured in outcomes delivered. We're a team that takes ownership of the work, stays curious about problems beyond our immediate scope, and builds relationships meant to grow, not just renew or end. We invest in our people with the training and support they need to grow their careers, and we back a culture where everyone, regardless of role, is encouraged to notice opportunities, ask one more question, and help lead the story for our clients, not just deliver it.

Job Summary

The Infrastructure & Platform Operations Engineer supports, maintains and improves the infrastructure and production platforms used to deliver Nephos services to customers. The jobholder works across cloud / on-premise infrastructure, Linux, networking, databases, containers, monitoring and enterprise platforms, with a strong initial focus on BAU production support, service stability and knowledge transfer. As capability develops, the role will also contribute to platform change, upgrades, automation and project delivery.

Job Responsibilities

Working hours: 9–5 UK hours during onboarding and transition, with flexibility for occasional out-of-hours upgrades, changes and major incidents.

  • Work as part of Service Operations to support enterprise infrastructure and production platforms, including health monitoring, issue resolution, proactive maintenance, resilience and continuous improvement.
  • Administer, support and troubleshoot customer environments, including core infrastructure, storage, networking, IAM awareness, compute resources and managed services relevant to supported platforms.
  • Support Linux environments, primarily Ubuntu and Red Hat, using command-line administration and troubleshooting techniques; provide appropriate support for Windows where required.
  • Diagnose complex network and connectivity issues across TCP/IP, DNS, routing, firewalls, VPNs, TLS, proxies, load balancers and cloud networking.
  • Administer and support PostgreSQL production environments, including configuration, roles and permissions, monitoring, backup and restore, recovery, performance investigation and upgrades.
  • Administer and support MongoDB production environments, including health, logs, performance, backup and recovery, upgrades and troubleshooting.
  • Support containerised and orchestrated environments using Docker and Kubernetes, including container lifecycle, logs, configuration, volumes, networking, registries, health checks, resource constraints and failure diagnosis.
  • Support cloud container services and clusters, including AWS with EKS and ECS where applicable, and contribute to platform configuration, health and capacity management.
  • Monitor infrastructure and applications using enterprise monitoring and observability tools, including alert investigation, log analysis, threshold review, root-cause identification and reduction of unnecessary alert noise.
  • Use scripting and automation, particularly Python and PowerShell, to improve repeatability, reduce manual effort and support operational tasks; Bash or other transferable scripting experience is also valuable.
  • Use Git and GitHub to manage scripts, configuration and operational content, following appropriate version-control practices.
  • Support enterprise platform installations, configuration, upgrades, testing, validation and recovery, including data platforms such as BigID and internally developed solutions such as Nephos-developed platforms where customer access is approved.
  • Learn and develop operational capability in new or unfamiliar enterprise platforms quickly, using structured knowledge transfer, documentation, practical shadowing and independent task completion.
  • Take ownership of Incident and Problem records, applying structured troubleshooting to identify issues, restore service as quickly and safely as possible, complete root-cause analysis and maintain accurate records.
  • Take ownership of Service Requests and technical operational tasks, ensuring work is completed accurately, securely and within agreed priorities.
  • Manage Change records relating to infrastructure, databases, platforms, upgrades and improvement activity, attending CAB where required and ensuring implementation, validation and rollback plans are appropriate.
  • Create and maintain clear knowledge articles, runbooks and work instructions, and actively participate in cross-training so that critical operational knowledge is not concentrated with one individual.
  • Work directly with customer and internal technical teams during incidents, changes, upgrades and investigations, explaining findings clearly and maintaining a professional, customer-focused approach.
  • Collaborate with DevOps, Data Services, Product, Delivery and other cross-functional teams to resolve issues and deliver end-to-end platform outcomes.
  • Contribute technical input to platform improvements, resilience, capacity, monitoring and operational design. Final architecture and governance decisions remain shared internal responsibilities.
  • Support backup, restore and disaster-recovery activities, including practical validation and recovery exercises.
  • Work in accordance with security principles including least privilege, secure credential handling, secrets, certificates, auditability, patching and vulnerability awareness.
  • Access to individual customer environments remains subject to the relevant customer contractual, security and access approval processes.
  • The role is focused on operating, supporting, deploying and improving production platforms rather than developing application functionality.
Qualifications
  • Typically 5+ years of experience in infrastructure engineering, platform operations, systems engineering or production support; demonstrable capability is more important than a fixed number of years.
  • Strong, practical BAU and production support experience, including ownership of live incidents, service restoration, monitoring, failed changes or deployments, root-cause analysis and controlled remediation.
  • Strong hands‑on AWS experience in production environments. AWS is the primary cloud requirement for this role.
  • Very strong demonstrable Linux administration and troubleshooting experience, particularly Ubuntu and Red Hat, including confident command-line use.
  • Deep practical networking and connectivity troubleshooting experience covering TCP/IP, DNS, routing, firewalls, VPNs, TLS and cloud networking.
  • Strong production MongoDB administration and troubleshooting experience. This is a critical capability for the role.
  • Meaningful PostgreSQL operational experience, including administration, backup and recovery, monitoring and troubleshooting; deeper production DBA experience is highly valued.
  • Hands‑on production experience with Docker and a strong working knowledge of container lifecycle, networking, logging, health and troubleshooting.
  • Strong hands‑on Kubernetes experience, including deployment health, pods, services, configuration, logs and failure diagnosis.
  • Experience supporting enterprise applications and platforms through installation, configuration, version upgrades, health monitoring and technical troubleshooting.
  • Strong scripting and automation capability, particularly Python and/or PowerShell; experience with Bash or other transferable scripting languages is also relevant.
  • Strong monitoring and observability fundamentals, including infrastructure metrics, log investigation, alert analysis and distinguishing symptoms from root cause.
  • Excellent issue analysis, troubleshooting and fault‑finding skills, with a methodical approach to complex production problems.
  • Working knowledge of Git and GitHub, including repositories, version control and collaborative working practices.
  • Practical understanding of backup, restore, recovery and service‑resilience principles.
  • Practical IT service management experience covering Incident, Problem, Change, Service Request and major‑incident processes. Formal ITIL certification is not mandatory.
  • Experience using a service‑management platform such as Jira/JSM, ServiceNow, Remedy, Freshservice or equivalent.
  • Good security awareness, including least privilege, IAM concepts, secrets, certificates, MFA, privileged access, patching, vulnerability awareness and secure credential handling.
  • Strong customer‑centric attitude with the ability to communicate clearly with technical and non‑technical stakeholders during normal operations and incidents.
  • Excellent written and verbal communication skills, including the ability to create clear technical documentation and operational procedures.
  • Demonstrable ability to learn complex unfamiliar technology and become independently effective following structured onboarding and knowledge transfer.
  • Able to work independently following onboarding, manage multiple priorities and recognise when escalation or wider technical input is required.
  • Flexible, organised and self‑motivated, with willingness to participate in occasional out-of-hours upgrades, changes and major incidents where required.
Nice to Have
  • Hands‑on Azure experience. Azure is expected to become increasingly relevant, although AWS remains the primary current requirement.
  • Experience with Google Cloud Platform and services such as Cloud Run or Cloud SQL.
  • Experience with infrastructure‑as‑code principles and tooling such as Terraform, CloudFormation, Ansible or equivalent.
  • Familiarity with CI/CD platforms and deployment pipelines such as GitHub Actions, Jenkins, GitLab CI or Azure DevOps.
  • Previous experience with BigID or another complex enterprise data‑governance platform. BigID experience is not required, but the ability to learn it is important.
  • Experience with RabbitMQ, Redis or similar messaging/cache technologies.
  • Experience with LogicMonitor or comparable enterprise monitoring platforms.
  • Experience using privileged‑access and secrets‑management technologies such as Delinea and HashiCorp tooling.
  • Experience with REST APIs and API troubleshooting.
  • Experience in applying, renewing and troubleshooting TLS certificates.
  • Experience with disaster‑recovery planning, RPO/RTO considerations and practical recovery testing.
  • Knowledge of data governance, information security or regulated customer environments.
  • Relevant AWS, Azure, Kubernetes, PostgreSQL, MongoDB or ITIL certifications. Certifications are advantageous but not mandatory.
  • Experience supporting microservices, serverless or distributed platform architectures.
What We Offer
  • Competitive compensation package commensurate with experience.
  • Government‑mandated benefits plus supplemental HMO coverage.
  • Collaborative and professional work environment.
  • Career growth opportunities within a growing technology organization.
  • Hybrid/Remote work arrangement (where applicable).
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Middle Tech Support / Junior DevOps specialist
Middle Tech Support / Junior DevOps specialist

Aether Biomedical • Deutschland

Hybrid
EUR 65.000 - 90.000
Health and life insurance (Luxmed)
MyBenefit platform with Multisport
english language classes
+2
DevOps Engineer – EU Night Shift
DevOps Engineer – EU Night Shift

Jobtailor • Deutschland

Hybrid
EUR 80.000 - 120.000
Senior Engineer (Platform)
Senior Engineer (Platform)

Jobgether • Deutschland

Vor Ort
EUR 120.000 - 160.000
Senior DevOps Engineer
Senior DevOps Engineer

EhsanLab • Deutschland

Hybrid
EUR 70.000 - 90.000
Flexible full-time and part-time arrangements
Exposure to compliance-heavy environments
Work across real fintech and banking platforms
DevOps Engineer
DevOps Engineer

Zero to One Search | Recruitment Agency • München

Vor Ort
EUR 90.000 - 130.000
Competitive benefits including travel
Wellness and gym discounts
Regular team and company events
+2
Staff / Senior DevOps Engineer
Staff / Senior DevOps Engineer

Akeno • Hamburg

Vor Ort
EUR 90.000 - 140.000
Generous Annual Compensation
35-40 Hours Working Week
Relocation Package
+10
Senior Infrastructure Engineer (Network, Datacenter)
Senior Infrastructure Engineer (Network, Datacenter)

erasys GmbH • Berlin

Hybrid
EUR 80.000 - 120.000
Hybrid work model
German language classes in Berlin
Wellpass gym access
+2
Senior Site Reliability Engineer (SRE / Backend)
Senior Site Reliability Engineer (SRE / Backend)

nilo.health • Berlin

Hybrid
EUR 110.000 - 140.000
Fully remote / in-office flexibility
Equity options
Office dogs in Berlin & Munich
+3
Platform Engineer
Platform Engineer

Aether Biomedical • Deutschland

Hybrid
EUR 90.000 - 120.000
Vacation days up to 26
Illness days: 10 per year
Health and life insurance
+6
Staff System Engineer (f/m/d) - TechOps DNS & Observability
Staff System Engineer (f/m/d) - TechOps DNS & Observability

IONOS SE • Karlsruhe

Hybrid
EUR 90.000 - 130.000
Hybrid working model
Flexible working hours
Subsidized canteen at some locations
+4