Network Reliability Engineer

MARGO

Warszawa

On-site

PLN 120,000 - 170,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

MARGO in Warszawa seeks a skilled Infrastructure Engineer to build and oversee AI infrastructure with robust monitoring, incident diagnosis, and remediation capabilities. You will collaborate across engineering teams to maintain service continuity and drive observability improvements.

The role emphasizes on-call readiness, deployment governance, and knowledge sharing, with tools and platforms spanning Prometheus, Grafana, and GPU/HPC environments. Proficiency in Go/Python and Linux is expected.

Qualifications

  • Experience with Go or Python.
  • Strong scripting skills (Bash, Python).
  • Hands-on experience with Linux systems (Ubuntu/Debian).
  • Preferred hands-on experience with GPU & HPC infrastructure.
  • Knowledge of networking (VLAN/LAN, TCP/IP, DNS, BGP, load-balancing, IPv6).
  • Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic).
  • Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX).
  • Experience managing relational databases (MariaDB).
  • Understanding of CI/CD pipelines (GitLab).
  • Comfortable with English (written and spoken).

Responsibilities

  • Build a large AI infrastructure with monitoring, diagnosis, and remediation of production incidents
  • Troubleshoot high-impact production issues in collaboration with other engineering teams
  • Participate in an on-call rotation to handle incidents and ensure service continuity
  • Implement and maintain observability solutions to monitor AI infrastructure and application health
  • Contribute to AI infrastructure lifecycle management across different environments and countries
  • Promote and apply best practices in stability, resiliency, scalability, and security
  • Maintain clear technical documentation for tools and procedures
  • Contribute to system and tool evolution based on production feedback
  • Collaborate closely with development teams to ensure infrastructure readiness
  • Participate in team rituals and knowledge-sharing initiatives

Skills

Proactive mindset
Automation & CI
Collaboration & communication
Independent work
Mentoring & knowledge sharing

Tools

Go
Python
Bash
Linux (Ubuntu/Debian)
GPU & HPC infra
Networking (VLAN/LAN, TCP/IP, DNS, BGP, IPv6)
Monitoring: Prometheus/Grafana/Elastic
Infrastructure-as-Code: Ansible/Salt/AWX
MariaDB
GitLab CI/CD

Job description

#HPC #AI #GPU #CLUSTERS

YOUR DAILY ROUTINE
  • Build a large AI infrastructure with monitoring, diagnosis, and remediation of production incidents
  • Troubleshoot high-impact production issues in collaboration with other engineering teams
  • Participate in an on-call rotation to handle incidents and ensure service continuity
  • Implement and maintain observability solutions to monitor AI infrastructure and application health
  • Contribute to AI infrastructure lifecycle management across different environments and countries
  • Promote and apply best practices in terms of stability, resiliency, scalability, and security
  • Maintain clear technical documentation for tools and procedures
  • Contribute to system and tool evolution based on production feedback
  • Collaborate closely with development teams to ensure infrastructure readiness
  • Participate in team rituals and knowledge-sharing initiatives
ABOUT YOU
  • SOFTSKILLS :
    • Proactive and solution-oriented mindset
    • Passion for automation and continuous improvement
    • Strong collaboration and communication skills
    • Ability to work independently and in a team
    • Willingness to mentor and share knowledge
  • HARDSKILLS :
    • Experience with Go or Python
    • Strong scripting skills (Bash, Python)
    • Hands-on experience with Linux systems (Ubuntu/Debian)
    • Preferred hands-on experience with GPU & HPC infrastructure
    • Knowledge of networking (VLAN/LAN, TCP/IP, DNS, BGP, load-balancing, IPv6, etc.)
    • Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.)
    • Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.)
    • Experience managing relational databases (MariaDB)
    • Understanding of CI/CD pipelines (GitLab)
    • Comfortable with English (written and spoken)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Network Reliability Engineer
Network Reliability Engineer

Margo Group • Warszawa

Hybrid
PLN 180,000 - 260,000
Network Reliability Engineer
Network Reliability Engineer

MARGO • Polska

On-site
PLN 180,000 - 300,000
SW Platform Engineer (M/K)
SW Platform Engineer (M/K)

Experis ManpowerGroup Sp. z o.o. • Województwo pomorskie

On-site
PLN 180,000 - 270,000
Senior Site Reliability Engineer (Python/Kubernetes)
Senior Site Reliability Engineer (Python/Kubernetes)

Luxoft Poland • Poland

On-site
PLN 180,000 - 240,000
Senior System Engineer
Senior System Engineer

AIDA projektai, MB • Warszawa

On-site
PLN 180,000 - 260,000
Principal HPC Network Engineer (remote in the EU)
Principal HPC Network Engineer (remote in the EU)

Mirantis • Poznań

On-site
PLN 180,000 - 260,000
DevOps / SRE Engineer
DevOps / SRE Engineer

Spyrosoft Ltd • Wrocław

Remote
PLN 180,000 - 240,000
HPC Network Engineer
HPC Network Engineer

Mirantis • Poznań

On-site
PLN 60,000 - 80,000
Exposure to NVIDIA GPU technologies
Opportunity to work with advanced AI infrastructure
Involvement in complex networking challenges
Senior AI Infrastructure & Platform Operations Engineer (remote in the EU)
Senior AI Infrastructure & Platform Operations Engineer (remote in the EU)

Mirantis • Poznań

On-site
PLN 180,000 - 300,000
Technical Lead - GPU Infrastructure
Technical Lead - GPU Infrastructure

Tether • Warszawa

On-site
PLN 300,000 - 520,000
Remote-friendly
Global collaboration