Network Reliability Engineer

MARGO

Polska

On-site

PLN 180,000 - 300,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

MARGO is seeking a hands-on Infrastructure Engineer to build and operate a scalable AI infrastructure in production. You will monitor, diagnose, and remediate incidents, collaborate across engineering teams, and participate in on-call rotations to ensure service continuity.

Ideal candidates bring strong Linux experience, scripting (Bash/Python), and familiarity with Prometheus, Grafana, and CI/CD pipelines, plus networking and security focus. English proficiency required.

Qualifications

  • Experience with Go or Python.
  • Hands-on Linux administration (Ubuntu/Debian).
  • Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic).
  • Experience with Infrastructure-as-Code (Ansible, Salt, AWX).
  • Relational databases experience (MariaDB).
  • Understanding of CI/CD pipelines (GitLab).
  • Networking knowledge (VLAN/LAN, TCP/IP, DNS, load balancing, IPv6).
  • Comfortable with English (written and spoken).

Responsibilities

  • Build and maintain AI infrastructure with monitoring, diagnosis, and remediation of production incidents.
  • Troubleshoot high-impact production issues in collaboration with other engineering teams.
  • Participate in on-call rotation to handle incidents and ensure service continuity.
  • Implement observability solutions to monitor AI infrastructure and application health.
  • Contribute to AI infrastructure lifecycle management across environments and countries.
  • Promote best practices for stability, resiliency, scalability, and security.
  • Maintain technical documentation for tools and procedures.
  • Collaborate with development teams to ensure infrastructure readiness.
  • Participate in team rituals and knowledge-sharing initiatives.

Tools

Go
Python
Bash
Linux (Ubuntu/Debian)
GPU & HPC infrastructure
Networking fundamentals
Prometheus
Grafana
Elastic
Ansible
Salt/AWX
MariaDB
GitLab CI/CD

Job description

YOUR DAILY ROUTINE


  • Build a large AI infrastructure with monitoring, diagnosis, and remediation of production incidents

  • Troubleshoot high-impact production issues in collaboration with other engineering teams

  • Participate in an on-call rotation to handle incidents and ensure service continuity

  • Implement and maintain observability solutions to monitor AI infrastructure and application health

  • Contribute to AI infrastructure lifecycle management across different environments and countries

  • Promote and apply best practices in terms of stability, resiliency, scalability, and security

  • Maintain clear technical documentation for tools and procedures

  • Contribute to system and tool evolution based on production feedback

  • Collaborate closely with development teams to ensure infrastructure readiness

  • Participate in team rituals and knowledge-sharing initiatives


ABOUT YOU

SOFTSKILLS


  • Proactive and solution-oriented mindset

  • Passion for automation and continuous improvement

  • Strong collaboration and communication skills

  • Ability to work independently and in a team

  • Willingness to mentor and share knowledge


HARDSKILLS


  • Experience with Go or Python

  • Strong scripting skills (Bash, Python)

  • Hands-on experience with Linux systems (Ubuntu/Debian)

  • Preferred hands-on experience with GPU & HPC infrastructure

  • Knowledge of networking (VLAN/LAN, TCP/IP, DNS, BGP, load-balancing, IPv6, etc.)

  • Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.)

  • Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.)

  • Experience managing relational databases (MariaDB)

  • Understanding of CI/CD pipelines (GitLab)

  • Comfortable with English (written and spoken)

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Network Reliability Engineer
Network Reliability Engineer

MARGO • Warszawa

On-site
PLN 120,000 - 170,000
Network Reliability Engineer
Network Reliability Engineer

Margo Group • Warszawa

Hybrid
PLN 180,000 - 260,000
DevOps / SRE Engineer
DevOps / SRE Engineer

Spyrosoft Ltd • Wrocław

Remote
PLN 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hard Rock Digital • Polska

Hybrid
PLN 260,000 - 380,000
Senior System Engineer
Senior System Engineer

AIDA projektai, MB • Warszawa

On-site
PLN 180,000 - 260,000
Senior Site Reliability Engineer (Python/Kubernetes)
Senior Site Reliability Engineer (Python/Kubernetes)

Luxoft Poland • Poland

On-site
PLN 180,000 - 240,000
Staff Engineer I, DevOps Engineering
Staff Engineer I, DevOps Engineering

Bain & Company • Warszawa

On-site
PLN 260,000 - 360,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Luxoft • Poland

On-site
PLN 180,000 - 260,000
Private Medical & Dental care
Life Insurance covered
Internal Mobility program
Network Engineer
Network Engineer

Astreya • Poland

On-site
PLN 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

SIX • Warszawa

On-site
PLN 240,000 - 360,000