Senior Infrastructure Engineer

ToolsGroup Inc.

Italia

In loco

EUR 55.000 - 68.000

Tempo pieno

14 giorni+

Ricevi più risposte dai datori di lavoro

Invia un CV specifico per questa offerta in pochi minuti.

Vantaggi offerti da questo lavoro

Bonus up to 10%

Descrizione del lavoro

ToolsGroup Inc. is seeking an experienced Site Reliability Engineer to lead IT service operations and drive reliability across production services from our Italy site.

You will diagnose failures, automate recurring tasks, and mentor a small IT/Ops team while coordinating with Engineering, Security, and Support to protect service availability.

Competenze

  • Experience leading production engineering teams and owning incidents.
  • Strong troubleshooting across Windows and Linux, DNS, networks, and apps.
  • Hands-on cloud ops with Azure, OCI, and related services.
  • Proficiency with containers, Kubernetes, and storage concepts.
  • Experience with observability, metrics, logs, traces, and alerting.

Mansioni

  • Lead major incidents from containment through recovery and post-incident actions.
  • Troubleshoot complex distributed systems across multiple stacks and cloud environments.
  • Operate and improve cloud environments (Azure, OCI) including IAM and networking.
  • Design and implement reliable, observable services with SLI/SLO guidance.
  • Automate tasks using PowerShell, Python, IaC, and CI/CD pipelines.

Conoscenze

Production engineering
Incident management
Cloud operations
Kubernetes
Observability
PowerShell
Python
CI/CD
Networking
Windows/Linux troubleshooting

Formazione

ITIL (preferred)

Strumenti

Azure
OCI
Terraform
PowerShell
Python
Kubernetes
CI/CD pipelines

Descrizione del lavoro

About Us

We are a dynamic, rapidly growing global company and the innovators of service-driven supply chain planning software. We help companies make better, faster supply chain decisions that reduce inventory, improve customer satisfaction, and deliver powerful financial results amid increasing complexity, product proliferation, and uncertainty.

Our solutions have been recognized by customers globally and analyst firms, such as Gartner, for our ability to support service and inventory trade-offs, while dramatically improving planner productivity. ToolsGroup has been successfully deployed worldwide in more than 44 countries, and we have one of the highest customer retention rates in our industry.

About the Role

We are looking for an experienced Site Reliability Engineer who can also lead IT service operations. You will lead a small IT/Ops team,remainthe senior technical escalation point for production services, and act as the operational interface between IT/Ops, Engineering, Product, Security, Support, and business teams.

This is not a coordination-only service management role or a generalist infrastructure position. You will diagnose distributed-system failures using logs, metrics, traces, commands, and platform tooling; make safe recovery decisions; automate recurring work; and engineer lasting reliability improvements.

Main Responsibilities
  • Lead major incidents from impact assessment and containment through recovery, stakeholder communication, root-cause analysis, and corrective actions.
  • Troubleshoot complex issues across Windows and Linux systems, Kubernetes and container workloads, hybrid networking and DNS, cloud infrastructure, identity, authentication, databases, storage, APIs, and service dependencies.
  • Operate and improve Azure, OCI, or comparable cloud environments, including monitoring, access controls, backup and recovery, reliability, and cost-aware scaling.
  • Define and improve service-level indicators andobjectives, observability, alert quality, capacity, resilience, dependency mapping, and recovery readiness for critical services.
  • Automate operational tasks and controls using PowerShell, Python, infrastructure as code, or CI/CD pipelines, with validation, logging, secure credential handling, and rollback.
  • Apply incident, change, and problem management pragmatically, protecting service availability without introducing unnecessaryprocess.
  • Connect technical and business teams: clarify service ownership and dependencies, translate business needs into reliability and infrastructure requirements, frame risk and tradeoffs, align priorities, and ensure decisions have accountable owners and realistic commitments.
What We Are Looking For
  • We care more aboutdemonstratedproduction engineering experience than a checklist of certifications. Strong candidates will bringall ofthe following.
  • A strong SRE or Production Engineering background, typically5+yearsoperating business-critical, customer-facing, or high-availability services. Recent work must include direct technical ownership, not only coordination or people management.
  • Recent ownership of high-severity incidents, including technical triage, recovery decisions, clear communications, and measurable follow-through.
  • Strong systems and network troubleshooting fundamentals: Windows and Linux, TCP/IP, DNS, routing, firewalls, proxies or load balancers, and the ability to isolate faults across service layers.
  • Hands-on cloud operations experience in Azure, OCI, or a similar platform, includingcompute, storage, networking, IAM, observability, backup, and recovery.
  • Production experience with containers and Kubernetes, including workload health, scheduling, networking, persistent storage, secrets, deployment and rollback, scaling, and backup or recovery considerations.
  • Deep observability and reliability engineering practice: metrics, logs, distributed tracing, actionable alerting, SLI/SLO design, capacity and saturation analysis, failure-mode thinking, and post-incident engineering.
  • Ability to diagnose database-backed and API-driven services across application, query, connection-pool, storage, certificate, network, and downstream dependency layers.
  • Practical identity and access management experience with Active Directory and Microsoft Entra ID or equivalent, including hybrid identity, privileged access, MFA, service identities orgMSAs, lifecycle controls, dependency mapping, and controlled recovery from identity failures.
  • Evidence of safe automation and infrastructure-as-code work using PowerShell, Python, Terraform or comparable tooling and CI/CD. You should be able to explain testing, idempotency, error handling, credential security, rollout, rollback, and measurable impact.
  • Experience leading, mentoring, or acting as the senior escalation point for other technical professionals.
  • Strong business-facing and cross-functional leadership. You can translate technical complexity into business impact and options, challenge unsafe or unrealistic requests constructively, negotiate priorities, and communicate decisions clearly to engineers, executives, customers, and non-technical stakeholders.
Additional Relevant Experience
  • Microsoft 365, endpoint management, EDR, device compliance, and hybrid workplace operations.
  • Formal ITIL, cloud, security, or infrastructure certifications.
What success looks like
  • Incidents are diagnosed and resolved with greater speed, structure, and confidence.
  • Monitoring, runbooks, automation, recovery controls, and change practices reduce repeat failures and operational toil.
  • The team becomes more capable and accountable without relying on a single hero.
  • Technical and business teams share clear service ownership, priorities, risk decisions, and delivery expectations.
Our hiring process

The process includes a scenario-based SRE technical discussion. We will ask you to think aloud through realistic production incidents involving cloud, Kubernetes, identity, networking, databases, storage, APIs, and service dependencies. You will be expected to describe the logs, metrics, traces, commands, tools, tradeoffs, and recovery criteria you would use. We will also assess how you align technical and business stakeholders when priorities, risk, and customer commitments conflict.

Our Vision, Purpose, and Values

Our VISION: Unparalleled control over demand and supply to deliver certainty.
Our PURPOSE: Problem Solvers Welcome.
Our VALUES: Deliver the Goods - Have Deep Care - Find the Right Answer, Not the First Answer - Creativity That Endures - Brilliant But Not Loud.

Salary range:

55-68k/year, plus 10% bonus based on personal and company objectives.

Equal Opportunity Employer

ToolsGroup provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
ToolsGroup is an E-Verify employer, to learn more please visit E-Verify.gov

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior Infrastructure Engineer
Senior Infrastructure Engineer

ToolsGroup • Milano

In loco
EUR 55.000 - 68.000
Bonus up to 10%
Site Reliability Engineering Manager
Site Reliability Engineering Manager

CoreView • Milano

In loco
EUR 90.000 - 130.000
Senior IT/Ops Lead (SRE) — Global SaaS Reliability
Senior IT/Ops Lead (SRE) — Global SaaS Reliability

ToolsGroup • Milano

In loco
EUR 55.000 - 68.000
Bonus up to 10%
SRE/DevOps Engineer
SRE/DevOps Engineer

E80 Group • Reggio Emilia

Remoto
EUR 70.000 - 110.000
Permanent contract
Full time
Full remote
+1
DevOps Engineer
DevOps Engineer

Bitrock • Turbigo

In loco
EUR 45.000 - 60.000
Sre Expert
Sre Expert

ING Bank N.V., Milan branch • Milano

Ibrido
EUR 48.000 - 60.000
Flexible Smart Working
Competitive Base Salary
Meal Vouchers
+3
Production Engineer
Production Engineer

Hyperproof • Lombardia

In loco
EUR 115.000 - 167.000
Equity
Bonuses
Comprehensive benefits
Site Reliability Engineer
Site Reliability Engineer

Cacheflow • Milano

In loco
EUR 90.000 - 130.000
DevOps/SRE
DevOps/SRE

xFarm Technologies • Milano

Ibrido
EUR 70.000 - 100.000
Hybrid work model
Senior Customer Reliability Engineer
Senior Customer Reliability Engineer

Sysdig • Italia

In loco
EUR 44.000 - 56.000
Extra days off
Mental health support
Great compensation package