Product Reliability Engineer | US

OpsMill

Paris

Sur place

EUR 50 000 - 70 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Collaborative work environment
Opportunity to shape product features
Diversity and inclusion commitment

Résumé du poste

OpsMill seeks a reliability-focused engineer to work on Infrahub, solving customer escalations and enhancing system automation. This role bridges practical implementation and reliability, requiring experience in production environments and strong software engineering skills.

The ideal candidate will have 4-7 years of relevant experience, practical expertise in Kubernetes, and proficiency in Python, Go, or Rust. Join a dedicated team shaping automation solutions for complex infrastructure.

Qualifications

  • 4-7 years of experience in SRE or similar roles with a focus on reliability.
  • Strong software engineering fundamentals with a focus on production-quality code.
  • Practical Kubernetes expertise sufficient to debug real deployments.
  • Deep troubleshooting instincts using observability tools.
  • Experience with one or more of: Python, Go, or Rust.
  • Excellent communication skills for breaking down complex issues.

Responsabilités

  • Partner with customers on escalations and resolve complex issues.
  • Build diagnostics tooling to streamline troubleshooting.
  • Own the test automation infrastructure roadmap.
  • Establish performance baselines and regression tests.
  • Write production-quality code for internal tooling.

Connaissances

Production engineering skills
Kubernetes expertise
Python
Go
Rust
Problem decomposition
Communication skills
Remote work capability

Description du poste

Shipping infrastructure software is only half the job. The other half is making it work in environments you don’t control—across messy reality, strict security constraints, and endless platform variations. The difference between a good product and a trusted one is how quickly you can diagnose issues and how effectively you prevent them from happening again.

At OpsMill, we're building Infrahub, a schema-driven infrastructure source of truth that helps teams unify data and scale automation reliably. Our customers deploy Infrahub on-prem, which means reliability is a product feature, not just an operational concern. When something breaks in the field, it's not just a support ticket—it's a signal about what we need to fix, test, or instrument better.

Why This Role Exists

We need someone who can operate in both worlds: diving deep on gnarly customer escalations while systematically eliminating entire classes of problems. You'll be the crucial bridge between "customer is blocked right now" and "this type of issue can't happen again." You'll build the diagnostics, tests, and automation that turn on-prem deployment chaos into predictable, debuggable, fixable reliability.

What You'll Be Doing
  • Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments.

  • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements

  • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster

  • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do

  • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early

  • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails

  • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability

  • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents

What You Bring
  • 4-7 years of experience in production engineering, SRE, platform engineering, or similar roles where you've owned reliability and customer escalations

  • Strong software engineering fundamentals including design, debugging, testing, code review, and a focus on maintainable, production-quality code

  • Practical Kubernetes expertise sufficient to debug real deployments: troubleshooting resources, networking, storage, RBAC, and platform-specific quirks across different distributions

  • Deep troubleshooting instincts and observability experience using logs, metrics, and traces to diagnose issues quickly in complex, distributed systems

  • Experience with at least one of: Python, Go, or Rust for building tooling and contributing to product code (you don't need to be expert in all three)

  • Excellent problem decomposition and communication skills—you can break down messy, ambiguous issues and clearly explain your findings and recommendations

  • Self-directed remote work capability with strong async communication skills and the ability to operate independently in a fast-moving environment where priorities shift based on customer needs

  • Collaborative mindset with experience partnering across product, engineering, and customer-facing teams to drive systematic improvements

Nice-to-Haves
  • Experience with packaging and distribution systems (containers, Helm charts, installers) and managing upgrade/migration flows

  • Background running CI/CD at scale including test parallelization, hermetic environments, and artifact management

  • Familiarity with performance tooling such as profiling, load generation, and benchmark harnesses

  • Previous experience in customer-facing technical roles like escalation engineering, support engineering, or solutions engineering

  • Contributions to open source projects, especially in infrastructure, observability, or reliability tooling

Why OpsMill?
  • The people: Work alongside world-class engineers who've built and scaled automation platforms in production. Daily technical challenges with smart colleagues who push you to grow.

  • The product: Shape Infrahub based on real customer needs. Your input directly influences features, integrations, and roadmap priorities.

  • The mission: We're making enterprise-grade infrastructure automation accessible to any organization. Open-source at the core, production-ready out of the box. This is a multi-year journey, not a quarterly sprint.

  • The impact: You'll work with teams managing some of the world's most complex infrastructure deployments, solving problems that ripple across entire organizations.

Our Commitment to Diversity and Inclusion

OpsMill is committed to building a diverse and inclusive team. We believe different perspectives make us stronger and more innovative. We encourage applications from candidates of all backgrounds and experiences, and we're committed to providing an inclusive environment where everyone can do their best work.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Field Automation Engineer
Field Automation Engineer

OpsMill • Paris

Hybride
EUR 60 000 - 85 000
Comprehensive health, dental, and vision coverage
Flexible PTO
Home office setup allowance
+2
On-Prem Reliability Engineer: Kubernetes & Diagnostics
On-Prem Reliability Engineer: Kubernetes & Diagnostics

OpsMill • Paris

Hybride
EUR 50 000 - 70 000
Collaborative work environment
Opportunity to shape product features
Diversity and inclusion commitment
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Replit • France

Sur place
EUR 80 000 - 110 000
Competitive Salary & Equity
Health, Dental, Vision and Life Ins.
Paid Parental, Medical, CaregiverLeave
+6
Senior Site Reliability Engineer: Scale, Automate, Observe
Senior Site Reliability Engineer: Scale, Automate, Observe

Replit • France

Sur place
EUR 80 000 - 110 000
Manager I, Engineering - Observability Pipelines (OP)
Manager I, Engineering - Observability Pipelines (OP)

United States Digital Space LLC • Paris

Hybride
EUR 95 000 - 130 000
New hire stock equity (RSUs)
Employee stock purchase plan (ESPP)
Professional development
Sr. Devops Engineer II
Sr. Devops Engineer II

DoubleVerify • Paris

Sur place
EUR 90 000 - 125 000
Infrastructure & Cloud Engineering Manager
Infrastructure & Cloud Engineering Manager

StrangeBee • Paris

Sur place
EUR 70 000 - 100 000
Manager 1, Technical Escalations Engineering - Paris
Manager 1, Technical Escalations Engineering - Paris

United States Digital Space LLC • Paris

Sur place
EUR 60 000 - 80 000
Stock equity (RSUs)
Career development opportunities
Inclusive workplace culture
CI Engineer
CI Engineer

Metabase • France

À distance
EUR 68 000 - 94 000
Flexible working hours
Fully remote environment
Ownership over CI systems
+1
Staff Site Reliability Engineer (x/f/m)
Staff Site Reliability Engineer (x/f/m)

EngineersOfAI • Paris

Hybride
EUR 120 000 - 180 000