Senior Software Engineer - Paris

H Company

Paris

Hybride

EUR 70 000 - 110 000

Plein temps

Il y a 2 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV personnalisé et une lettre de motivation en environ une minute.

Passez les filtres ATS

Résumé du poste

H Company in Paris is hiring for a backend-focused role that builds and runs the evaluation framework powering benchmarks and models. You will optimize integration, scheduling, observability, and reproducibility across runs, working with researchers and forward deployed engineers.

Ideal candidates have 5+ years in backend development with Python, Kubernetes, Docker, and experience with APIs, databases and cloud services. Hybrid Paris posting with a competitive package.

Qualifications

  • 5+ years in backend development with production Python at the core.
  • Built test/QA or evaluation tooling used by other teams.
  • Operated distributed systems on Kubernetes in a public cloud; AWS helps.
  • Built APIs (REST/GraphQL) and integrations with external services.
  • Know relational and non-relational databases; queues like SQS, RabbitMQ, Kafka.
  • Instrument your work with metrics, tracing and monitoring from the start.

Responsabilités

  • Integration support for researchers and forward deployed engineers bringing in a benchmark.
  • Set the standard for how a benchmark enters the framework, and enforce checks.
  • Scheduling and observability to keep cluster capacity utilized.
  • Ensure reproducible results across trials to support release decisions.
  • Work with various stacks and benchmarks; adapt to different environments.
  • Engage with customers to understand what they want measured and automate it.

Connaissances

Python
Backend development
Kubernetes
REST/GraphQL APIs
AWS
Databases

Outils

Docker
Virtual Machines
Temporal
Dask
FastAPI
PostgreSQL
Grafana
Datadog

Description du poste

About H:

When we released Holo4 on 28 September, we published every trajectory behind its public benchmark scores at trajectories.hcompany.ai. For OSWorld that's 369 desktop tasks, each run 3 times, with the steps, tokens and time of every attempt. This role builds and runs the evaluation framework that produces runs like these.

H builds computer-use agents and the models behind them. Developers use them through a managed API, and our forward deployed engineers take them into enterprise workflows.

What this team owns

The evaluation framework: orchestration, runtimes and observability. Researchers and forward deployed engineers bring the benchmarks, across web apps, desktop applications and the command line. Your job is to make the framework that runs them reliable, fast and cheap, and to make adding a new one quick. Research uses the results to choose checkpoints and decide whether a model ships. Product and the forward deployed engineers use them to measure agents on customer workflows. It carries roughly 50 benchmarks now. That number should be between 100 and 200 soon, and the framework has to keep up.

What you'd be doing
  • Integration support for researchers and forward deployed engineers bringing in a benchmark, with a shorter path each time.

  • Setting the standard for how a benchmark enters the framework, and building the checks that enforce it.

  • Scheduling and observability, so cluster capacity isn't left idle while evaluation jobs queue.

  • Reproducible results across trials, so a release decision rests on numbers that hold.

  • Whatever stack a benchmark calls for. One week that's cluster tuning; the next it's a browser extension or desktop environments.

  • Time with customers, from single developers to large companies, to find out what they want measured, then automating it so the results flow back into our harnesses and models.

The first few months

By 3 months you'll have helped researchers or forward deployed engineers integrate 5 benchmarks, and started fixing what slows the framework down. By 6 months one part of it is yours, for example scaling the runs, observability, or a group of related benchmarks, and a release will have gone out on your numbers. By 12 months you'll know the design and trade-offs of the whole evaluation system, and be the person the rest of H asks about evaluations.

Who you'd work with

Ceiran Chapman, our VP Engineering, is hiring for this role. You'd join the evaluation team. The people relying on your work day to day are H's researchers and forward deployed engineers.

What we think it takes

Likely a good fit if you

  • Have spent 5+ years in backend development, with production Python at the core, and use coding agents to go faster without letting quality drop.

  • Have built test, QA or evaluation tooling that other teams depended on, and care whether a number is right.

  • Have operated distributed systems on Kubernetes in a public cloud. AWS experience helps most.

  • Have built and shipped systems end to end, including APIs (REST or GraphQL) and integrations with outside services.

  • Know relational and non-relational databases, and message queues such as SQS, RabbitMQ or Kafka.

  • Instrument what you build, with metrics, tracing and monitoring from the start.

Stronger still if you have

  • Measured LLM quality before, or built agents yourself.

  • Packaged and run workloads in Docker and on virtual machines.

  • Used Temporal, Dask, FastAPI, PostgreSQL, Grafana or Datadog.

  • Automated web or desktop software with Playwright, Selenium or a browser extension you wrote.

  • Set standards other engineers follow, through code review, design review or mentoring.

How we hire

A 30 minute call with our Talent team, a 60 minute technical challenge, a 60 minute system design interview, and a 30 minute final conversation with Ceiran. About 3.5 hours in total.

Practicalities

Paris posting: Hybrid in Paris. That means 3 office days a week and a London trip about once every 4 to 6 weeks. There is a London posting for the same role. We offer a competitive package.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Research Engineer (Evals)
Research Engineer (Evals)

White Circle • Paris

Hybride
EUR 90 000 - 130 000
Equity
Flexible time off
Hybrid office Paris/London
+4
PRINCIPAL FORWARD-DEPLOYED ENGINEER
PRINCIPAL FORWARD-DEPLOYED ENGINEER

STATION F • Paris

Hybride
EUR 83 000 - 152 000
Member of technical staff (Infrastructure) - Paris
Member of technical staff (Infrastructure) - Paris

H Company • Paris

Sur place
EUR 90 000 - 130 000
Competitive salary
Career growth and professional develop
Member of technical staff (Infrastructure) - Paris
Member of technical staff (Infrastructure) - Paris

hcompany • Paris

Hybride
EUR 90 000 - 120 000
Competitive salary
Career development
Collaborative team
Research Engineer (Evals)
Research Engineer (Evals)

Aisafety • Paris

Sur place
EUR 80 000 - 110 000
Relocation package
Hybrid work (Paris)
Comprehensive medical insurance
+1
Member of technical staff (Infrastructure) - London
Member of technical staff (Infrastructure) - London

H Company • Paris

Sur place
EUR 90 000 - 120 000
In-office collaboration
Forward Deployed Engineer
Forward Deployed Engineer

hcompany • Paris

Hybride
EUR 90 000 - 140 000
Salary and equity
Founding journey
Global team
Infrastructure Engineer (Platform) - London
Infrastructure Engineer (Platform) - London

H Company • Paris

Hybride
EUR 65 000 - 100 000
Competitive salary
Professional growth opportunities
Collaborative multicultural team
+1
Infrastructure Engineer (Platform) - London
Infrastructure Engineer (Platform) - London

Creandum • Paris

Hybride
EUR 65 000 - 90 000
Competitive salary
Hybrid work model
Career development
Founding Backend & AI Engineer
Founding Backend & AI Engineer

Parsio AI • Paris

Sur place
EUR 90 000 - 130 000