THE TECHNICAL CHALLENGE
KEY RESPONSIBILITIES
- Set the technical direction of the harness and the interfaces between its levels of abstraction
- Establish the evaluation regime and the fortnightly score, and hold the team to measured results
- Run the R&D team on a two-week rhythm: planning, review with working software, retrospective, written post-mortems for failed experiments and incidents
- Coach, level and hire the engineers; interface with product management on the backlog and with the platform team on what the harness exposes
DESIRED QUALIFICATIONS
- Managed both a research-cadence team and a sprint-cadence team, and can articulate the interface between them
- Built an agent harness or an evaluation harness for computer-use systems, agents or reinforcement learning, with an evaluation culture recognised outside the team, and can say which abstraction was right and which was wrong
- Has taken apart the harnesses of coding agents (OpenHands, SWE-agent, Aider, DeepSeek Harness or comparable) and can explain how they work in plain terms
- Has fine-tuned an LLM, from data preparation to evaluation on a held-out set
EXPECTED QUALIFICATIONS
- T-shaped: deep in one domain, with working breadth in a neighbouring one
- Structures a large, incompletely specified problem and drives it to a working result independently
- Managed a team building on LLMs or machine-learning models in production, with evaluation pipelines as a management requirement and a clear account of how quality was gated
- Turns research or prototype code into reference implementations others reuse, using AI coding tools daily and verifying their output
HOW WE WORK
- Product engineering: we own what we build and run it in production
- Small teams, two-week cycles, working software at every review
- AI coding tools are part of the standard workflow
WHAT WE OFFER
- AI-augmented engineering environment
- Access to on-premise Nvidia B200s
- Flexible work environment
PROCESS
- Practical session; the format is agreed with you
In coding exercises, AI tools are allowed and expected. No LeetCode.
WHO WE ARE
New product organisation as part of a large semi-government in Abu Dhabi. International, ex-FAANG team. Completely greenfield, with a modern tech stack.
REQUIREMENTS TO BE CONSIDERED
- Clear written and spoken English; has reported technical progress to non-technical executives on a fixed cadence
- 6+ years in engineering, of which 2+ leading a team of engineers where work was accepted on measured results
- Still coding, able to review TypeScript and Python; own code operated in production at a product company or a startup
- Bachelor's degree in any field, or self-taught with a track record of open-source contributions
RELATED TECHNOLOGIES AND CONCEPTS
- Agent harnesses and frameworks: OpenHands, SWE-agent, Aider, DeepSeek Harness, Hermes Agent, LangGraph, smolagents
- Programmatic prompt and pipeline optimisation: DSPy, TextGrad, Ax
- Evaluation and observability: evaluation harnesses, LLM-as-judge, trajectory evaluation, Inspect, Langfuse, OpenTelemetry; benchmarks such as Online-Mind2Web, WebArena, OSWorld
- Models and serving: open-weight LLMs, vLLM, SGLang
- State and orchestration: state machines, event sourcing, XState, Trigger.dev, Temporal, Kubernetes, TypeScript, Python