Turn this role into an interview — a resume and cover letter built around what this employer wants.
Day One Partners is partnering with an early-stage AI startup to hire 1–2 Applied AI Engineers. You will own the intelligence and infrastructure behind production agents, focusing on memory, orchestration, and system reliability.
You’ll work on frontier models, develop Evals and benchmarks, and build scalable sandboxed agent infrastructure, taking ownership from experimentation through deployment.
We’re partnering with an early-stage, venture-backed AI startup building the collaboration layer for humans and AI agents.
The company is developing systems that allow agents to maintain context, work across long-running tasks, collaborate with people and other agents, and become more effective over time. This is a small, highly technical team tackling problems at the edge of what production agent systems can reliably do today.
They’re hiring 1–2 Applied AI Engineers to take significant ownership of the intelligence and infrastructure behind the product.
You’ll build the systems surrounding frontier models that determine how agents remember, act, coordinate, and improve.
The core scope spans agent memory, task optimization, agent orchestration, harnesses, evals, and benchmarks, alongside the infrastructure required to run these systems securely and reliably at scale.
This is a hands‑on engineering role for someone who wants to work below the application layer. You should be excited by the hard parts of making agents actually work in production, not just integrating an LLM into an existing product.
We care considerably more about the depth and quality of what you’ve built than a particular number of years of experience.
The strongest candidates will have done things like:
You’ll work on problems where the underlying models are improving rapidly, but the infrastructure around them is still being invented.
How should an agent decide what to remember? How do you optimize performance across a task that may run for hours? How do you benchmark something nondeterministic? How do you safely run large numbers of agents capable of taking real actions?
If you’ve already spent meaningful time wrestling with problems like these and want substantially more ownership over them, we’d love to hear from you.