Stand out for this role — generate a tailored resume and cover letter in about a minute.
Sourcegraph is seeking a staff ML and agent engineer to lead production ML systems and evaluation for the Code Understanding surfaces. You will own the hardest problems, set standards for model and evaluation, and drive agentic product improvements at enterprise scale.
You will shape model choices, containment of costs, and latency while partnering with teams across the company to deliver reliable, measurable enhancements to code understanding experiences for developers and their AI tools.
Our mission is to bring clarity and control to the world’s most complex codebases. AI is accelerating code creation, but the infrastructure to understand, oversee, and evolve that code hasn’t kept pace. Sourcegraph gives engineering organizations full visibility across their systems, precise context for their agents, and the ability to execute coordinated code changes at scale. As agentic development becomes the dominant engineering paradigm, we provide the context layer teams need to take control of their codebase. With Code Search, Deep Search, MCP, and Agentic Batch Changes, we deliver on that mission today – giving engineering teams and their AI tools the cross-repo context to navigate massive codebases with confidence, and the ability to make changes across hundreds of repositories at once. Companies like Stripe, Reddit, and Leidos rely on Sourcegraph to ship faster and with higher quality. We’re backed by a16z, Sequoia, and Redpoint, and proud to operate as a globally distributed team that values high agency, direct communication, and customer love. If you want to build the infrastructure that lets every engineering team – and every agent they deploy – operate on their codebase with confidence, join us.
While we hire almost anywhere in the world, we have a preference for someone to reside in the following locations for this role. However, if you feel qualified, we welcome you to apply regardless of location. No matter what, working hours must overlap with EST for at least 20 hours/week.
Sourcegraph is at the forefront of building AI tools to solve the biggest problems in the software industry, problems that only get bigger as codebases grow and as more of the work is done by agents. The Code Understanding team owns the surfaces where that intelligence meets the developer: Deep Search, our agentic, multi-step answer engine across an enterprise’s entire lineup of codebases, Query Assist, turning natural language queries into Sourcegraph query syntax, Smart Hovers, concisely summarizing symbols right where devs need it, guided diff review, and the APIs that both humans and AI agents rely on every day.
Most of what makes those surfaces tick is agent engineering: a blend of software engineering, machine learning, and statistics. Agent engineering tells us which model to use and when, how to retrieve and pack context, how to measure answer quality, where to fine-tune or distill a smaller model to cut costs and latency, and how to expand a single LLM call into a reliable multi-step agent. As the staff ML and agent engineer on Code Understanding, you’ll be the technical owner of it: setting the tam’s direction for models, evaluations, and agentic systems, making our products measurably better, faster, and cheaper, and raising the team’s fluency in building with models.
This is a staff-level role: we’re hiring a technical leader, not just a strong individual contributor. The primary need is a production ML and evaluation authority who also builds production agent systems. You’ll own the hardest, most ambiguous problems in this space, set standards others follow, and influence direction beyond your immediate team. You’ll get the exhilarating chance to drive the vision on how we can provide the best code understanding experience on the market by combining our deterministic, large-scale systems and AI into experiences never seen before.
You are a staff engineer and technical leader with hard-won skills across production machine learning, evaluation, and agent systems. This high-leverage role relies on your ability to make sound model and evaluation decisions for a fast-moving product, build the production systems around them, steer technical direction, and be a force multiplier for a talented, product-minded team. You’re equally comfortable reasoning about an eval harness, a fine-tuning run, a retrieval pipeline, and the multi-step agent loop that ties them together, and you make everyone around you better at all of it.
You operate at staff scope: you own the most ambiguous, highest-risk problems in your domain, go into whatever codebase a problem requires, set standards and patterns others adopt, and translate fluidly between engineering goals and business objectives. You influence direction beyond your immediate team. You lead through technical excellence and mentorship.
You have personally owned a production model lifecycle. You have trained or fine-tuned at least one model and taken it from dataset construction through evaluation, production rollout, and monitoring. You can explain how you chose between training or fine-tuning and prompting or retrieval, how you chose baselines and metrics, what error analysis revealed, and how production evidence affected the next version.
You build agents, fluently and opinionatedly. You’ve designed multi-step agentic systems and made them reliable, observable, and cost-bounded. You have a point of view on where agents shine and where deterministic code or human judgment is required.
You have strong evaluation judgment. You build representative datasets, meaningful baselines, useful error taxonomies, and release criteria that connect offline measurements to production behavior. You know where rigorous evaluations earn their keep, where lightweight smoke tests or qualitative review are enough, and where a precise-looking metric is misleading.
You treat cost and latency as product constraints. You make measured quality, latency, and cost tradeoffs and use the appropriate combination of model selection, prompting, retrieval, caching, distillation, and fine-tuning rather than reaching reflexively for a more complex model.
You operate autonomously on ambiguous problems. Given a rough product idea, a few customer quotes, and a Slack thread, you come back with a plan, a prototype, milestones, and a point of view on tradeoffs, without it being pre-scoped. You own high-technical-risk projects end-to-end.
You contribute beyond your domain. As a senior IC, you go into whatever part of the codebase a problem requires, recognize issues beyond your immediate area, and translate between engineering goals and business objectives.
You up-level the people around you. You mentor by pairing on hard problems, providing substantive design and code reviews, and spreading agent engineering literacy across the team. You see investing in your teammates’ growth as part of the job.
You’re customer and product-driven. You’re comfortable on customer calls and in feedback threads, you turn raw signals into requirements, scopes, and milestones, and you push back when feedback would lead the product astray.
You’re pragmatic, not a perfectionist. You ship the smallest correct thing, prefer robust solutions over complicated ones, and keep a high-quality bar with simplicity.
This job is an IC4. You can read more about our job leveling philosophy in our Handbook.
We pay above-market salaries because we want to hire exceptional people who can focus on building great products, not worrying about paying bills. As an open and transparent company, our compensation philosophy and pay bands are visible to every Sourcegraph teammate, and we strive to make our approach equitable, explainable, and competitive. Your base salary is determined by the IC4 pay band for your location zone (1-4). Our pay bands are informed by market data and designed to ensure competitive compensation wherever you live. During the recruiting process, we’ll discuss the range applicable to you based on job level, relevant skills, experience, qualifications, and location zone.
In addition to competitive cash compensation, we offer meaningful equity and generous perks & benefits.
Below is the interview process you can expect for this role (you can read more about the types of interviews in our Handbook). It may look like a lot of steps, but rest assured that we move quickly and the steps are designed to help you get the information needed to determine if we’re the right fit for you… Interviewing is a two-way street, after all!
Please note – you are welcome to request additional conversations with anyone you would like to meet, but didn’t get to meet during the interview process.
You can learn more about what it is like to work at Sourcegraph by reading our handbook.
We are an ambitious team who are collectively working hard to build the most influential company in the world. You can read more