Get more replies from employers
Send a job-specific resume in minutes.
OpenTeams is seeking a Senior AI/ML Test and Evaluation Engineer in a hybrid role based in Washington, DC or CO locations. You will build and operate evaluation harnesses and automated metrics alongside expert judgment to stress-test models and agentic workflows.
You will document limitations, surface critical failure modes, and produce defensible reports for senior stakeholders. This hands-on engineering role uses open-source toolchains and may require travel up to 15% to government facilities
Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.
OpenTeams exists to make ownership possible.
Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.
If that sounds like your kind of work, we'd like to meet you.
Location: Washington, DC; Denver, CO; or Colorado Springs, CO preferred (hybrid). Highly qualified candidates outside these locations may also be considered.
Work Authorization: U.S. citizenship required
Clearance: An active U.S. security clearance is strongly preferred. Candidates without an active clearance may be considered for unclassified work but must be eligible to obtain and maintain a clearance.
Salary Range: $145,000-$250,000 USD, dependent on experience level and location
We're looking for a Senior AI/ML Test and Evaluation Engineer to build and operate the benchmarking and evaluation capability at the core of an AI platform. This is a role for someone who is more interested in what a model gets wrong than in what it gets right.
You build the evaluation harnesses - automated metrics paired with structured human expert judgment, applied to candidate models and to the agentic workflows built on top of them. You develop repeatable methodologies for comparing performance against current operational baselines, which means the comparison holds up when someone runs it again in six months with a different model. And you document the limitations and surface the failure modes that matter, including the ones nobody asked about.
Your reports go to senior stakeholders and inform decisions about which capabilities are ready to field. That's the weight of the job: a benchmark that looks good and hides a failure mode is worse than no benchmark at all, and you're the check against that.
This is hands-on engineering on an open-source toolchain.
This position is contingent upon contract award. Travel of up to 15% may be required, primarily to Government facilities and between company locations.