A complete application in a minute — tailored resume and cover letter, ready to send.
BMC Software, Inc. is seeking an experienced AI Evaluation Engineer to build and own evaluation frameworks for Generative AI and agentic AI applications.
You will create golden datasets, rubrics, and automated CI evaluation harnesses to ensure quality before shipping new features. You will collaborate across DS, engineering, and product teams, uphold responsible-AI practices, and help shape governance and model validation.
You may occasionally be required to travel for business
Additional Locations:
This role can be based remotely in United States
BMC empowers nearly 80% of the Forbes Global 100 to accelerate business value, faster than humanly possible. Our industry-leading portfolio unlocks human and machine potential to drive business growth, innovation, and sustainable success. BMC does this in a simple and optimized way by connecting people, systems, and data that power the world’s largest organizations so they can seize a competitive advantage.
The IZOT product line includes BMC’s Intelligent Z Optimization & Transformation products, which help the world’s largest companies to monitor and manage their mainframe systems. The modernization of mainframe is the beating heart of our product line, and we achieve this goal by developing products that improve the developer experience, the mainframe integration, the speed of application development, the quality of the code and the applications’ security, while reducing operational costs and risks. We acquired several companies along the way, and we continue to grow, innovate, and perfect our solutions on an ongoing basis.
You build the systems that tell us whether our agentic AI is good enough to ship: evaluation frameworks, golden datasets, and rubrics; automated eval harnesses in CI; and drive down failure modes such as hallucination, drift, and unsafe or non-repeatable output. You are the reason customers can trust what our agents produce. At this level you independently own well-scoped work from definition through delivery. Scope at this level: Independently owns well-scoped eval work for a feature/team from definition through delivery. Organizational impact expected: Improves outcomes for one team or project.
BMC’s culture is built around its people. We have 6000+ brilliant minds working together across the globe. You won’t be known just by your employee number, but for your true authentic self. BMC lets you be YOU!
BMC is committed to equal opportunity employment regardless of race, age, sex, creed, color, religion, citizenship status, sexual orientation, gender, gender expression, gender identity, national origin, disability, marital status, pregnancy, disabled veteran or status as a protected veteran. If you need a reasonable accommodation for any part of the application and hiring process, visit the accommodation request page.
BMC Software maintains a strict policy of not requesting any form of payment in exchange for employment opportunities, upholding a fair and ethical hiring process.
The annual base salary range represents the low and high end of the BMC salary range for this position. Actual salaries depend on a wide range of factors that are considered in making compensation decisions, including but not limited to skill sets; experience and training, licensure, and certifications; and other business and organizational needs.
The range listed is just one component of BMC's employee compensation package. Other rewards may include a variable plan and country specific benefits.
At BMC, it is not typical for an individual to be hired at /near the top of the range. A reasonable estimate of the current range is $152,925 - $254,875
We use AI technology to support parts of our recruitment process, but people—not algorithms—make all final hiring decisions. AI may assist with tasks like scheduling, screening for role alignment, or helping us manage large volumes of applications more efficiently. However, candidates are reviewed by a member of our recruitment team, and interviews and hiring decisions are always made by people. We’re committed to ensuring that technology enhances fairness, efficiency, and the candidate experience—never replaces genuine human judgment.