Get more replies from employers
Send a job-specific resume in minutes.
OpenTrain AI is seeking an LLM Red Team Specialist to support the development of benchmark challenges. You will probe large language models for coding, ML, and analysis tasks, turning weaknesses into difficult but fair tasks with clear reproducible steps.
You will work independently on ambiguous, open-ended problems while documenting evidence and contributing to a feedback loop that strengthens benchmark quality.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized projects, build an AI training profile, and apply in minutes while contributing to the fast-growing field of human-led AI development.
OpenTrain AI is hiring and contracting for this role. Creating an OpenTrain account is free.
AI training is the human side of building artificial intelligence. Specialists evaluate model behavior, identify weaknesses, and create high-quality examples and feedback that help AI systems become more accurate, reliable, and useful.
Red teaming is a particularly challenging form of model evaluation: you deliberately probe systems for vulnerabilities, edge cases, shortcuts, and failures that ordinary testing may miss.
OpenTrain is seeking an LLM Red Team Specialist to support the development of next-generation agentic evaluation benchmarks. You will probe large language models on coding, machine learning, and analysis tasks, then turn discovered weaknesses into difficult but fair benchmark challenges.
You will work independently on ambiguous, open-ended problems while documenting evidence clearly and contributing to an ongoing feedback loop that strengthens benchmark quality.
Your work will combine adversarial testing, benchmark authoring, technical investigation, and precise written communication. Findings should be clear enough for researchers and collaborators to reproduce and act on.
This role requires strong familiarity with how large language models work, where they fail, and how their performance can be evaluated. Equivalent practical research experience may substitute for the stated academic background.
Experience in AI training, model evaluation, or benchmark and task authoring is helpful. The listing identifies the experience level as entry level, while the required skills call for demonstrated research, security, or AI-evaluation capability.
AI training and evaluation work lets specialists help shape how cutting-edge models behave. Projects are often remote and flexible, making it possible to build experience in a rapidly expanding technology field while applying skills in research, coding, security, and analysis.