A leading AI solutions firm is seeking an AI Safety and Evaluations Engineer for a 12-month contract, focusing on designing evaluation frameworks to ensure AI models are bias-free and compliant. The candidate will create automated datasets and develop specific metrics in RAG-based systems. With a requirement of 3+ years in AI Research or Quality Engineering, strong communication skills, and autonomy, this role is integral to maintaining the safety of AI models. This is a remote opportunity with a pay range of $50/hr to $175/hr.
Qualifications
3+ years of experience in AI Research or Quality Engineering.
Deep expertise in model evaluation techniques and NLP metrics.
Demonstrated ability to work autonomously and manage time effectively.
Experience with Python, data analysis tools, and LLM-as-a-Judge frameworks.
Strong communication skills for team updates.
Responsibilities
Design and build evaluation frameworks for model bias.
Create automated datasets to benchmark models before production.
Develop metrics for 'Grounding' and 'Faithfulness' in RAG-based systems.
Build monitoring tools for harmful AI outputs.
Partner with legal and ethics teams for safety constraints.
Skills
AI Research
Quality Engineering
Model evaluation techniques
NLP metrics (ROUGE, BLEU, BERTScore)
Python
Data analysis tools
Self-motivated
Communication skills
Job description
A leading AI solutions firm is seeking an AI Safety and Evaluations Engineer for a 12-month contract, focusing on designing evaluation frameworks to ensure AI models are bias-free and compliant. The candidate will create automated datasets and develop specific metrics in RAG-based systems. With a requirement of 3+ years in AI Research or Quality Engineering, strong communication skills, and autonomy, this role is integral to maintaining the safety of AI models. This is a remote opportunity with a pay range of $50/hr to $175/hr.