Test Engineer-AI/LLM

OPPO

Palo Alto (CA)

On-site

USD 100,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Bonus
Long-term incentives
Equal opportunity workplace

Job summary

A leading tech company is seeking a full-time AI/LLM Test Engineer to evaluate the performance and safety of Large Language Models. In this critical role, you will innovate testing methodologies, ensure AI-powered features meet rigorous standards, and collaborate with cross-functional teams. Candidates should have a Bachelor's degree in a technical field, 1+ years of relevant experience, and proficiency in Python and testing frameworks. The position offers a competitive salary range of $100,000–$200,000 plus bonuses and benefits.

Qualifications

  • 1+ years of experience in software testing or ML validation, with exposure to AI/ML systems.
  • Experience with automated evaluation tools or LLM-specific test suites.
  • Familiarity with cloud platforms (GCP, Azure, or AWS) and MLOps tooling.

Responsibilities

  • Evaluate performance, reliability, and safety of LLMs in real-world scenarios.
  • Develop automated test frameworks to evaluate LLM outputs.
  • Collaborate with teams to define test requirements and acceptance criteria.

Skills

Meticulous testing
Innovation in testing methodologies
Python proficiency
Collaboration with engineering teams
Analytical skills

Education

Bachelor's degree in a technical field
Master's degree in AI or Machine Learning

Tools

PyTest
Selenium
Jupyter
Postman

Job description

OPPO US Research Center is seeking a full‑time meticulous and innovative AI/LLM Test Engineer to join our cutting‑edge AI team. In this critical role, you will evaluate the performance, reliability, and safety of Large Language Models (LLMs) in real‑world product scenarios and test end‑to‑end generative AI solutions. Your work will directly shape how users experience AI‑powered features by ensuring robustness, accuracy, and alignment with product goals. This is a unique opportunity to pioneer testing methodologies for next‑generation AI systems at the forefront of technology.

Full‑Time Position Requirements
Core Testing & Evaluation
  • Design and execute performance tests for LLMs across diverse product use cases (e.g., chatbots, content generation etc.)
  • Develop automated test frameworks to evaluate LLM outputs for accuracy, bias, safety, and coherence
  • Conduct end‑to‑end testing of integrated generative AI solutions, including APIs, data pipelines, and user interfaces
Optimization & Validation
  • Collaborate with ML engineers to validate fine‑tuned models and optimize prompts for target scenarios
  • Analyze model failures, edge cases, and adversarial inputs to identify risks and improvement areas
  • Benchmark LLM performance against industry standards and product‑specific KPIs
Collaboration & Quality Assurance
  • Partner with product, engineering, and research teams to define test requirements and acceptance criteria
  • Document defects, performance metrics, and test results to drive data‑driven improvements
  • Advocate for AI ethics and safety through rigorous testing of fairness, bias mitigation, and content moderation
Innovation & Tooling
  • Build scalable tools for synthetic test data generation, prompt variation testing, and automated evaluation workflows
  • Stay current with advancements in generative AI testing, including red‑teaming techniques and evaluation frameworks (e.g., HELM, Dynabench)
  • Propose novel testing strategies for emerging challenges (e.g., hallucinations, context drift)
Basic Qualifications
  • Bachelor’s degree in Computer Science, Data Science, Engineering, or a related technical field, or equivalent practical experience
  • 1+ years of experience in software testing, data science, or ML validation, with exposure to AI/ML systems
  • Proficiency in Python and testing frameworks (e.g., PyTest, Selenium)
  • Hands‑on experience evaluating LLMs in production environments (e.g., GPT, Claude, Llama, Gemini)
  • Strong analytical skills for dissecting model behavior, statistical performance, and failure modes
  • Familiarity with cloud platforms (GCP, Azure, or AWS) and MLOps tooling (e.g., MLflow, Weights & Biases)
  • Experience with version control (Git) and agile development methodologies
Preferred Qualifications
  • Master’s degree in AI, Machine Learning, or a related field
  • Expertise in prompt engineering, LLM fine‑tuning (e.g., LoRA, RLHF), or optimization techniques
  • Experience with automated evaluation tools (e.g., LangChain, TruLens) or LLM‑specific test suites
  • Knowledge of data pipelines, SQL/NoSQL databases, and API testing (e.g., Postman)
  • Background in statistics, quantitative analysis, or data visualization for test insights
  • Contributions to AI safety/ethics initiatives or open‑source LLM evaluation projects
  • Experience testing mobile‑integrated AI solutions (Android/iOS)

We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies, execute evaluation workflows, and assist in model performance validation across diverse generative AI use cases. This contract role is ideal for someone with hands‑on experience in AI/ML evaluation, QA engineering, or data analysis who wants to deepen their exposure to generative AI systems.

Contractor Position Requirements
Testing & Evaluation Support
  • Execute pre‑defined performance tests for LLMs across various tasks (e.g., summarization, Q&A, chatbot flows).
  • Run scripted evaluations to assess outputs for factuality, coherence, and safety.
  • Perform manual and automated test execution on APIs and LLM-integrated user interfaces.
Prompt & Model Validation
  • Assist ML engineers in evaluating prompt variations and prompt‑tuning outcomes.
  • Log and analyze failure cases, anomalies, and edge cases based on provided guidelines.
Collaboration & Documentation
  • Work with QA leads, product managers, and ML engineers to understand test goals and criteria.
  • Report defects, compile evaluation summaries, and maintain testing logs.
Tooling & Automation
  • Use existing internal tools or frameworks to automate test runs and result collection.
  • Contribute to prompt generation, input templating, or result tagging processes.
Basic Qualifications
  • Bachelor’s degree or equivalent work experience in a technical field (e.g., Computer Science, Engineering, Data Science).
  • 6+ months experience in software QA, data labeling, LLM evaluation, or ML testing projects.
  • Basic Python proficiency, especially for data processing and automation tasks.
  • Familiarity with LLMs (e.g., GPT, Claude, Gemini) and prompt‑based outputs.
  • Comfortable working with tools like Jupyter, Postman, or testing dashboards.
  • Detail‑oriented with good documentation habits.
Contractor Details
  • Duration: Long term
  • Rate: Commensurate with experience
  • Conversion Opportunity: High‑performing contractors may be considered for full‑time roles
Benefits

OPPO is proud to be an equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements.

The US base salary range for this full‑time position is $100,000–$200,000 + bonus + long‑term incentives benefits. Our salary ranges are determined by role, level, and location.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
AI Engineer: LLM & Multimodal Agent Architect
AI Engineer: LLM & Multimodal Agent Architect

OPPO US Research Center • Palo Alto (CA)

On-site
AI Researcher - Efficient AI (Contractor)
AI Researcher - Efficient AI (Contractor)

LG Electronics North America • Santa Clara (CA)

Hybrid
USD 84,000 - 92,000
No-cost health premiums for you and 1+
401(k) retirement plan with company-m
Paid time off and holidays
+4
Senior AI Test Automation Engineer
Senior AI Test Automation Engineer

Genuine Parts Company • Alabama

On-site
USD 120,000 - 150,000
AI Engineer
AI Engineer

OPPO US Research Center • Palo Alto (CA)

On-site
USD 100,000 - 200,000
Bonus
Long-term incentives
Equal opportunity workplace
LLM Engineer (Remote)
LLM Engineer (Remote)

Cognizant • Louisville (KY)

Remote
USD 81,000 - 142,000
Medical/Dental/Vision/Life Insurance
Paid holidays plus Paid Time Off
401(k) plan and contributions
+3
AI Researcher - Efficient AI (Contractor)
AI Researcher - Efficient AI (Contractor)

LG Electronics USA • Santa Clara (CA)

Hybrid
USD 120,000 - 180,000
Senior LLM Evaluation Engineer
Senior LLM Evaluation Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000
Senior AI Test Automation Engineer
Senior AI Test Automation Engineer

Motion • Birmingham (AL), Northern (KY)

Hybrid
USD 120,000 - 180,000
AI Engineer
AI Engineer

Fluency • San Francisco (CA)

On-site
USD 180,000 - 250,000
US$1,000 per month food and commuting allowance
Laptop of choice
ESOP available