LLM Agent Evaluation & Evolution Architect

ByteDance

San Jose (CA)

On-site

USD 162,000 - 317,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
Short-/long-term disability coverage
Life insurance
Wellbeing benefits
Paid holidays

Job summary

ByteDance's Applied Machine Learning Ark team develops end-to-end MaaS platforms that combine system engineering with large language models to serve businesses globally. The US team designs and operates MaaS solutions across the US and international markets beyond mainland China.

We seek PhD-level researchers and engineers with a solid ML background, Python skills, and hands-on LLM experience to build evaluation pipelines, benchmarks, and production-grade components for scalable AI systems.

Qualifications

  • PhD in CS/AI/ML/Data Science or related field is required or recently completed.
  • Solid foundation in machine learning and deep learning.
  • Hands-on experience with LLM-based systems (agents, tool calls, retrieval, multi-agent systems).
  • Strong Python skills and experience with a mainstream ML or agent evaluation framework.
  • Demonstrated research or engineering ability through publications, projects, internships, or open-source work.

Responsibilities

  • Design evaluation systems for LLM-based agents, covering task success, tool use, reasoning quality, and reliability.
  • Build benchmarks and automated judging pipelines with rule-based checks, model-based judging, and human review.
  • Analyze agent execution traces and user feedback to identify failure patterns and drive improvements.
  • Collaborate with research, platform, and product teams to bring methods into production.

Skills

Solid ML foundation
Hands-on with LLM-based systems
Strong Python skills
Publications/open-source work
Demonstrated research/engineering

Education

PhD in Computer Science/AI/ML/Data Science

Job description

ByteDance's Applied Machine Learning Ark team develops end-to-end MaaS platforms that combine system engineering with large language models to serve businesses globally. The US team designs and operates MaaS solutions across the US and international markets beyond mainland China.

We seek PhD-level researchers and engineers with a solid ML background, Python skills, and hands-on LLM experience to build evaluation pipelines, benchmarks, and production-grade components for scalable AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Agent Evaluation & Evolution Researcher
LLM Agent Evaluation & Evolution Researcher

Bytedance • San Jose (CA)

Hybrid
USD 180,000 - 240,000
LLM Agent Evaluation Engineer — Build Intelligent Agents
LLM Agent Evaluation Engineer — Build Intelligent Agents

Bytedance • San Jose (CA)

On-site
USD 154,000 - 256,000
LLM MaaS Engineer – Agents & Evaluation
LLM MaaS Engineer – Agents & Evaluation

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
LLM Agent Evaluation Engineer Intern (MaaS)
LLM Agent Evaluation Engineer Intern (MaaS)

ByteDance • Seattle (WA)

On-site
USD 58,000 - 61,000
Health insurance
10 paid holidays per year
Housing allowance
LLM Systems Engineer: AI Agents & Evaluation
LLM Systems Engineer: AI Agents & Evaluation

ByteDance • Seattle (WA)

On-site
USD 120,000 - 180,000
ML Engineer, LLM & Multimodal Systems
ML Engineer, LLM & Multimodal Systems

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 170,000
Pioneering Multimodal Visual Gen & Evaluation Scientist
Pioneering Multimodal Visual Gen & Evaluation Scientist

ByteDance • Seattle (WA)

On-site
USD 150,000 - 190,000
AI Evaluation Engineer for LLMs & Multimodal Systems
AI Evaluation Engineer for LLMs & Multimodal Systems

ByteDance • San Jose (CA)

On-site
USD 120,000 - 190,000
Medical Insurance
Dental Insurance
Vision Insurance
+9
LLM Systems Intern: Build & Improve AI Agents
LLM Systems Intern: Build & Improve AI Agents

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 27,552,000 - 41,328,000
Multimodal Vision & AI Evaluation Researcher
Multimodal Vision & AI Evaluation Researcher

Bytedance • San Jose (CA)

Hybrid
USD 150,000 - 230,000