Visual Generation & Multimodal Evaluation Machine Learning Engineer Graduate (AML-Ark-US) - 202[...]

ByteDance

Seattle (WA)

On-site

USD 100,000 - 160,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance, located in Seattle, is seeking graduates to join the Applied Machine Learning Ark team. You will contribute to MaaS platforms and end-to-end visual and multimodal model evaluation across text and multimodal LLMs.

You will work on large-scale data pipelines and automated evaluation methods to drive product improvements. Successful applicants have a Bachelor's or Master’s degree in CS/AI/ML/CV, with strong Python and PyTorch skills, and a track record in research or projects.

Qualifications

  • Bachelor’s or Master’s degree in CS, AI, ML, CV, or related field.
  • Solid foundation in deep learning and computer vision fundamentals.
  • Experience in visual generation, multimodal LLMs, or video understanding.
  • Strong Python skills and PyTorch or equivalent framework.

Responsibilities

  • Build evaluation systems for image/video models and agents, covering generation quality and safety.
  • Develop automated metrics and reproducible human evaluation protocols.
  • Design and develop video generation and debugging agents for multi-step workflows.
  • Create large-scale image/video data pipelines to drive model improvements.

Skills

Python
PyTorch
Deep learning
Computer vision
Multimodal evaluation
Multimodal LLMs
Video understanding

Education

Bachelor's degree in Computer Science, AI, ML, CV or related field

Job description

Join us as we work together to inspire creativity and enrich life around the globe.

Location:

Seattle

Team:

Technology

Employment Type:

Regular

Job Code:

A19835B

Share this listing
Responsibilities

The Applied Machine Learning Ark team combines system engineering and machine learning to develop and operate Large Language Model (LLM) service platforms that offer businesses Model-as-a-Service (MaaS) solutions, serving both large model providers and downstream users. The US team drives the design, development, and operation of MaaS solutions across the US and international markets outside mainland China. We are building full-stack, end-to-end solutions spanning text and multimodal LLM algorithms, LLM training/fine-tuning/inference frameworks, prompt engineering, model alignment, and intelligent agent systems. Beyond model serving, we operate large-scale log analytics pipelines that process massive volumes of invocation logs from text models, multimodal models, and agent systems — extracting usage patterns, quality signals, and actionable insights to inform model improvement, system optimization, and product decisions through continuous, data-driven feedback loops. We are actively seeking talented engineers and researchers specializing in Large Language Models and AI Agent systems to join our dynamic team.We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.

Successful candidates must be able to commit to an onboarding date by the end of the year.

  • Build evaluation systems for image and video models/agents, covering generation quality, instruction following, multimodal understanding, and safety.
  • Develop automated metrics and model-based evaluators, and design reproducible human evaluation protocols.
  • Design and develop video generation/debugging agents that orchestrate multi-step creative workflows.
  • Build large-scale image and video data pipelines, and turn evaluation findings into model and product improvements.
Qualifications
Minimum Qualifications
  • Individuals who are completing or have recently completed a Bachelor's/ Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field.
  • Solid foundation in deep learning and computer vision, including generative modeling fundamentals.
  • Practical experience in at least one of: visual generation, multimodal LLMs, video understanding, or visual quality assessment.
  • Strong Python skills and proficiency with PyTorch or an equivalent framework, or multimodal evaluation framework.
  • Demonstrated research or engineering ability through publications, substantial projects, internships, or open-source work.
Preferred Qualifications
  • Publications at top-tier vision or ML venues, e.g., NeurIPS, ICML, CVPR, ICCV, ECCV, etc.
  • Hands-on experience with modern visual generation stacks, including diffusion-based models and their post-training.
  • Familiarity with visual generation benchmarks, or experience building evaluation frameworks.
  • Experience applying agent frameworks to creative workflows, or working with large-scale video data infrastructure.
Job Information
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an \"Always Day 1\" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Reasonable Accommodation

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Visual Generation & Multimodal Evaluation Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)
Visual Generation & Multimodal Evaluation Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

ByteDance • Seattle (WA)

On-site
USD 150,000 - 190,000
Visual Generation & Multimodal Evaluation Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start Bachelor/Master Graduate - 2027 Start San Jose Regular
Visual Generation & Multimodal Evaluation Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start Bachelor/Master Graduate - 2027 Start San Jose Regular

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 170,000
Visual Generation & Multimodal Evaluation Machine Learning Engineer Intern (AML-Ark-US) - 2027 [...]
Visual Generation & Multimodal Evaluation Machine Learning Engineer Intern (AML-Ark-US) - 2027 [...]

Pangle • San Jose (CA), Northern (KY)

On-site
USD 18,000 - 36,000
Visual Generation & Multimodal Evaluation Machine Learning Engineer Intern (AML-Ark-US) - 2027 [...]
Visual Generation & Multimodal Evaluation Machine Learning Engineer Intern (AML-Ark-US) - 2027 [...]

ByteDance • Seattle (WA)

On-site
USD 20,000 - 27,000
Visual Generation & Multimodal Evaluation Machine Learning Engineer Graduate (AML-Ark-US) - 202[...]
Visual Generation & Multimodal Evaluation Machine Learning Engineer Graduate (AML-Ark-US) - 202[...]

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Visual Generation & Multimodal Evaluation Machine Learning Engineer Intern (AML-Ark-US) - 2027 [...]
Visual Generation & Multimodal Evaluation Machine Learning Engineer Intern (AML-Ark-US) - 2027 [...]

ByteDance • San Jose (CA)

On-site
USD 51,000 - 73,000
Health insurance
Wellbeing benefits
Housing allowance
Visual Generation & Multimodal Evaluation Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)
Visual Generation & Multimodal Evaluation Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Health insurance
401(k) with company match
Paid parental leave
+6
Agent Evaluation & Evolution Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start
Agent Evaluation & Evolution Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start

ByteDance • Seattle (WA)

On-site
USD 120,000 - 180,000
Agent Evaluation & Evolution Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)
Agent Evaluation & Evolution Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

ByteDance • Seattle (WA)

On-site
USD 180,000 - 240,000
Agent Evaluation & Evolution Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start Bache[...]
Agent Evaluation & Evolution Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start Bache[...]

Bytedance • San Jose (CA)

On-site
USD 154,000 - 256,000