ML Engineer — LLM Evaluation

Capitolis

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation techniques, along with the ability to adapt to new findings. Join a top AI startup committed to responsible development and user privacy in machine learning.

Qualifications

  • Knowledge of LLM evaluation and data curation methods.
  • Experience in designing LLM benchmarking methods.
  • Ability to shift focus as new findings emerge.

Responsibilities

  • Own LLM evaluation processes and methods for benchmarks.
  • Generate synthetic data and conduct benchmarking.
  • Deliver scalable and reproducible production code.
  • Develop new benchmarking methods for safety and helpfulness.
  • Co-author academic papers and presentations.

Skills

LLM evaluation
Data curation techniques
Designing benchmarking methods
Adaptability and flexibility

Job description

At Dynamo AI, we believe that LLMs must be developed with safety, privacy, and real-world responsibility in mind. Our ML team comes from a culture of academic research driven to democratize AI advancements responsibly. By operating at the intersection of ML research and industry applications, our team empowers Fortune 500 companies’ adoption of frontier research for their next generation of LLM products.

Join us if you:

  • Wish to work on the premier platform for private and personalized LLMs. We provide the fastest end to end solution to deploy research in the real world with our fast-paced team of ML Ph.D.’s and builders, free of Big Tech / academic bureaucracy and constraints.
  • Are excited at the idea of democratizing state-of-the-art research on safe and responsible AI.
  • Are motivated to work at a 2023 CB Insights Top 100 AI Startup and see your impact on end customers in the timeframe of weeks not years.
  • Care about building a platform to empower fair, unbiased, and responsible development of LLMs and don’t accept the status quo of sacrificing user privacy for the sake of ML advancement.
Responsibilities
  • Own LLM evaluation processes and methods with a focus on generating benchmarks representative of real-world usage and safety vulnerabilities.
  • Generate high quality synthetic data, curate labels, and conduct rigorous benchmarking.
  • Deliver robust, scalable, and reproducible production code.
  • Push the envelope by developing methods for benchmarking that revamps how we assess the best LLMs for harmlessness and helpfulness. Your research will directly empower our customers to more feasibly deploy safe and responsible LLMs.
  • Co-author papers, patents, and presentations with our research team by integrating other members’ work with your vertical.
Qualifications
  • Domain knowledge in LLM evaluation and data curation techniques.
  • Extensive experience in designing and implementing LLM benchmarking, extending previous methods. Comfortability with leading end-to-end projects.
  • Adaptability and flexibility. In both the academic and startup world, a new finding in the community may necessitate an abrupt shift in focus. You must be able to learn, implement, and extend state-of-the-art research.
  • Preferred: past research or projects in benchmarking LLMs.

Dynamo AI is committed to maintaining compliance with all applicable local and state laws regarding job listings and salary transparency. This includes adhering to specific regulations that mandate the disclosure of salary ranges in job postings or upon request during the hiring process. We strive to ensure our practices promote fairness, equity, and transparency for all candidates.

Salary for this position may vary based on several factors, including the candidate's experience, expertise, and the geographic location of the role. Compensation is determined to ensure competitiveness and equity, reflecting the cost of living in different regions and the specific skills and qualifications of the candidate.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Evaluation & Benchmarking Engineer
LLM Evaluation & Benchmarking Engineer

Capitolis • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
ML Engineer - LLM Evaluation & Automation
ML Engineer - LLM Evaluation & Automation

Grid Dynamics • United States

On-site
USD 140,000 - 170,000
Flexible schedule
Medical insurance
Vision and dental
+3
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
LLM Applications Engineer
LLM Applications Engineer

SupportFinity™ • New York (NY)

Hybrid
USD 130,000 - 175,000
AI Engineer - OH
AI Engineer - OH

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
ML Engineer: LLM Evaluation & Observability
ML Engineer: LLM Evaluation & Observability

Gleanwork • Mountain View (CA)

Hybrid
USD 200,000 - 300,000
Health insurance
401(k) plan
Home office improvement stipend
+3
LLM Evaluation Engineer
LLM Evaluation Engineer

ThirdLaw | Runtime AI Safety • United States

Hybrid
USD 120,000 - 160,000
Market cash compensation
Above-market equity
Generous benefits
AI Engineer
AI Engineer

AllyNd Partners • Chicago (IL)

Remote
USD 120,000 - 160,000
Annual learning stipend
Flexible hours
Comprehensive health benefits
+1