Remote AI Benchmark Engineer & Multilingual Specialist

Lilt

Town of Belgium (WI)

Hybrid

USD 90,000 - 150,000

Part time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote freelance
Flexible schedule
Competitive rates

Job summary

LILT is seeking experienced native-speaking software engineers to design, build, and validate multilingual benchmark tasks. You will create high-signal tasks in your native language, ensuring assets remain in the target language to measure multilingual handling without English translation crutches.

This is a remote, freelance opportunity with flexible scheduling. You’ll work on high-quality evaluation suites and participate in a rigorous QA process across creation, reviews, calibration, and

Qualifications

  • 5+ years of industry experience in software engineering.
  • Proven track record at leading technology companies and/or graduation from top-tier engineering universities.
  • Native or near-native fluency with deep understanding of grammar, register and phrasing rules. High English proficiency.
  • Strong proficiency in Python, standard shell scripting, and data processing.
  • Extensive experience with Terminal/CLI-based development workflows and familiarity with coding agents.
  • Deep technical understanding of multilingual text processing pitfalls, incl. encoding/decoding robustness, Unicode normalization, locale conventions, and RTL handling.

Responsibilities

  • Task Engineering: Evaluating Coding Agents.
  • Asset Creation: Build realistic task environments using native-language datasets and files.
  • Prompting & Translation: identify failure points in the native language.
  • Implementation & Verification: support robust solutions and deterministic verifier scripts.
  • Calibration & Execution: analyze logs and adjust task difficulty across model tiers.
  • Quality Assurance: participate in 4-layer human and automated checks to ensure fairness and benchmark integrity.

Skills

Python
Shell scripting
Data processing
CLI workflows
Multilingual text processing

Education

Higher education degree in CS

Tools

Coding agents

Job description

LILT is seeking experienced native-speaking software engineers to design, build, and validate multilingual benchmark tasks. You will create high-signal tasks in your native language, ensuring assets remain in the target language to measure multilingual handling without English translation crutches.

This is a remote, freelance opportunity with flexible scheduling. You’ll work on high-quality evaluation suites and participate in a rigorous QA process across creation, reviews, calibration, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmark Engineer | Native Language Specialist - French (Belgium) - Remote
AI Benchmark Engineer | Native Language Specialist - French (Belgium) - Remote

Lilt • Town of Belgium (WI)

Hybrid
USD 90,000 - 150,000
Remote freelance
Flexible schedule
Competitive rates
AI Engineer & Technical SME - Quant/Science/Corp
AI Engineer & Technical SME - Quant/Science/Corp

LILT AI • Germany (OH)

On-site
USD 85,000 - 120,000
Technical SME & AI Engineer — Quantitative & Scientific
Technical SME & AI Engineer — Quantitative & Scientific

LILT AI • Germany (OH)

On-site
USD 80,000 - 120,000
Remote AI Benchmark Technical Writer
Remote AI Benchmark Technical Writer

YO AI Labs • Maryland

Remote
USD 55,000 - 96,000
Senior AI Benchmark SME for Quant, Science & Corporate
Senior AI Benchmark SME for Quant, Science & Corporate

Lilt • United States

Remote
USD 90,000 - 130,000
Subject Matter Expert — AI Benchmarking (German/English)
Subject Matter Expert — AI Benchmarking (German/English)

LILT (Production) • United States

Remote
USD 70,000 - 100,000
Flexible working hours
Competitive pay
Access to diverse projects
Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 150,000 - 210,000
Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options
Remote AI Localization & Language Specialist
Remote AI Localization & Language Specialist

SDLC Technologies • United States

On-site
USD 30,000 - 60,000
Remote AI Benchmark Technical Writer (Contractor)
Remote AI Benchmark Technical Writer (Contractor)

YO AI Labs • California (MO)

Remote
USD 55,000 - 110,000
Remote Technical Writer for AI Benchmark Tasks
Remote Technical Writer for AI Benchmark Tasks

YO AI Labs • Town of Texas (WI)

Remote
USD 60,000 - 90,000