Remote AI Benchmark Engineer — Multilingual Native Specialist

LILT (Production)

United States

Remote

USD 70,000 - 120,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

LILT (Production) is seeking experienced software engineers to design, build, and validate multilingual benchmarks for evaluating language models in native languages. You will create high-signal tasks, maintain assets in the target language, and test prompts without English translation crutches.

The role is remote and freelance, focusing on task engineering, dataset creation, and robust verifier scripts. Collaboration spans across a global linguistics and AI team, with flexible scheduling.

Qualifications

  • 1+ years of industry experience in software or prompt engineering.
  • Native or near-native fluency in the target language with high English proficiency.
  • Strong proficiency in Python, shell scripting, and data processing.
  • Extensive experience with Terminal/CLI-based development workflows and coding agents.
  • Deep understanding of multilingual text processing pitfalls.

Responsibilities

  • Task Engineering: Evaluating Coding Agents.
  • Asset Creation: Build language-native task environments.
  • Prompting & Translation: identify failure points in native language.
  • Implementation & Verification: develop reference implementations and verifier scripts.
  • Calibration & Execution: analyze logs and adjust task difficulty across model tiers.
  • Quality Assurance: participate in 4-layer quality control and automated checks.

Skills

Python
Shell scripting
Data processing
Terminal/CLI workflows

Job description

LILT (Production) is seeking experienced software engineers to design, build, and validate multilingual benchmarks for evaluating language models in native languages. You will create high-signal tasks, maintain assets in the target language, and test prompts without English translation crutches.

The role is remote and freelance, focusing on task engineering, dataset creation, and robust verifier scripts. Collaboration spans across a global linguistics and AI team, with flexible scheduling.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Benchmark Engineer & Multilingual Specialist
Remote AI Benchmark Engineer & Multilingual Specialist

Lilt • Town of Belgium (WI)

Hybrid
USD 90,000 - 150,000
Remote freelance
Flexible schedule
Competitive rates
Remote Native-Language AI Benchmark Engineer
Remote Native-Language AI Benchmark Engineer

LILT (Production) • United States

Remote
USD 83,000 - 138,000
Remote work
Flexible schedule
Global collaboration
AI Benchmark Engineer | Native Language Specialist - Chinese Mandarin - Remote
AI Benchmark Engineer | Native Language Specialist - Chinese Mandarin - Remote

LILT (Production) • United States

Remote
USD 83,000 - 138,000
Remote work
Flexible schedule
Global collaboration
Independent SME AI Engineer – Domain Benchmarking
Independent SME AI Engineer – Domain Benchmarking

Lilt • United States

Remote
USD 90,000 - 140,000
Independent contractor
Competitive rates
Global community
+1
AI Benchmark Engineer | Native Language Specialist - French (Belgium) - Remote
AI Benchmark Engineer | Native Language Specialist - French (Belgium) - Remote

Lilt • Town of Belgium (WI)

Remote
USD 90,000 - 150,000
Remote freelance
Flexible schedule
Competitive rates
Remote Senior AI Translation Engineer — Benchmark Lead
Remote Senior AI Translation Engineer — Benchmark Lead

Smartcat • United States

Remote
USD 140,000 - 180,000
Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 150,000 - 210,000
Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options
Remote Natural Sciences SME for AI Benchmarking
Remote Natural Sciences SME for AI Benchmarking

Lilt • United States

Remote
USD 83,000 - 138,000
Independent contractor
Flexible schedule
Competitive pay
+1
Remote AI Code Engineer — Benchmark & Validate Models
Remote AI Code Engineer — Benchmark & Validate Models

Turing • United States

Remote
USD 69,000 - 138,000
Remote AI Analyst: Multilingual Research & LLM Tuning
Remote AI Analyst: Multilingual Research & LLM Tuning

Turing • San Francisco (CA)

Remote
USD 55,000 - 83,000
Fully remote work environment
Opportunity for contract extension
Engagement with leading AI projects