AI Benchmark Engineer Native Language Specialist - Chinese Hong Kong - Remote

LILT AI

Victoria

A distancia

ARS 83.977.000 - 178.451.000

Jornada completa

Hace 8 días
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Flexible schedule
Remote work
Competitive rates

Descripción de la vacante

Lilt seeks experienced native-speaking software engineers to design, build, and validate multilingual benchmarks. You will create high-signal tasks in your native language to test a model’s multilingual handling without relying on English translation.

This remote, freelance opportunity offers flexible scheduling and collaboration with a global team, focusing on robust evaluation suites and language-specific challenges across encoding, locale rules, and UI text processing.

Formación

  • 5+ years of industry experience in software engineering.
  • Proven track record at leading technology companies and/or graduation from top-tier engineering universities.
  • Native or near-native fluency, with a deep understanding of grammar, register, and phrasing rules. High English proficiency.
  • Strong proficiency in Python, standard shell scripting, and data processing.
  • Extensive experience with Terminal/CLI-based development workflows and familiarity with coding agents.
  • Deep technical understanding of multilingual text processing pitfalls, including encoding/decoding robustness, locale conventions, and Unicode normalization.

Responsabilidades

  • Task Engineering: Evaluating Coding Agents.
  • Asset Creation: Build realistic task environments using datasets and files in your native language.
  • Prompting & Translation: identify failure points where AI does not work in your native language.
  • Implementation & Verification: support reference implementations and write deterministic verifier scripts.
  • Calibration & Execution: analyze logs and calibrate task difficulty across model tiers.
  • Quality Assurance: participate in a 4-layer quality control process plus automated checks.

Conocimientos

Python
Shell scripting
Data processing
Multilingual NLP

Descripción del empleo

Job Description:

About The Opportunity

We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.

We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a models ability to handle multilingual environments without relying on English translation crutches.

Note this is a remote, freelance opportunity

What You’ll Deliver
  • Task Engineering: Evaluating Coding Agents.

  • Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling.

  • Prompting & Translation: finding failure points where AI does not work, in your native language

  • Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary).

  • Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).

  • Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.

Qualifications
  • Experience: 5+ years of industry experience in software engineering.

  • Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities.

  • Language: Native or near-native fluency, with a deep understanding of its grammar, register, and phrasing rules. High English proficiency.

  • Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing.

  • Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents.

  • Domain Expertise: Deep technical understanding of multilingual text processing pitfalls, including:

    • Encoding/decoding robustness and Unicode normalization.

    • Locale-dependent conventions (collation, casing, non-Gregorian dates).

    • Text I/O, toolchain interoperability, and safe string operations.

    • Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts. (For specific languages)

Why Collaborate with Lilt?
  • Your schedule, your rules. As an independent contractor, work when you want, as much or as little as you want. No fixed hours, no check-ins, no micromanaging.

  • Get paid quickly and fairly. We respect your time and your expertise. Competitive rates, prompt payments, no chasing invoices.

  • Work on projects that actually matter. Contribute to cutting-edge AI and language technology that is shaping how humans and machines communicate.

  • Be part of something bigger. Join a global community of linguists, subject matter experts, and language professionals who are advancing human knowledge together.

  • Grow without limits. As a Lilt contractor you get access to diverse, innovative projects that expand your portfolio and sharpen your skills across industries and domains.

  • Have fun doing what you love. Bring your language skills to life on projects that are as interesting as they are impactful. We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.

What to Consider Before Applying
  • Not ideal as a full time job or primary income source. Work availability fluctuates with project demand, making this better suited as a supplemental income stream. As a 1099 contractor, you wont receive benefits such as health insurance, paid time off, or retirement contributions, and hours are not guaranteed.

  • Requires reliable availability and commitment. Once you accept a task, we expect quality work and on-time delivery. Most tasks require a minimum of 2 hours per day or 15-20 hours per week. If your schedule is unpredictable, this may not be the right fit.

  • Geographic restrictions may apply. We cannot engage contractors in regions subject to international embargo or sanctions. As a 1099 contractor, you are solely responsible for your own tax obligations. We recommend consulting a tax professional before engaging.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

AI Training Contributor - Spanish (Latin America) - Remote
AI Training Contributor - Spanish (Latin America) - Remote

LILT, Inc. • Argentina

A distancia
ARS 51.974.000 - 103.948.000
Flexible schedule
Competitive rates
Global projects
+2
Remote Native Chinese AI Benchmark Engineer (Freelance)
Remote Native Chinese AI Benchmark Engineer (Freelance)

LILT AI • Victoria

A distancia
ARS 83.977.000 - 178.451.000
Flexible schedule
Remote work
Competitive rates
Remote Native-Language AI Benchmark Engineer
Remote Native-Language AI Benchmark Engineer

Lilt • Argentina

A distancia
ARS 124.818.000 - 249.637.000
Remote work
Flexible schedule
Prompt payments
Senior Software Engineer, AI Training - Argentina
Senior Software Engineer, AI Training - Argentina

G2i Inc. • Argentina

A distancia
ARS 209.054.000 - 418.107.000
Senior Full Stack AI Engineer - Remote Work | REF#304064
Senior Full Stack AI Engineer - Remote Work | REF#304064

BairesDev • Comisión de Fomento de Perú

Presencial
ARS 181.190.000 - 271.784.000
100% remote work
Competitive USD compensation
Home office setup provided
+3
Staff AI Engineer - Remote Work | REF#304055
Staff AI Engineer - Remote Work | REF#304055

BairesDev • Comisión de Fomento de Perú

Presencial
ARS 211.388.000 - 317.082.000
Remote work
Competitive salary USD
Home office hardware provided
+3
Staff Software Engineer, Artificial Intelligence/LLM
Staff Software Engineer, Artificial Intelligence/LLM

Beacon AI • San Carlos

Híbrido
ARS 274.156.000 - 365.542.000
Healthcare coverages
Time off (PTO)
401(k) plan
Senior Software Engineer, Artificial Intelligence/LLM
Senior Software Engineer, Artificial Intelligence/LLM

Beacon AI • San Carlos

Híbrido
ARS 228.464.000 - 319.849.000
Healthcare coverage
Paid time off
401(k)
Staff Software Engineer, Cloud Infrastructure
Staff Software Engineer, Cloud Infrastructure

Beacon AI • San Carlos

Híbrido
ARS 274.156.000 - 380.773.000
Healthcare
PTO
401(k)
AI Reviewer & Evaluator — Spanish (LatAm)
AI Reviewer & Evaluator — Spanish (LatAm)

LILT, Inc. • Argentina

A distancia
ARS 51.974.000 - 103.948.000
Flexible schedule
Competitive rates
Global projects
+2