AI Benchmark Engineer | Native Language Specialist - Arabic (UAE) - Remote

LILT

United States

Remote

USD 83,000 - 152,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Flexible schedule
Competitive rates
Fast payments
Global community

Job summary

LILT seeks experienced software engineers to design, build, and validate multilingual benchmarks in a remote freelance capacity. You will create high-signal tasks in native languages and test model handling without English crutches.

You’ll work on task engineering, asset creation, prompting, and verification, with a rigorous 4-layer QA process and deterministic rubrics to ensure fairness and accuracy.

Qualifications

  • 5+ years of software engineering experience.
  • Proven track record at leading tech firms or top-tier engineering degree.
  • Native or near-native fluency with strong English proficiency.
  • Strong Python, shell scripting, and data processing skills.
  • Extensive experience with Terminal/CLI workflows and coding agents.
  • Deep understanding of multilingual text processing pitfalls (Unicode, locale rules, RTL).

Responsibilities

  • Task Engineering: Evaluating Coding Agents.
  • Asset Creation: Build native-language task environments with authentic datasets.
  • Prompting & Translation: identify failure points in your language.
  • Implementation & Verification: develop reference implementations and deterministic verifier scripts.
  • Calibration & Execution: analyze logs and adjust task difficulty across model tiers.
  • Quality Assurance: participate in 4-layer review and automated checks.

Skills

Python
Shell scripting
Data processing
Multilingual text processing
Terminal/CLI workflows
Coding agents

Education

Bachelor's degree or higher

Tools

CLI tooling

Job description

About The Opportunity

We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.

We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches.

Note this is a remote, freelance opportunity

What You’ll Deliver

- Task Engineering: Evaluating Coding Agents.
- Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling.
- Prompting & Translation: finding failure points where AI does not work, in your native language
- Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary).
- Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).
- Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.

Qualifications

- Experience: 5+ years of industry experience in software engineering.
- Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities.
- Language: Native or near-native fluency, with a deep understanding of its grammar, register, and phrasing rules. High English proficiency.
- Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing.
- Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents.
- Domain Expertise: Deep technical understanding of multilingual text processing pitfalls, including:

Encoding/decoding robustness and Unicode normalization.
- Locale-dependent conventions (collation, casing, non-Gregorian dates).
- Text I/O, toolchain interoperability, and safe string operations.
- (For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts.

Why Collaborate with Lilt?

- Your schedule, your rules. As an independent contractor, work when you want, as much or as little as you want. No fixed hours, no check-ins, no micromanaging.
- Get paid quickly and fairly. We respect your time and your expertise. Competitive rates, prompt payments, no chasing invoices.
- Work on projects that actually matter . Contribute to cutting-edge AI and language technology that is shaping how humans and machines communicate.
- Be part of something bigger. Join a global community of linguists, subject matter experts, and language professionals who are advancing human knowledge together.
- Grow without limits. As a Lilt contractor you get access to diverse, innovative projects that expand your portfolio and sharpen your skills across industries and domains.
- Have fun doing what you love. Bring your language skills to life on projects that are as interesting as they are impactful.We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.

What to Consider Before Applying

- Not ideal as a full time job or primary income source. Work availability fluctuates with project demand, making this better suited as a supplemental income stream. As a 1099 contractor, you won't receive benefits such as health insurance, paid time off, or retirement contributions, and hours are not guaranteed.
- Requires reliable availability and commitment. Once you accept a task, we expect quality work and on-time delivery. Most tasks require a minimum of 2 hours per day or 15-20 hours per week. If your schedule is unpredictable, this may not be the right fit.
- Geographic restrictions may apply. We cannot engage contractors in regions subject to international embargo or sanctions. As a 1099 contractor, you are solely responsible for your own tax obligations. We recommend consulting a tax professional before engaging.

AI is changing how the world communicates — and LILT is leading that transformation.

LILT's mission is to make the world's information available to everyone, no matter the language they speak. Join our global community who thrive on innovation and excellence. Our collective knowledge, uniqueness, and skills deliver multilingual AI and human-verified services to Enterprises, Governments, and AI Developers around the world.

Earn money. Have fun. Advance human knowledge. Work on diverse projects from anywhere, any time you want. Get paid quickly and fairly, and build your professional network in a supportive community—all through a streamlined application process tailored to your expertise.

Information collected and processed as part of your application process, including any job applications you choose to submit, is subject to LILT's Privacy Policy at https://lilt.com/legal/privacy .

At LILT, we are committed to a fair, inclusive, and transparent hiring process. As part of our recruitment efforts, we may use artificial intelligence (AI) and automated tools to assist in the evaluation of applications, including résumé screening, assessment scoring, and interview analysis. These tools are designed to support human decision-making and help us identify qualified candidates efficiently and objectively. All final hiring decisions are made by people. If you have any concerns, require accommodations, or would like to opt-out of the use of AI in our hiring process, please let us know at recruiting@lilt.com.

LILT is an equal opportunity employer. We extend equal opportunity to all individuals without regard to an individual’s race, religion, color, national origin, ancestry, sex, sexual orientation, gender identity, age, physical or mental disability, medical condition, genetic characteristics, veteran or marital status, pregnancy, or any other classification protected by applicable local, state or federal laws. We are committed to the principles of fair employment and the elimination of all discriminatory practices.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Benchmark Engineer | Native Language Specialist - Arabic (Saudi Arabia) - Remote
AI Benchmark Engineer | Native Language Specialist - Arabic (Saudi Arabia) - Remote

LILT • United States

Remote
USD 83,000 - 152,000
Remote freelance opportunity
Competitive rates
AI Benchmark Engineer | Native Language Specialist - Chinese (Taiwan) - Remote
AI Benchmark Engineer | Native Language Specialist - Chinese (Taiwan) - Remote

LILT • United States

Remote
USD 120,000 - 180,000
AI Benchmark Engineer | Native Language Specialist - Chinese (Hong Kong) - Remote
AI Benchmark Engineer | Native Language Specialist - Chinese (Hong Kong) - Remote

LILT • United States

Remote
USD 96,000 - 165,000
Remote work
Flexible schedule
Competitive rates
+1
AI Benchmark Engineer | Native Language Specialist - French (Belgium) - Remote
AI Benchmark Engineer | Native Language Specialist - French (Belgium) - Remote

Lilt • Town of Belgium (WI)

Remote
USD 90,000 - 150,000
Remote freelance
Flexible schedule
Competitive rates
AI Benchmark Engineer | Native Language Specialist - French (France) - Remote
AI Benchmark Engineer | Native Language Specialist - French (France) - Remote

LILT • United States

Remote
USD 83,000 - 165,000
Flexible schedule
Competitive freelance rates
Global project exposure
+3
AI Benchmark Engineer | Native Language Specialist - Chinese Mandarin - Remote
AI Benchmark Engineer | Native Language Specialist - Chinese Mandarin - Remote

LILT (Production) • United States

Remote
USD 83,000 - 138,000
Remote work
Flexible schedule
Global collaboration
AI Training Contributor - English (Great Britain) - Remote
AI Training Contributor - English (Great Britain) - Remote

LILT (Production) • United States

Remote
USD 60,000 - 100,000
Content Writer - Tamil - Remote
Content Writer - Tamil - Remote

LILT (Production) • United States

Remote
USD 25,000 - 70,000
Content Writer - Somali - Remote
Content Writer - Somali - Remote

LILT • United States

Remote
USD 35,000 - 60,000
Content Writer - Kirundi - Remote
Content Writer - Kirundi - Remote

LILT (Production) • United States

Remote
USD 50,000 - 75,000
Flexible schedule
Competitive pay
Meaningful projects
+3