AI Benchmark Engineer | Native Language Specialist Arabic (UAE)

LILT

Abu Dhabi

Remote

AED 202,000 - 405,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Flexible schedule
Prompt payments
Work on impactful AI benchmarks
Global community of language experts

Job summary

LILT is building a rigorous verifiable evaluation suite to test multilingual software challenges and model robustness across non-English data processing. We seek experienced native-speaking software engineers to design, build, and validate benchmarks that avoid English translation crutches.

This remote freelance opportunity focuses on creating high-signal tasks, translating prompts, and developing reliable verifier scripts while participating in a multi-layer QA process.

Qualifications

  • 5+ years of industry software engineering experience.
  • Native or near-native fluency with high English proficiency.
  • Strong proficiency in Python, shell scripting and data processing.

Responsibilities

  • Task Engineering: Evaluating Coding Agents.
  • Asset Creation: Build realistic task environments using datasets in your native language.
  • Prompting & Translation: identify failure points where AI struggles in your native language.
  • Implementation & Verification: develop robust reference implementations and deterministic verifier scripts.
  • Calibration & Execution: analyze logs and tune task difficulty across model tiers.
  • Quality Assurance: participate in 4-layer human QA plus automated checks for fairness and accuracy.

Skills

Python
Shell scripting
Data processing
Terminal/CLI workflows
Multilingual NLP

Education

Engineering degree from a top-tier university

Job description

About The Opportunity

We are building a rigorous verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects non-English data processing and complex locale/encoding edge cases in terminal workflows.

We are seeking experienced native-speaking software engineers to design build and validate these benchmarks. You will create high-signal high-quality tasks that genuinely test a models ability to handle multilingual environments without relying on English translation crutches.

Note this is a remote freelance opportunity
What Youll Deliver
  • Task Engineering: Evaluating Coding Agents.

  • Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially these assets must remain in the target language to genuinely measure multilingual handling.

  • Prompting & Translation: finding failure points where AI does not work in your native language

  • Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable deterministic verifier scripts (using rubric-based judging only when strictly necessary).

  • Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku Sonnet Opus).

  • Quality Assurance: Participate in a rigorous 4-layer human quality control process (creation human review calibration review and audit) alongside automated LLM-based checks to ensure fairness grammatical accuracy and benchmark integrity.

Qualifications
  • Experience: 5 years of industry experience in software engineering.

  • Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities.

  • Language: Native or near-native fluency with a deep understanding of its grammar register and phrasing rules. High English proficiency.

  • Technical Stack: Strong proficiency in Python standard shell scripting and data processing.

  • Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents.

  • Domain Expertise: Deep technical understanding of multilingual text processing pitfalls including:

  • Encoding/decoding robustness and Unicode normalization.

  • Locale-dependent conventions (collation casing non-Gregorian dates).

  • Text I/O toolchain interoperability and safe string operations.

  • (For specific languages) Bidirectional/RTL handling font fallbacks and rendering/typography in UI or artifacts.

Why Collaborate with Lilt
  • Your schedule your rules. As an independent contractor work when you want as much or as little as you want. No fixed hours no check-ins no micromanaging.

  • Get paid quickly and fairly. We respect your time and your expertise. Competitive rates prompt payments no chasing invoices.

  • Work on projects that actually matter. Contribute to cutting-edge AI and language technology that is shaping how humans and machines communicate.

  • Be part of something bigger. Join a global community of linguists subject matter experts and language professionals who are advancing human knowledge together.

  • Grow without limits. As a Lilt contractor you get access to diverse innovative projects that expand your portfolio and sharpen your skills across industries and domains.

  • Have fun doing what you love. Bring your language skills to life on projects that are as interesting as they are are building a rigorous verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects non-English data processing and complex locale/encoding edge cases in terminal workflows.

What to Consider Before Applying
  • Not ideal as a full time job or primary income source. Work availability fluctuates with project demand making this better suited as a supplemental income stream. As a 1099 contractor you wont receive benefits such as health insurance paid time off or retirement contributions and hours are not guaranteed.

  • Requires reliable availability and commitment. Once you accept a task we expect quality work and on-time delivery. Most tasks require a minimum of 2 hours per day or 15-20 hours per week. If your schedule is unpredictable this may not be the right fit.

  • Geographic restrictions may apply. We cannot engage contractors in regions subject to international embargo or sanctions. As a 1099 contractor you are solely responsible for your own tax obligations. We recommend consulting a tax professional before engaging.

Information collected and processed as part of your application process including any job applications you choose to submit is subject to LILTs Privacy Policy at

LILT is an equal opportunity employer. We extend equal opportunity to all individuals without regard to an individuals race religion color national origin ancestry sex sexual orientation gender identity age physical or mental disability medical condition genetic characteristics veteran or marital status pregnancy or any other classification protected by applicable local state or federal laws. We are committed to the principles of fair employment and the elimination of all discriminatory practices.


Required Experience:

IC

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Arabic AI Benchmark Engineer (Native)
Remote Arabic AI Benchmark Engineer (Native)

LILT • Abu Dhabi

Remote
AED 202,000 - 405,000
Flexible schedule
Prompt payments
Work on impactful AI benchmarks
+1
AI Engineer (Applied)
AI Engineer (Applied)

BlackStone eIT • Dubai

On-site
AED 120,000 - 150,000
Paid Time Off
Performance Bonus
Training & Development
AI Trainers Network - English
AI Trainers Network - English

Jobgether SRL • United Arab Emirates

On-site
AED 89,000 - 134,000
Flexible projects
Remote work
Global network
+2
Agentic AI Engineer | Systems Ltd | Dubai, UAE
Agentic AI Engineer | Systems Ltd | Dubai, UAE

Systems Ltd • Dubai

On-site
AED 360,000 - 600,000
Bilingual Customer Support Specialist - AI Trainer
Bilingual Customer Support Specialist - AI Trainer

DataAnnotation • United Arab Emirates

Remote
AED 126,000 - 202,000
AI Engineer (Vietnamese speaker)
AI Engineer (Vietnamese speaker)

AGAPI Technologies • United Arab Emirates

On-site
AED 180,000 - 280,000
Competitive compensation
Visa processing and UAE benefits
Generous holidays
+3
Remote Coding Expertise for AI Training
Remote Coding Expertise for AI Training

Outlier AI, Inc. • United Arab Emirates

On-site
MXN 465,000 - 745,000
Forward deployed engineer (Software Engineer)
Forward deployed engineer (Software Engineer)

Instrumental • Dubai

Hybrid
AED 469,000 - 692,000
Health insurance
Paid annual leave
Ticket
AI Specialist- Native Arabic Speakers
AI Specialist- Native Arabic Speakers

Confidential Company • Abu Dhabi

On-site
AED 320,000 - 520,000
AI Engineer – Computer-Use Agents
AI Engineer – Computer-Use Agents

Confidential Careers • Abu Dhabi

On-site
AED 240,000 - 420,000