Remote Native-Language AI Benchmark Engineer

Lilt

Argentina

A distancia

ARS 124.818.000 - 249.637.000

A tiempo parcial

Hace 6 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Remote work
Flexible schedule
Prompt payments

Descripción de la vacante

LILT is seeking experienced software engineers to design, build, and validate benchmarks for multilingual models. This remote, freelance role focuses on creating high-quality tasks in the candidate’s native language, evaluating agents, and writing deterministic verifier scripts.

You’ll work across data, prompts, and evaluation rubrics, collaborating with a global linguistics and ML team. Ideal candidates have 5+ years of software engineering experience, strong Python and shell skills, and a deep

Formación

  • 5+ years of software engineering experience.
  • Proven track record at leading tech companies or top-tier universities.
  • Native or near-native fluency in the target language with high English proficiency.
  • Strong proficiency in Python, shell scripting, and data processing.

Responsabilidades

  • Task Engineering: evaluating coding agents.
  • Asset Creation: build datasets and files in the native language.
  • Prompting & Translation: identify failure points in the native language.
  • Implementation & Verification: develop reference implementations and verifier scripts.
  • Calibration & Execution: analyze logs and calibrate task difficulty across model tiers.
  • Quality Assurance: participate in a 4-layer QA process.

Conocimientos

Python
Shell scripting
Data processing
Terminal/CLI workflows
Multilingual text processing
Unicode handling

Educación

Engineering degree

Descripción del empleo

LILT is seeking experienced software engineers to design, build, and validate benchmarks for multilingual models. This remote, freelance role focuses on creating high-quality tasks in the candidate’s native language, evaluating agents, and writing deterministic verifier scripts.

You’ll work across data, prompts, and evaluation rubrics, collaborating with a global linguistics and ML team. Ideal candidates have 5+ years of software engineering experience, strong Python and shell skills, and a deep

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

LLM Evaluation Scientist: Frontier Benchmarking
LLM Evaluation Scientist: Frontier Benchmarking

Anyone AI • Buenos Aires

Presencial
ARS 178.909.000 - 268.365.000
Remote AI Quality Evaluator & Support for LLM Benchmarking
Remote AI Quality Evaluator & Support for LLM Benchmarking

Innodata Inc. • Argentina

Presencial
ARS 29.181.000 - 52.109.000
Remote Freelance AI Evaluation Engineer
Remote Freelance AI Evaluation Engineer

AI Chopping Block • Buenos Aires

Híbrido
ARS 31.166.000 - 62.331.000
Flexible hours
Remote freelance project
AI Engineer
AI Engineer

Aqusag-Technologies • Argentina

A distancia
USD 120.000 - 170.000
Freelance Agent Evaluation Engineer
Freelance Agent Evaluation Engineer

AI Chopping Block • Buenos Aires

Híbrido
ARS 31.166.000 - 62.331.000
Flexible hours
Remote freelance project
Remote Spanish Linguistic Specialist, AI Data & Evaluation
Remote Spanish Linguistic Specialist, AI Data & Evaluation

Mercor • Buenos Aires

A distancia
ARS 42.127.000 - 75.828.000
Lead AI Engineer — Remote RAG/NLP/LLM Systems
Lead AI Engineer — Remote RAG/NLP/LLM Systems

N-iX • Argentina

Híbrido
ARS 83.351.000 - 111.136.000
Senior Python Backend Engineer - AI/ML & LLMs, Remote
Senior Python Backend Engineer - AI/ML & LLMs, Remote

Intellectsoft • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Awesome projects with an impact
Udemy courses of your choice
Team-building events
+2
Data & Machine Learning Engineer
Data & Machine Learning Engineer

IDT • Buenos Aires

Presencial
ARS 104.200.779 - 133.972.431
Senior AI/ML Engineer - Build LLM Agents & Guardrails
Senior AI/ML Engineer - Build LLM Agents & Guardrails

Ciklum • Argentina

Presencial
ARS 3.000.000 - 6.000.000
Medical insurance
Mental health programs
Udemy license
+1