Frontend Code Evaluation Specialist

Obsidian

Santiago

Presencial

CLP 13.392.000 - 23.436.000

Jornada completa

14 días+
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Obsidian is seeking an experienced web evaluator to compare reference pages against two model attempts across a full desktop viewport. You will document precise layout structures, spacing, typography, and asset handling, then justify every preference with evidence from the source pages.

You will review responsiveness, predict how changes affect rendering, and provide a detailed, example-driven rationale for your judgments, mirroring a professional frontend critique workflow.

Formación

  • 3–8 years of professional web development experience.
  • Fluency across web eras and traditional markup (HTML/CSS, tables, etc.).
  • Strong hand-written HTML/CSS with semantic markup knowledge.
  • Browser DevTools as core skill for layout/DOM inspection.
  • Ability to read and reason about legacy layouts and non-component structures.
  • Proficient written English for detailed justifications.

Responsabilidades

  • Render and compare reference pages with two model attempts at 1920×1080.
  • Diff visual fidelity: box model, typography, color, borders, images, and z-order.
  • Read source of attempts and judge construction quality vs rendering accuracy.
  • Test responsiveness and explain where layouts break or adapt poorly.
  • Provide precise, evidence-based justifications for each preference.

Conocimientos

Frontend web development
HTML/CSS expertise
DevTools proficiency
JavaScript understanding
English writing

Herramientas

Local server setup
HTML/CSS markup reading

Descripción del empleo

About the work

We're building a high-quality dataset of human preference judgments on AI-generated frontend code. Each task hands you a reference web page — crawled from the real internet, delivered as a full zipped site tree plus screenshots of its default view and, on some pages, additional states reached by hovering, clicking, or scrolling. Alongside it come two model attempts, A and B, each a zipped self-contained site tree. The models only ever saw the screenshots; they never had the source.

You download all three, run them locally, view each at a 1920×1080 viewport, interact with them to reach every required state, then open the source of both attempts and grade them against each other — on visual fidelity per state, and on how the code is actually constructed. Structure and responsiveness are explicitly part of the rubric, not just the render.

This is evaluation work, not authoring. The defining skill is not that you can build a page — it's that you can open someone else's page and tell how it was built and where it cheats.

Please read before applying
  • Each unit takes roughly 2–3 hours and is timed. This is not microtask work; if you can only offer scattered 15-minute windows, you will not be able to finish a unit.

  • You need a real local development environment. A tablet, a Chromebook, or a locked-down work machine that cannot run a local static server will not work for this project.

  • You need at least one completed Mercor engagement, delivered in full. We are not onboarding net-new experts to this project.

What you'll do
  • Render a reference page and two candidate replications at 1920×1080 and judge which is the closer reproduction, state by state.

  • Diff visual fidelity in detail: box model and spacing, typography (family, size, weight, line-height, letter-spacing), color and border treatment, image and asset handling, z-order and overflow.

  • Read the source of both attempts and grade construction quality — distinguishing a replication that is genuinely correct from one that merely looks correct at one viewport. Hardcoded pixel offsets, absolute positioning standing in for real layout, inline style soup, a single undifferentiated div tree, or a screenshot pasted in as an instead of a rebuilt section.

  • Test responsiveness: a nav bar that looks right at 1920px but collapses at 1400px is a defect, and you should be able to say precisely why.

  • Write a specific, evidence-cited justification for every preference. We need "B nests the article body in a single absolutely-positioned div, so the text overlaps the footer below 1600px, while A uses normal document flow" — not "A looks closer."

  • Use the "this task is broken" escape hatch with judgment: distinguish a task that genuinely cannot be completed from the screenshots provided from one that is merely hard. Over-flagging and under-flagging are both failure modes.

You're a fit if you have
  • 3–8 years of professional web development experience, shipping web interfaces for a living, primarily in frontend or full-stack work.

  • Fluency across web eras. The reference pages are real crawled sites — one is a small charter-fishing business built in the table-and-image-map tradition, another a corporate press-release page with stacked navigation rows and social share widgets. If your entire career happened inside a modern component framework and you have never authored raw CSS or seen a table used for layout, you will misjudge many of these pages.

  • Command of hand-written HTML and CSS: semantic markup, flexbox, grid, media queries, and legacy float- and table-based layouts you can read and reason about. You should be able to look at a rendered layout and predict what's holding it together before opening DevTools.

  • Browser DevTools as muscle memory — setting an exact viewport, walking the element tree, checking computed styles, watching what a hover handler mutates.

  • Command-line comfort: unzipping an archive, standing up a static local server because the relative asset paths demand it, and untangling a broken image reference rather than giving up and grading from the screenshot.

  • Enough JavaScript to read a page's scripts and understand what they do to the DOM, even if you don't write JS daily.

  • Professional written English. Every task ends in a free-text justification, and a rating without a specific rationale is worth very little.

Equipment
  • A desktop or laptop that displays a 1920×1080 viewport.

  • Administrator rights on your own machine, so you can install and run a local server.

Nice to have
  • Prior RLHF, preference-labeling, model-evaluation, or structured code-review work — the strongest single signal. Rubric-driven comparison at volume needs almost no ramp here.

  • Pixel-perfect design-to-code experience: agency work, design systems, template production. Anyone who has had a designer reject a build over four pixels has exactly the fidelity eye this needs.

  • Accessibility expertise (ARIA, semantic landmarks, heading hierarchy) — you'll notice immediately when an attempt renders a heading as a styled ..

  • Familiarity with how LLMs fail at code generation.

  • Web scraping, archiving, or DOM-parsing background — comfort with messy crawled site trees.

  • More than one completed Mercor project, and availability in contiguous multi-hour blocks.

Note:

this seat is for practicing web developers. Backend-only, ML/data-science-only, mobile-native-only, and DevOps-only engineers do not have the UI instincts this requires, however strong they are otherwise. Designers who do not code cannot grade the source axes at all. Framework-only engineers who have never authored CSS outside a component library will struggle with the legacy reference pages.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Frontend Code Auditor — UI Fidelity & Layout Critique
Frontend Code Auditor — UI Fidelity & Layout Critique

Obsidian • Santiago

Presencial
CLP 13.392.000 - 23.436.000
AI Engineer, Automation and Developer Tooling Santiago, Chile · Remote
AI Engineer, Automation and Developer Tooling Santiago, Chile · Remote

Arcadia Power, Inc. • Santiago

Híbrido
CLP 74.931.000 - 107.308.000
Remote first
Flexible PTO
11 holidays
+5
Full-Stack Developer (AI Focused)
Full-Stack Developer (AI Focused)

Brand & Bot • Chile

A distancia
CLP 56.444.000 - 94.073.000
Sr Web Analyst | Remote For Marketing Agency For Education
Sr Web Analyst | Remote For Marketing Agency For Education

Atomic Hr • Concepcion

Híbrido
CLP 41.876.000 - 83.752.000
Remote work
Senior AI Research Engineer
Senior AI Research Engineer

Improving South America • Chile

Presencial
CLP 83.488.000 - 111.317.000
Contrato a largo plazo
100% Remoto
Vacaciones y PTOs
+7
Product Designer
Product Designer

Jobtailor • Santiago

Presencial
CLP 26.000.000 - 38.000.000
Fusion 360 CAM Programming Expert
Fusion 360 CAM Programming Expert

Mercor • Santiago

Presencial
CLP 20.088.000 - 33.480.000
Mid-Level Frontend Developer (Blazor) - Remote - Latin America
Mid-Level Frontend Developer (Blazor) - Remote - Latin America

FullStack • Santiago

Presencial
CLP 65.851.000 - 94.073.000
Competitive pay
100% remote work
Continuing education opportunities
+1
Mid-Level Frontend Developer (Blazor) - Remote - Latin America
Mid-Level Frontend Developer (Blazor) - Remote - Latin America

FullStack • Concepcion

A distancia
CLP 83.488.000 - 111.317.000
100% remote work
Competitive pay
Continuing education
WordPress Developer (AI First)
WordPress Developer (AI First)

Brand & Bot • Chile

A distancia
CLP 20.088.000 - 31.248.000
Remote work
AI tools access
Flexible PTO
+2