Senior Applied Scientist: GenAI Eval & Memory (Remote)

Caseware

Bogotá

Presencial

COP 120.000.000 - 160.000.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Permanent contract
Prepaid medicine
Life insurance
Internet allowance
Home office stipend
Competitive compensation above market
Remote work

Descripción de la vacante

Caseware is a leading fintech company with vast experience in audit and accounting software. We seek an experienced applied scientist to design experiments, build evaluation methodologies, and develop AI-powered tools for an end-to-end agent builder platform.

You will contribute to memory systems and self-learning capabilities while upholding legal obligations and guiding junior staff. Join a remote-friendly team in Colombia and influence technical direction at a staff level, mentoring others

Formación

  • 6+ years (senior) to 8+ years (staff) of professional experience in applied science, machine learning, research, or data-intensive engineering.
  • 2+ years working on production GenAI or LLM-based systems (applications, agents, or the tooling and evaluation around them).
  • Proven experience designing and building evaluation frameworks for ML or LLM systems: metrics, benchmarks, scoring approaches, and rigorous experiment design.
  • A strong experimentation mindset. You form clear hypotheses, design sound experiments, and draw defensible conclusions from noisy, real-world data.
  • Ability to build and operate your own production tooling and services, ideally on AWS.
  • Strong understanding of GenAI system tradeoffs including quality, latency, cost, reliability, and safety.
  • A strong foundation in probability and statistics: experimental design, significance testing, and reasoning under uncertainty.
  • Strong English language communication and collaboration skills.
  • Comfortable operating in fast-moving environments with ambiguity and evolving requirements.

Responsabilidades

  • Design and run experiments that measure and improve the quality of LLM-based applications and agents, turning ambiguous quality questions into measurable, reproducible results.
  • Build out and own the evaluation methodology in practice that internal teams and customers depend on: benchmarks, regression suites, scoring methods, acceptance thresholds across task success, faithfulness, safety, latency, and cost, and the scientific foundations for comparing offline and online performance, partnering with platform engineering on the automated infrastructure that runs those comparisons and alerts on drift or regression, and wiring the signals into release gates.
  • Build AI-based tools and capabilities that power the end-to-end agent builder developer experience used by other Caseware teams and by our customers.
  • Build the synthetic data generation and eval-builder capabilities that let internal teams and customers create, evaluate, and deploy their own agents.
  • Turn domain procedures into task-plus-verifier structures, drawing on deep intuition for how LLM-based systems fail.
  • Contribute to applied science on agentic memory: use statistical methods grounded in audit domain knowledge to surface which experiences are worth remembering, and define and validate the promotion and demotion criteria that let agents compound that knowledge over time, partnering with platform engineering on the infrastructure that executes it.
  • Advance self-learning, self-improving systems that get better both offline and online as they are used.
  • Help ensure memory promotion and demotion criteria uphold the legal and contractual obligations owed to customers and their clients, partnering with Security, Legal, and Domain SMEs.
  • Design experiments and evaluations that demonstrate agent quality improves the more customers use the platform.
  • Translate advances in GenAI (models, retrieval, agent frameworks, evaluation techniques) into practical, maintainable capabilities.
  • At the staff level, influence technical direction through RFCs and design reviews, and mentor other scientists and engineers.

Conocimientos

Applied science
GenAI systems
Experimentation mindset
Production tooling
AWS familiarity
ML/LLM knowledge
English communication

Educación

PhD or MS in quantitative field

Herramientas

LangFuse
LangSmith
LangChain
LangGraph
AWS

Descripción del empleo

Caseware is a leading fintech company with vast experience in audit and accounting software. We seek an experienced applied scientist to design experiments, build evaluation methodologies, and develop AI-powered tools for an end-to-end agent builder platform.

You will contribute to memory systems and self-learning capabilities while upholding legal obligations and guiding junior staff. Join a remote-friendly team in Colombia and influence technical direction at a staff level, mentoring others

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Applied Scientist, GenAI & Evaluation Leader (Remote)
Senior Applied Scientist, GenAI & Evaluation Leader (Remote)

Alcor • Medellín

Presencial
COP 378.310.000 - 472.888.000
Life insurance
Funeral assistance
Home office stipend
+1
Senior GenAI Scientist - Evaluation & Agentic Memory Remote
Senior GenAI Scientist - Evaluation & Agentic Memory Remote

Caseware International • Bogotá

Presencial
COP 120.000.000 - 200.000.000
100% remote work
Home office stipend
Competitive compensation
+1
Remote Senior Agentic AI Scientist
Remote Senior Agentic AI Scientist

Caseware • Medellín

Presencial
COP 60.000.000 - 90.000.000
Remote work 100%
Competitive compensation
Training budget
+1
Senior AI Scientist: Evaluation & Agentic Memory
Senior AI Scientist: Evaluation & Agentic Memory

Caseware • Medellín

Presencial
COP 264.197.000 - 372.984.000
Remote work
Prepaid medicine
Life insurance
+5
Senior GenAI Scientist - Remote (Colombia)
Senior GenAI Scientist - Remote (Colombia)

Caseware • Medellín

Presencial
COP 285.307.000 - 443.810.000
Contrato a termino Indefinido
Prepaid Medicine
Life insurance and funeral assistance
+5
Senior Applied Scientist — AI Agents & Evaluation (Remote)
Senior Applied Scientist — AI Agents & Evaluation (Remote)

CaseWare • Bogotá

Presencial
COP 120.000.000 - 210.000.000
Remote work
Home office stipend
Prepaid medicine
+5
Senior AI Platform Engineer (Remote Colombia)
Senior AI Platform Engineer (Remote Colombia)

CaseWare • Bogotá

Presencial
COP 120.000.000 - 180.000.000
100% remote work environment
Competitive compensation
Home office stipend
+1
Senior AI Platform Engineer - Remote Colombia
Senior AI Platform Engineer - Remote Colombia

Alcor • Bogotá

Presencial
COP 120.000.000 - 180.000.000
Contrato a termino Indefinido
Prepaid Medicine
Life insurance and funeral assistance
+3
Senior AI Platform Architect - Remote (Colombia)
Senior AI Platform Architect - Remote (Colombia)

Caseware • Medellín

Presencial
COP 401.082.000 - 601.625.000
Contrato a término Indefinido with all
Prepaid Medicine
Life insurance and funeral assistance
+4
Senior AI Platform Engineer — Remote Colombia
Senior AI Platform Engineer — Remote Colombia

Caseware-International • Colombia

Presencial
COP 279.677.000 - 372.902.000
Permanent contract with all legal bene
Prepaid medical insurance