Senior Applied Scientist

Caseware International

Bogotá

Presencial

COP 120.000.000 - 200.000.000

Jornada completa

Hace 10 días

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

100% remote work
Home office stipend
Competitive compensation
Flexible work options

Descripción de la vacante

Caseware International is hiring for a staff-to-senior level applied science role focused on evaluating and improving GenAI-powered audit and accounting tools. You will design experiments, build evaluation frameworks, and develop AI-driven capabilities behind our agent platforms.

The role emphasizes self-learning systems with strong collaboration across Security, Legal, and Domain SMEs. The position is fully remote, based in Colombia, with opportunities to influence technical direction and

Formación

  • 6+ to 8+ years of professional experience in applied science, ML, research, or data-intensive engineering.
  • 2+ years working on production GenAI or LLM-based systems (applications, agents, or tooling).
  • Proven experience designing evaluation frameworks: metrics, benchmarks, scoring, and experiments.
  • Ability to build and operate production tooling, ideally on AWS.
  • Strong English communication and collaboration skills.
  • Comfortable operating in fast-moving environments with ambiguity.

Responsabilidades

  • Design and run experiments to measure and improve quality of LLM-based applications and agents.
  • Own evaluation methodology: benchmarks, regression suites, thresholds for success, and drift monitoring.
  • Build AI-based tools powering the end-to-end agent builder experience.
  • Develop synthetic data generation and evaluation capabilities for internal teams and customers.
  • Translate domain procedures into task-plus-verifier structures and validate memory promotion criteria.
  • Contribute to self-learning systems that improve offline and online with guardrails for legal obligations.
  • Influence technical direction via RFCs and design reviews; mentor other scientists and engineers.

Conocimientos

Applied science
Machine learning
Research
Data engineering
Experiment design

Educación

MS in quantitative field
PhD preferred

Herramientas

AWS
LangChain
LangSmith
OpenSearch
Textract

Descripción del empleo

Caseware is one of Canada's original Fintech companies, having led the global audit and accounting software industry for over 30 years, with more than 500,000 users across 130 countries and available in 16 different languages. While you might not have heard of us over 36,000 accounting and audit professionals list Caseware as a skill on their LinkedIn profiles!

We are building the agentic platform that powers Caseware's next generation of audit and accounting products, and applied science is how we measure and improve quality across our agents. This is a role for someone who thinks in experiments, has built evaluations for LLM-based systems, and wants to build the tools that other teams and our customers rely on to ship their own agents.

You will design and run experiments that turn ambiguous quality questions into measurable results, and own the evaluation methodology that internal teams and customers depend on. You will build the AI-based capabilities behind our end-to-end agent builder, including the synthetic data generation and eval builders that let teams and customers create, evaluate, and deploy their own agents. You will also work on something genuinely novel: self-learning, self-improving systems that get better both offline and online. A significant part of that is agentic memory: deciding what is worth remembering and validating the criteria that let agents compound useful knowledge on behalf of customers over time, while rigorously upholding the legal and contractual obligations owed to them and their clients. At the staff level, you will also influence technical direction and mentor others.

Location: This is a fully remote position located in Colombia.
Maira Russo- Senior Talent Acquisition Partner

The domain: financial audit

This role sits squarely in the financial audit domain, and that domain matters more than most people realize. Independent audit is one of the quiet foundations of the global economy. When investors, lenders, regulators, and the public can trust that a company's financial statements are accurate, capital can flow, markets can function, and organizations can be held accountable. That trust rests on the quality of audit and assurance work, which makes the tools that audit professionals use genuinely consequential.

You do not need prior knowledge of accounting or auditing to succeed here. What you will gain is deep, durable expertise in how audits are performed and where AI can make them faster, more reliable, and more insightful. You will learn the domain on the job alongside experienced domain experts, and that fluency will make you a stronger applied scientist. If you already bring that background, even better.

What you'll be doing:
  • Design and run experiments that measure and improve the quality of LLM-based applications and agents, turning ambiguous quality questions into measurable, reproducible results.

  • Build out and own the evaluation methodology in practice that internal teams and customers depend on: benchmarks, regression suites, scoring methods, acceptance thresholds across task success, faithfulness, safety, latency, and cost, and the scientific foundations for comparing offline and online performance, partnering with platform engineering on the automated infrastructure that runs those comparisons and alerts on drift or regression, and wiring the signals into release gates.

  • Build AI-based tools and capabilities that power the end-to-end agent builder developer experience used by other Caseware teams and by our customers.

  • Build the synthetic data generation and eval-builder capabilities that let internal teams and customers create, evaluate, and deploy their own agents.

  • Turn domain procedures into task-plus-verifier structures, drawing on deep intuition for how LLM-based systems fail.

  • Contribute to applied science on agentic memory: use statistical methods grounded in audit domain knowledge to surface which experiences are worth remembering, and define and validate the promotion and demotion criteria that let agents compound that knowledge over time, partnering with platform engineering on the infrastructure that executes it.

  • Advance self-learning, self-improving systems that get better both offline and online as they are used.

  • Help ensure memory promotion and demotion criteria uphold the legal and contractual obligations owed to customers and their clients, partnering with Security, Legal, and Domain SMEs.

  • Design experiments and evaluations that demonstrate agent quality improves the more customers use the platform.

  • Translate advances in GenAI (models, retrieval, agent frameworks, evaluation techniques) into practical, maintainable capabilities.

  • At the staff level, influence technical direction through RFCs and design reviews, and mentor other scientists and engineers.

What you'll bring:
  • 6+ years (senior) to 8+ years (staff) of professional experience in applied science, machine learning, research, or data-intensive engineering. Staff-level candidates bring demonstrated impact beyond a single team.

  • 2+ years working on production GenAI or LLM-based systems (applications, agents, or the tooling and evaluation around them).

  • Proven experience designing and building evaluation frameworks for ML or LLM systems: metrics, benchmarks, scoring approaches, and rigorous experiment design.

  • A strong experimentation mindset. You form clear hypotheses, design sound experiments, and draw defensible conclusions from noisy, real-world data.

  • Ability to build and operate your own production tooling and services, ideally on AWS.

  • Strong understanding of GenAI system tradeoffs including quality, latency, cost, reliability, and safety.

  • A strong foundation in probability and statistics: experimental design, significance testing, and reasoning under uncertainty.

  • Strong English language communication and collaboration skills.

  • Comfortable operating in fast-moving environments with ambiguity and evolving requirements.

Nice to have
  • PhD or MS in a quantitative field (Computer Science, Statistics, Machine Learning, or similar) preferred, though equivalent industry experience is welcome.

  • Prior machine learning experience (classical ML, model training, or MLOps).

  • Experience with AI guardrails, governance, or safety mechanisms.

  • Experience with distributed, SaaS, cloud-native, multi-tenant platforms at scale.

  • Familiarity with Infrastructure as Code (CDK, CloudFormation, or Terraform).

  • Experience operating in regulated or compliance-heavy domains.

  • Familiarity with financial audit, accounting, or assurance workflows.

Tech stack you'll be working with
  • Backend & Platform: TypeScript, NestJS, Python

  • Cloud & Infrastructure: AWS EKS, AWS Lambda, AWS Bedrock, AWS AgentCore

  • Search & Retrieval: AWS OpenSearch and S3 Vectors

  • Document & Data Processing: AWS Textract, DynamoDB, S3

  • AI Evaluation & Observability: LangFuse, LangSmith, LangChain, LangGraph

  • AI-Assisted Development: GitHub Copilot, Claude Code, Devin

  • Developer Tooling: GitHub, GitHub Actions, Nx Monorepo

Perks & Benefits
  • Contrato a termino Indefinido with all the legal benefits
  • Prepaid Medicine
  • Life insurance and funeral assistance
  • Internet allowance
  • Home office stipend
  • Competitive compensation — above the market average
  • 100% remote work environment and an excellent work-life balance
  • 5 Personal Time Off days per year
  • Sick Leave Top up to total 100% of salary paid by the employer from Day 3 to 90.
  • Recognition Award, additional paid time off in recognition of the corresponding year of service
  • Upgrade vacation starting at 5 years of service
  • Opportunity to work for a growing global SaaS leader company
  • A culture that promotes independence, innovation, trust, and accountability
  • Open space to be creative, innovative and strategize for the future
  • Mentorship by highly experienced professional
  • Budget for training, we want you to grow
  • AI-first environment: Be part of an AI-first engineering organization that embraces modern tools, automation, and AI-driven ways of working.

What's in it for you:

Innovation is at our core. We work with cutting-edge technology in accounting and financial reporting, constantly pushing the boundaries to create impactful software solutions.

We are committed to a collaborative culture, where your ideas are valued, and knowledge sharing is encouraged within a supportive, inclusive team.

Work-life balance is important to us. We offer flexible work options, remote opportunities, and generous time-off policies to ensure a healthy work-life balance.

We offer competitive compensation, including a competitive salary and comprehensive benefits

We are driven by impactful work. Your contributions directly affect how our clients manage financial processes and drive their success.

Recognition and rewards matter to us. We celebrate hard work through recognition programs, performance bonuses, and opportunities for career growth.

We embrace global opportunities. Work on international projects and collaborate with a diverse, global team.

About Caseware: Caseware's cutting-edge software products are meticulously designed for accounting firms, corporations, and governments. Our teams are continually collaborating, innovating, and building upon our existing suite of products. With a customer-focused mindset, we are building technology that is shaping what the future of audits, financial reporting, and financial data analytics will look like.

With a recent strategic investment from Hg Capital in 2020, Caseware is now in its next major growth phase as we double down on the people and products that have made Caseware so successful to date.

One of Caseware's core values is Many Voices, One Team and with that in mind, we’re dedicated to building teams as diverse as our customers in an equitable and inclusive way. We welcome and encourage candidates of all backgrounds to apply. Should you require accommodations or have any questions at any point during the application or interview process, please e-mail our People Operations team at talent@caseware.com.

Background Check: Any candidates successful in obtaining an offer for a position will need to successfully complete a background check through Certn.co which typically includes an Identity Verification and Criminal Record Check. Executives and Senior Managers will undergo a Soft Credit Check as well. Candidates residing in the Netherlands and Germany are excluded from undergoing background checks via Certn.co

Security and Fraud: Caseware takes the security of candidates seriously. All legitimate communication from us will come from email addresses ending in @caseware.com and our open positions are always listed on reputable job boards and on our website https://jobs.lever.co/caseware. We will NEVER ask for payment or financial information from you. If you receive an unsolicited job offer, proceed with extreme caution.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Applied Scientist
Senior Applied Scientist

Caseware • Medellín

Presencial
COP 285.307.000 - 443.810.000
Contrato a termino Indefinido
Prepaid Medicine
Life insurance and funeral assistance
+5
Senior Software Developer III
Senior Software Developer III

Caseware International • Bogotá

Presencial
COP 120.000.000 - 180.000.000
100% remote work
Competitive compensation
Training budget
+1
Fullstack Senior Software Developer
Fullstack Senior Software Developer

Caseware International • Bogotá

Híbrido
COP 133.920.000 - 223.200.000
Prepaid Medicine
Life insurance and funeral assistance
Internet allowance
+4
Java Staff Software Developer IV
Java Staff Software Developer IV

Caseware International • Bogotá

Híbrido
COP 90.000.000 - 170.000.000
Home office stipend
Internet allowance
Competitive compensation—above market
Staff Software Developer IV
Staff Software Developer IV

Caseware • Bogotá

Presencial
COP 180.000.000 - 320.000.000
Remote work
Home office stipend
Training budget
+2
SDET II
SDET II

Caseware • Medellín

Presencial
COP 9.000.000 - 15.000.000
Prepaid Medicine
Life insurance and funeral assistance
Internet allowance
+2
Principal Software Engineer
Principal Software Engineer

Caseware • Colombia

A distancia
COP 221.934.000 - 295.913.000
Flexible work options
Generous time-off policies
Comprehensive health insurance
+2
Software Develop in Test
Software Develop in Test

Caseware • Bogotá

Presencial
COP 54.000.000 - 90.000.000
100% remote work
Home office stipend
Internet allowance
+3
Software Develop in Test
Software Develop in Test

Caseware International • Bogotá

Presencial
COP 60.000.000 - 120.000.000
Prepaid Medicine
Life insurance & funeral assistance
Internet allowance
+2
Java Senior Software Developer
Java Senior Software Developer

Caseware International • Bogotá

Híbrido
COP 283.921.000 - 441.654.000
Permanent contract
Prepaid Medicine
Life insurance and funeral assistance
+3