Technical Lead

SmartDev

Województwo kujawsko-pomorskie

Hybrid

PLN 420,000 - 640,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SmartDev, an AI-powered software development company, is seeking a founding technical lead for Veris EvalOps. You will architect and ship two commercial modules—a Release Gate and a Knowledge Health Monitor—while hiring and leading 5–7 engineers across three streams, with the first pilot client live by month 3.

You’ll own the evaluation engine, tracing layer, and knowledge health pipeline, making architecture calls in a zero-to-one build without a principal above you, and you’ll ensure the

Qualifications

  • 5+ years in engineering with leadership experience in a complete build cycle.
  • Experience with LLM evaluation methods, regression testing, and hallucination detection.
  • Strong grasp of LLM observability, tracing, and cost/latency attribution.
  • Proven experience with RAG systems, embeddings, and vector stores.
  • Production-grade backend engineering with async Python (FastAPI, Celery) and multi-tenant SaaS.
  • Ability to balance build-versus-integrate decisions and work with pilot clients.

Responsibilities

  • Own the evaluation engine: LLM-as-judge scoring, rule checks, groundedness, and regression tracking.
  • Develop tracing and observability across LLM calls, RAG, and agent workflows.
  • Lead knowledge health pipeline: ingest and analyze enterprise knowledge sources; detect stale content and coverage gaps.
  • Own platform core and integrations: multi-tenant RBAC, API connectors, dashboards, and CI/CD hooks.
  • Build and lead a 5–7 engineer team across streams: Platform Core, Release Gate, Knowledge Health.

Skills

Engineering leadership
LLM evaluation
LLM observability
RAG architecture
AI agent systems
Backend Python
FastAPI
PostgreSQL/Redis

Tools

OpenTelemetry
Pinecone
Weaviate
Qdrant
pgvector
Langfuse
Braintrust
Celery

Job description

SmartDev is an AI-powered software development company headquartered in Vietnam, part of the Verysell Group. We help global businesses deliver faster and build smarter — combining AI-driven development practices with deepexpertiseacrossfintech, healthcare, retail, and enterprise technology. Our team of engineers, architects, and AI specialistsworksacross the full stack: from custom software and cloud solutions to generative AI,MLOps, and intelligent automation. At SmartDev, we believe that great architecture and AI-first thinking are a competitive advantage — and we build our teams accordingly.

The company is at a pivotal point in our journey: transitioning from a pure IT outsourcing provider to anAI-enabled solutions partner.

We have over 200 talented employees working both on-site and remote/hybrid at locations:

  • Danang: 81 Quang Trung, Hai Chau 1 ward, Hai Chau District, Danang
  • Hanoi: Zodiac building, 19 Duy Tan Street, Dich Vong Hau ward, Cau Giay District, Hanoi

You’ll be the founding technical lead for Veris EvalOps, building the platform that answers the two questions every AI-deploying business needs answered: is this system safe to launch, and is it still working correctly a month later. You’ll take it from first line of code to first paying clients in 6–7 months.

Role Summary

This is a zero-to-one build, not a maintenance role. You’ll architect and ship two commercial modules — a pre-production Release Gate that turns “looks good” into a reproducible readiness score, and a Knowledge Health Monitor that continuously audits the knowledge base an AI draws from — while hiring and leading the engineers who build them alongside you. There's no principal architect above you to escalation: you make the calls and live with them, with the first pilot client live by month 3.

Key Responsibilities
  • Own the evaluation engine. LLM-as-judge scoring, rule-based checks, groundedness verification, hallucination detection, and regression comparison — every readiness score comes from here.
  • Build the tracing and observability layer. Distributed tracing across LLM calls, RAG retrievals, and agent workflows, built on OpenTelemetry, capturing every token, tool call, cost, and latency metric.
  • Ship the knowledge health pipeline. Ingestion and continuous analysis of enterprise knowledge sources — stale-content detection, contradiction analysis, and coverage-gap mapping.
  • Own platform core and integrations. Multi-tenant architecture, RBAC, API connectors, dashboards, and the CI/CD hooks that let the Release Gate plug into client engineering workflows.
  • Build and lead the team. Hire and run 5–7 engineers across three streams — Platform Core, Release Gate, Knowledge Health — and own every architecture decision end to end.
Must-have
  • 5+ years in engineering. Including 2+ years leading a team of 3–8 through a complete build cycle — architecture to shipping to paying users. Not a first-time lead role.
  • LLM evaluation methodology. LLM-as-judge design, RAGAS/DeepEval-style metrics, golden dataset construction, regression testing for AI systems, and hallucination detection — not just the library calls, the mechanics behind them.
  • LLM observability and tracing. OpenTelemetry-based tracing across LLM calls, RAG retrievals, and multi-step agent trajectories; cost/latency attribution; drift and anomaly detection.
  • RAG system architecture. Production experience across the full pipeline — chunking, embeddings, a vector store (Pinecone, Weaviate, Qdrant, or pgvector), retrieval, re-ranking.
  • AI agent systems. Production experience with agent patterns (ReAct, Plan-and-Execute, supervisor/sub-agent), tool-call evaluation, and guardrails.
  • Backend platform engineering. Production-grade async Python (FastAPI, Celery), multi-tenant SaaS architecture, PostgreSQL/Redis, and CI/CD integration.
  • Build-vs-integrate judgment and client-facing comfort. Can weigh integrating Langfuse/Braintrust vs. building from scratch, and work directly with pilot clients during onboarding and results review.
Nice-to-have
  • LLM APIs and model ecosystem. Multi-provider experience (OpenAI, Anthropic, Azure OpenAI, Bedrock) and routing/prompt-management at scale.
  • MLOps and experiment tracking. Background with MLflow, Weights & Biases, or equivalent experiment-tracking tooling.
  • Security, compliance, and AI governance. EU AI Act and NIST AI RMF awareness, PII handling in AI pipelines, and red-teaming basics — increasingly a qualification question in enterprise security reviews.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Lead
Technical Lead

SmartDev • Warszawa, Wierzchy

On-site
PLN 240,000 - 360,000
Technical Lead
Technical Lead

SmartDev • gmina Osie

On-site
PLN 240,000 - 420,000
AI Architect (Voice AI)
AI Architect (Voice AI)

Neurons Lab • Poland

On-site
PLN 260,000 - 420,000
Forward Deployed Engineer
Forward Deployed Engineer

Tenarai Europe • Poland

On-site
PLN 120,000 - 180,000
Onsite parking
Life Insurance
Development budget
+5
Lead AI Engineer & Technical Architect (Remote)
Lead AI Engineer & Technical Architect (Remote)

Bold Business • Warszawa

On-site
PLN 80,000 - 110,000
Lead AI Engineer & Technical Architect (Remote)
Lead AI Engineer & Technical Architect (Remote)

Bold Business • Województwo małopolskie

On-site
PLN 254,000 - 382,000
AI Architect
AI Architect

Luxoft • Poland

On-site
PLN 240,000 - 360,000
Private Medical & Dental care
Life Insurance
Internal Mobility program
AI Architect
AI Architect

Devapo • Województwo mazowieckie

On-site
PLN 50,000 - 70,000
Certifications and training funded
Private medical care (Medicover)
Multisport card
+4
AI Architect / Tech Lead (mahjong game)
AI Architect / Tech Lead (mahjong game)

Neurons Lab • Poland

On-site
PLN 201,000 - 357,000
Software Engineer - AI Platform
Software Engineer - AI Platform

CaptivateIQ, Inc. • Poland

Hybrid
PLN 212,000 - 298,000