Advantest is seeking a highly experienced Senior Principal AI Systems Engineer to architect and deliver advanced production AI capabilities. This role is for a hands‑on technical leader with deep expertise in agentic AI, open‑source model development, model fine‑tuning, evaluation systems, and reliable AI infrastructure. The ideal candidate combines advanced AI knowledge with strong software‑engineering discipline and has experience building systems that reason across multiple steps, use tools, evaluate results, recover from failures, and improve from feedback.
Key Responsibilities
Agentic AI Systems
- Architect production‑grade systems for tool‑using and multi‑step AI agents.
- Design orchestration, planning, memory, session management, retries, timeouts, observability, and failure recovery.
- Establish structured and typed interfaces between models, tools, data services, and applications.
- Build controls that prevent agents from bypassing required validation and approval stages.
- Develop reusable agent frameworks that support multiple products and deployment environments.
Model Development and Fine‑Tuning
- Fine‑tune open‑source language models and specialized models for complex technical applications.
- Apply LoRA, QLoRA, full fine‑tuning, distillation, preference optimization, and related post‑training methods.
- Develop efficient strategies for adapting foundation models to new applications and datasets.
- Evaluate tradeoffs among model quality, inference performance, deployment cost, security, and maintainability.
- Build optimized models and inference profiles for constrained deployment environments.
Evaluation and Quality
- Define measurable standards for model accuracy, reliability, safety, and production readiness.
- Build offline evaluation datasets, automated regression suites, judge systems, and promotion gates.
- Develop methods for measuring confidence, consistency, tool‑use accuracy, and action quality.
- Establish processes for model comparison, controlled release, rollback, and continuous improvement.
- Ensure model outputs remain grounded in available evidence and approved data sources.
Training Data and Learning Pipelines
- Design reproducible pipelines for training‑data generation, cleaning, labeling, versioning, and validation.
- Develop synthetic‑data and preference‑data strategies where appropriate.
- Implement controls for data leakage, contamination, duplication, provenance, and customer isolation.
- Convert expert feedback and observed outcomes into high‑quality training and evaluation datasets.
- Maintain traceability between datasets, experiments, model versions, and production results.
AI Platform and MLOps
- Build repeatable training, evaluation, and deployment workflows for multi‑GPU infrastructure.
- Establish model lifecycle practices, including model cards, release criteria, monitoring, and rollback.
- Support secure, private, air‑gapped, and customer‑controlled deployment environments.
- Develop observability for model behavior, tool execution, latency, cost, and failure conditions.
- Partner with platform and CI/CD teams to make AI workflows repeatable, testable, and auditable.
Technical Leadership
- Set technical direction for a small, highly skilled AI engineering team.
- Review architectures, models, training methods, and production implementation decisions.
- Mentor engineers and establish durable AI engineering practices.
- Work with domain experts to translate complex technical requirements into reliable AI capabilities.
- Communicate technical risks, tradeoffs, progress, and recommendations to engineering and executive stakeholders.