Machine Learning Operations / Language Model Operations Engineer
As part of our continued expansion in AI-driven platforms and data-intensive systems, we are looking for a Machine Learning Operations / Language Model Operations Engineer to join our advanced engineering team. In this role, you will be responsible for designing, building, and operating the end-to-end infrastructure that supports production AI systems, ensuring reliability, observability, data quality, and continuous improvement of machine learning and LLM-based solutions.
Key Responsibilities
- Design and implement scalable, source-agnostic ingestion pipelines for production ML and LLM data.
- Build and maintain data processing workflows including ingestion, redaction, storage, classification, and slicing of production signals.
- Define and implement data storage strategies, including tiering, retention policies, and privacy-aware data handling.
- Develop observability and debugging tools, including dashboards and query systems for production AI systems.
- Implement evaluation and monitoring frameworks, including offline evaluation sets, online autoraters, and regression detection systems.
- Build automated triage systems to identify, classify, and surface production failures.
- Develop and maintain PII redaction mechanisms and enforce data governance and compliance policies at ingestion level.
- Design and operate LLM evaluation mining workflows to continuously improve model and prompt performance.
- Implement alerting systems to detect regressions across model and prompt deployments.
- Collaborate with AI, engineering, and product teams to ensure robustness and reliability of production AI systems.
- Evaluate and select tooling, infrastructure, and hosting strategies for ML/LLM platforms.
- Own the operational reliability of the entire ML/LLM data and evaluation pipeline.
Requirements
- Proven experience building and operating production‑grade data platforms or ML systems, including ingestion, storage, access control, monitoring, and on‑call responsibilities.
- Hands‑on experience developing ML/LLM evaluation systems (e.g., regression test sets, autoraters, LLM‑as‑a‑judge frameworks, or golden datasets).
- Strong understanding of LLM observability, tracing, and debugging tools.
- Experience implementing data privacy controls such as PII redaction in production environments.
- Deep understanding of failure modes in ML and LLM systems (hallucinations, retrieval failures, agent loops, ASR/TTS degradation, prompt/model regressions).
- Strong production‑level Python engineering skills, with a hands‑on mindset.
- Solid understanding of data pipelines, system reliability, and distributed data processing.
Nice-to-Have Skills
- Experience in multi‑tenant or SaaS architectures with strict data isolation requirements.
- Familiarity with Azure and/or AWS cloud ecosystems.
- Experience making infrastructure trade‑off decisions between managed services and self‑hosted solutions.
- Knowledge of vector databases, embedding techniques, and clustering or unsupervised failure detection methods.
- Experience with data versioning tools such as LakeFS, DVC, or Delta Lake.
- Familiarity with GDPR, data deletion workflows, and compliance‑driven data systems.
- Exposure to embedded, automotive, or constrained environments.
- Experience with non‑English language model evaluation or multilingual datasets.
- Experience working with LLM APIs such as Anthropic Claude, OpenAI models, or open‑source alternatives.
- Familiarity with CI/CD workflows, GitHub‑based development, and modern DevOps practices.
- Experience with dashboard development using TypeScript or similar frontend technologies.
Equal Opportunity Statement
At Tieto, we welcome applicants of all backgrounds, genders (m/f/d), and walks of life. We are committed to diversity, equity, and inclusion.