Role Summary
We are looking for an experienced engineer to design and develop agentic AI systems that automate complex HPC and enterprise workflows. In this role, you will build production‑grade agents capable of planning, invoking tools, managing workflow state, launching jobs, diagnosing failures, optimizing execution, and operating safely across shared compute and enterprise environments.
The successful candidate will have experience delivering reliable LLM‑powered agentic systems, a strong understanding of infrastructure and developer workflows, and the ability to bridge agentic AI, distributed systems, HPC scheduling, enterprise automation, observability, and performance optimization.
Key Responsibilities
- Design agents for HPC and enterprise workflows, incorporating planning, tool use, memory, retrieval, governance, and human approval mechanisms.
- Build production systems for job orchestration, enterprise workflow automation, monitoring, debugging, optimization, reproducibility, and rollback.
- Automate environments, containers, toolchains, schedulers, CI/CD, enterprise systems, telemetry, workflow composition, and launch processes.
- Implement safe execution through authentication, authorization, secrets handling, audit logging, sandboxing, compliance controls, and policy enforcement.
- Research and prototype agentic methods for long‑running workflows, failure recovery, optimization loops, evaluation, and human‑in‑the‑loop reliability.
Collaboration
- Partner with Platform teams to deploy solutions safely on shared clusters, enterprise infrastructure, and governed execution environments.
- Work with Performance teams to define benchmarks, variance controls, profiling methods, productivity metrics, and acceptance criteria.
- Collaborate with Compiler and Runtime teams to expose code generation, tuning, tracing, debugging, and execution controls.
- Engage workload owners, enterprise application teams, security, and compliance stakeholders to onboard use cases and ensure auditable operation.
What You'll Bring
- Advanced degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent practical experience.
- Experience delivering production agentic AI or LLM systems with orchestration, tool use, memory, evaluations, and long‑horizon reliability.
- Strong proficiency in Python and experience with at least one systems programming language such as C, C++, Rust, or Go.
- Experience integrating automation with real‑world tooling, including code execution, build/test systems, schedulers, enterprise systems, telemetry, CI/CD, or deployment pipelines.
- Experience building reliable production systems with observability, fallback mechanisms, regression gates, incident debugging, security controls, and operational ownership.
Additional Experience That May Be Helpful
- Experience with HPC workflows, including Slurm, Kubernetes, multi‑node GPU execution, containers, distributed launch, and shared clusters.
- Experience automating enterprise workflows, developer platforms, IT operations, business systems, approvals, audit trails, or governed tool execution.
- Experience with GPU profiling and performance analysis tools and trace‑based workflows.
- Experience with kernel authoring, tuning, or code generation using Triton, CUDA, HIP, MLIR, LLVM, or XLA‑like flows.
- Experience with LLM training, supervised fine‑tuning, reinforcement learning, evaluation workflows, inference serving, model optimization, or agent evaluation.
Location
Finland or Sweden. Remote work arrangements may be considered for candidates located outside the Helsinki or Stockholm metropolitan areas. Advanced degree in Computer Science, Computer Engineering, Electrical Engineering, or related technical field, or equivalent practical experience, Experience delivering production agentic AI or LLM systems with orchestration, tool use, memory, evaluations, and long‑horizon reliability, Strong proficiency in Python, Experience with at least one systems programming language (C, C++, Rust, or Go), Experience integrating automation with real‑world tooling (code execution, build/test systems, schedulers, enterprise systems, telemetry, CI/CD, or deployment pipelines), Experience building reliable production systems with observability, fallback mechanisms, regression gates, incident debugging, and security controls