Senior Software Engineer – Applied AI

Jobtailor

Deutschland

Remote

EUR 120.000 - 180.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, versandbereit.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor is seeking a senior AI/ML engineer in Germany to design and run experiments across routing, coverage, and budget headroom, while owning the analytical layer of our measurement program.

You will build and validate LLM-based judgment pipelines, extend AI capabilities across testing, deployment, and data workflows, and report findings to leadership to guide product delivery.

Qualifikationen

  • 7+ years in software engineering and quantitative analysis.
  • Production experience with LLM applications: prompting, tool calling, context management.
  • Experimental design and causal inference, including randomized and quasi-experimental designs, difference-in-differences, instrumental variables, and hierarchical models.
  • Strong Python and SQL, with a statistical stack such as pandas, statsmodels, scikit-learn, or R.
  • Data collection from operational systems and APIs; robust sampling design.
  • Identity resolution across systems and complex data joins.

Aufgaben

  • Design and run randomized experiments on routing, MCP coverage, permission configuration, repository context quality, and budget headroom.
  • Own the analytical layer of the measurement program, including work classification, evaluation design, longitudinal within-unit analysis, and staggered-adoption estimates.
  • Build and validate LLM-as-judge and classification pipelines with sampling, hand-labeled ground truth, precision and recall measurement, and revalidation.
  • Extend AI capabilities across testing, environment and data setup, migration and modernization, code review, security remediation, and certification evidence assembly.
  • Work directly with constrained teams to identify delivery bottlenecks and target capabilities accordingly.
  • Build evaluations for internal AI capabilities, including golden sets, regression suites, groundedness, answer-quality scoring, cost, and latency telemetry.
  • Identify effective practitioner behaviors, document and teach practices, and publish practices rather than rankings.
  • Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry.
  • Report findings to engineering leadership and finance, including what is working, delivery constraints, and the limits of each claim.

Kenntnisse

LLM Applications
Experimental Design
Causal Inference
Python
SQL
Data Collection
Git and GitLab
Executive Communication

Tools

Git
GitLab
Jira
Confluence
Amazon Bedrock
AWS
Ragas
DeepEval
LangChain
Airflow

Jobbeschreibung

  • Design and run randomized experiments on tier routing, MCP coverage, permission configuration, repository context quality, and budget headroom
  • Own the analytical layer of the measurement program, including work classification, evaluation design, longitudinal within-unit analysis, and staggered-adoption estimates
  • Build and validate LLM-as-judge and classification pipelines with sampling, hand-labeled ground truth, precision and recall measurement, and revalidation
  • Extend AI capabilities across testing, environment and data setup, migration and modernization, code review, security remediation, and certification evidence assembly
  • Work directly with constrained teams to identify delivery bottlenecks and target capabilities accordingly
  • Build evaluations for internal AI capabilities, including golden sets, regression suites, groundedness, answer-quality scoring, cost, and latency telemetry
  • Identify effective practitioner behaviors, document and teach practices, and publish practices rather than rankings
  • Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry
  • Report findings to engineering leadership and finance, including what is working, delivery constraints, and the limits of each claim
Requirements
  • 7+ years spanning software engineering and quantitative analysis
  • Production experience with LLM applications: prompting, tool and function calling, context management, evaluation, and knowing where models fail in practice
  • Experimental design and causal inference, including randomized and quasi-experimental designs, difference-in-differences, instrumental variables, and hierarchical models
  • Strong Python and SQL, with a statistical stack such as pandas, statsmodels, scikit-learn, or R
  • Data collection from operational systems and APIs; robust sampling design
  • Identity resolution across systems and complex data joins
  • Exploratory analysis, distributions, cohort and time-series analysis, and reporting of coverage and limitations
  • Familiarity with the software delivery lifecycle, including code review, CI/CD, test strategy, release and change management
  • Git and GitLab at instrumentation depth, including merge request and pipeline data models, diffs and SHAs, merge, squash, rebase, cherry-pick, APIs, and hooks
  • Jira and Confluence integration experience, including REST APIs, changelogs, page version history, GitLab–Jira development panel, fields, and labels
  • Care with personnel-adjacent data and aggregate reporting
  • Communication with executive and engineering audiences
  • Preferred MCP servers and clients or comparable connector frameworks
  • Agent frameworks such as LangGraph, LangChain, Bedrock Agents, Strands, or equivalent
  • Enterprise deployment of coding assistants and telemetry
  • Server-side Git hooks, GitLab CI, and system or webhook-driven capture on a self-managed instance
  • Confluence and Jira as MCP-connected systems, including permission propagation, scoped credentials, and audit logging
  • Evaluation tooling such as Ragas, DeepEval, or Bedrock model evaluation; LLM observability such as LangFuse, Arize, or OpenTelemetry-based tracing
  • Amazon Bedrock and AWS cost and usage data
  • Engineering productivity frameworks such as DORA, DX Core 4, or SPACE
  • Program analysis, test generation, or developer tooling research
  • dbt, Airflow, Dagster, or equivalent transformation and orchestration; warehouse or lakehouse modeling
  • BI and visualization tooling; summary-table-based reporting
  • Queueing and flow analysis, including utilization, batch economics, and constraint identification
Core Competencies

Demonstrates expertise in experimental design and causal inference, with a strong focus on LLM applications and quantitative analysis. Proficient in Python and SQL, with experience in data collection, exploratory analysis, and software delivery lifecycle.

Highest-signal resume keywords
  • LLM Application Development
  • Experimental Design and Causal Inference
  • Python and SQL Proficiency
  • Data Collection and Sampling Design
  • Git and GitLab Expertise
Hard Skills
  • Experimental Design
  • Causal Inference
  • Python
  • SQL
  • Data Collection
  • Exploratory Analysis
  • LLM Applications
  • Statistical Analysis
  • Sampling Design
  • Evaluation Tooling
Soft Skills
  • Communication with Executive Audiences
  • Documentation and Teaching Practices
Industry Keywords
  • Software Engineering
  • Quantitative Analysis
  • AI Capabilities
  • Software Delivery Lifecycle
  • Engineering Productivity Frameworks
Tools & Technologies
  • Git
  • GitLab
  • Jira
  • Confluence
  • Amazon Bedrock
  • AWS
  • Ragas
  • DeepEval
  • LangChain
  • Airflow
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Research Engineer
AI Research Engineer

Jobtailor • Berlin

Vor Ort
EUR 90.000 - 130.000
Junior Data Scientist
Junior Data Scientist

Jobtailor • Hamburg

Vor Ort
EUR 70.000 - 100.000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Jobtailor • Deutschland

Remote
EUR 120.000 - 170.000
AI Engineering Strategist
AI Engineering Strategist

Jobtailor • Deutschland

Remote
EUR 70.000 - 100.000
Senior Software Test Engineer
Senior Software Test Engineer

Jobtailor • Deutschland

Remote
EUR 90.000 - 120.000
AI Engagement Lead
AI Engagement Lead

Jobtailor • Deutschland

Remote
EUR 110.000 - 160.000
Senior GenAI Engineer
Senior GenAI Engineer

Jobtailor • Deutschland

Remote
EUR 90.000 - 150.000
Prompt Engineer
Prompt Engineer

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Manager, Machine Learning Engineer
Manager, Machine Learning Engineer

Jobtailor • Deutschland

Remote
EUR 120.000 - 180.000
AI Engineer – GenAI Platform, Mid-Level
AI Engineer – GenAI Platform, Mid-Level

Jobtailor • Deutschland

Remote
EUR 90.000 - 130.000