Job Description: Principal AI Engineer — Delivery Measurement & Evaluation
Company: Logic Hire
Role Type: Full-Time
Required: 10+ years
Location: Remote / Hybrid (US) & Should be US Citizen or Green card only - H1B and other visa's will be rejected.
Reports To: Engineering Leadership & Finance
About The Role
Logic Hire is applying AI across the software delivery lifecycle — not only to writing code, but to testing, review, documentation, migration, incident response, and the validation stages where delivery is usually constrained. Measuring what that produces is the starting project, not the whole of it.
You will own the analytical and applied half of that work: designing the experiments that establish what actually helps, building the classification and evaluation pipelines the measurement program depends on, and then taking the findings back into the delivery lifecycle as capabilities teams can use.
This role pairs with a data engineer who owns extraction, identity, and the metric pipeline. You own what the numbers mean and what to do about them.
Key Responsibilities
- Experimental Design & Causal Inference: Design and run controlled experiments across the software delivery lifecycle. Model tier routing, MCP coverage, permission configuration, repository context quality, and budget headroom — randomized across teams and reported with their limits stated. These are the cleanly causal questions available once a tool is deployed, and where the returns are.
- Select and justify the appropriate experimental methodology for each question — randomized controlled trials, quasi-experimental designs, difference-in-differences, instrumental variables, regression discontinuity, or hierarchical models — and state clearly when a design does not support the claim being asked of it.
- Own the analytical layer of the measurement program: work classification over model traffic, evaluation design, longitudinal within-unit analysis, and the staggered-adoption estimates that connect delivery outcomes to adoption timing.
- Define the counterfactual. For every claimed improvement, establish what would have happened without the intervention — through holdout groups, matched controls, pre/post within-unit comparison, or adoption-timing natural experiments — and publish the assumptions that make the estimate valid.
- Model uncertainty explicitly. Report confidence intervals, not point estimates. Distinguish statistical significance from practical significance. State what the experiment was and was not powered to detect.
- Design for heterogeneity. Segment results by team, role, repository, language, service criticality, and tenure with the tool. Surface where an intervention helps one group and harms another rather than reporting a single blended average.
- Evaluation & Classification Pipeline Development
- Build and validate LLM-as-judge pipelines — sampling strategy, hand-labeled ground truth, precision and recall measured and published, and revalidation whenever the taxonomy or the model changes.
- Design and maintain classification systems over model traffic: what work is being done, at what stage of the lifecycle, by whom, with what tool, and to what outcome. Own the taxonomy, its versioning, and its drift.
- Build evaluations for internal AI capabilities: golden sets, regression suites, groundedness and answer-quality scoring, and the cost and latency telemetry beside them.
- Establish ground truth rigorously. Define labeling protocols, measure inter-rater agreement, resolve disagreements, and maintain a labeled corpus that survives scrutiny from engineers who will dispute the results.
- Instrument for evaluation from the start. Work with platform and data engineering to ensure the events, context, and metadata required for evaluation are captured at source rather than reconstructed after the fact.
- Revalidate continuously. Any change to the model, prompt, taxonomy, or tool surface invalidates prior evaluation results. Own the revalidation schedule and publish when results are stale.
- AI Capability Extension Across the Delivery Lifecycle
- Extend AI beyond code authoring into the stages that constrain delivery: test authoring and maintenance, environment and data setup, migration and modernization, code review assistance, security remediation, and evidence assembly for certification.
- Work directly with the constrained teams. Where validation or certification is the bottleneck, coding assistance produces little regardless of how well it works — find where the constraint actually sits and aim the capability at it.
- Identify the real constraint. Apply queueing and flow analysis — utilization, batch economics, work-in-progress limits — to locate where delivery actually slows down, rather than assuming the bottleneck is where the tooling is easiest to deploy.
- Prototype, measure, and either scale or kill. Build the minimum viable capability, evaluate it against the established baseline, and make an explicit recommendation to scale, iterate, or stop.
- Translate findings into deployable capabilities. Work with platform and product engineering to turn a validated prototype into a supported capability teams can adopt without your involvement.
- Turning Findings into Practice
- Identify what the most effective practitioners do differently. Study the top decile of users, document their workflows, prompts, context strategies, and tool configurations, and publish practices rather than rankings.
- Teach and enable. Run working sessions, write internal documentation, and build onboarding material so effective practices propagate without requiring your direct involvement.
- Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry, so what you learn becomes the default rather than folklore.
- Close the loop between measurement and configuration. When an experiment identifies a better configuration, drive it into the platform defaults rather than leaving it as a recommendation.
- Publish practices, not scorecards. Avoid building artifacts that rank individuals. The output of this work is better delivery, not performance surveillance.
- Data Collection, Aggregation & Instrumentation
- Instrument and extract from operational systems and APIs. Design sampling strategies that survive scrutiny, and know when a data source cannot answer the question being asked of it.
- Resolve identity across systems. Join sources never designed to be joined — GitLab, Jira, Confluence, gateway telemetry, model registries, CI/CD — and model the summary tables that reporting reads from.
- Own the analytical data model. Define what a "unit of work" is, how it is attributed, and how it flows from raw events to summary tables. You need not own the pipeline, but you must be able to build one when the answer depends on it.
- Perform exploratory analysis with distributions rather than averages, cohort and time-series work, and reports that state their own coverage and limits.
- Guard against measurement artifacts. Recognize when a change in tooling, taxonomy, team composition, or reporting cadence is producing a change in the numbers rather than a change in reality.
- Reporting, Governance & Communication
- Report to engineering leadership and finance on what is working, what is not, where delivery is actually constrained, and what each claim does and does not establish.
- State the limits of every claim. Publish coverage, sample sizes, known confounders, and the conditions under which a finding would not hold.
- Communicate in both directions — to an executive audience that wants a number, and to engineers who will dispute it. Hold the line on rigor without becoming an obstacle to delivery.
- Exercise care with personnel-adjacent data. Aggregate reporting by default. Maintain a clear sense of what should not be built even when it is technically easy, and elevate when asked to build it anyway.
- Report to finance in financial terms. Connect AI investment to delivery cost, cycle time, rework, and capacity — with the uncertainty stated — so budget decisions rest on evidence rather than enthusiasm.
RequirementsExperience
- 10+ years spanning software engineering and quantitative analysis. This role needs both; a strong background in one and a passing acquaintance with the other will not carry it.
- Production experience with LLM applications: prompting, tool and function calling, context management, evaluation, and knowing where models fail in practice.
- Experimental design and causal inference — randomized and quasi-experimental designs, difference-in-differences, instrumental variables, hierarchical models — and the judgment to say when a design does not support the claim being asked of it.
Technical Skills
- Strong Python and SQL, with a statistical stack (pandas, statsmodels, scikit-learn, or R).
- Data collection: instrumenting and extracting from operational systems and APIs, designing sampling that survives scrutiny, and knowing when a source cannot answer the question being asked of it.
- Aggregation: resolving identity across systems, joining sources never designed to be joined, and modeling the summary tables reporting reads from. You need not own the pipeline, but you must be able to build one when the answer depends on it.
- Analytics: exploratory analysis, distributions rather than averages, cohort and time-series work, and reports that state their own coverage and limits.
Domain Knowledge
- Real familiarity with the software delivery lifecycle — code review, CI/CD, test strategy, release and change management — sufficient to hold a credible conversation with the teams you are measuring.
- Git and GitLab at instrumentation depth: merge request and pipeline data models, diffs and SHAs, what merge, squash, rebase, and cherry-pick do to line-level analysis, and the API and hook surfaces available for capturing it.
- Jira and Confluence integration experience — REST APIs, changelog and page version history, the GitLab–Jira development panel, and the field and label conventions that determine whether the resulting data means anything.
- Care with personnel-adjacent data: aggregate reporting by default, and a clear sense of what should not be built even when it is technically easy.
Communication
- Communication that works in both directions — an executive audience that wants a number, and engineers who will dispute it.
Tech StackCategoryTechnologies
Programming Languages Python, SQL, R (optional)
Statistical & ML Libraries pandas, statsmodels, scikit-learn, NumPy, SciPy
LLM & AI Frameworks LangGraph, LangChain, Bedrock Agents, Strands, OpenAI API, Anthropic Claude API
Evaluation & Observability Ragas, DeepEval, Bedrock Model Evaluation, LangFuse, Arize, OpenTelemetry-based tracing
Data Pipeline & Orchestration dbt, Airflow, Dagster, warehouse/lakehouse modeling
Version Control & DevOps Git, GitLab (merge requests, pipelines, diffs, SHAs, hooks), GitLab CI, server-side Git hooks
Project & Knowledge Management Jira (REST APIs, changelog, development panel), Confluence (REST APIs, page version history)
Cloud & Infrastructure Amazon Bedrock, AWS cost and usage data, self-managed GitLab instances
BI & Visualization BI tooling (Tableau, Looker, or equivalent), summary table modeling
Engineering Frameworks DORA, DX Core 4, SPACE
Connector Frameworks MCP servers and clients, webhook-driven capture
Preferred Qualifications
- MCP servers and clients, or comparable connector frameworks.
- Agent frameworks — LangGraph, LangChain, Bedrock Agents, Strands, or equivalent.
- Enterprise deployment of coding assistants, and their telemetry.
- Server-side Git hooks, GitLab CI, and system or webhook-driven capture on a self-managed instance.
- Confluence and Jira as MCP-connected systems — permission propagation, scoped credentials, and audit logging.
- Evaluation tooling — Ragas, DeepEval, Bedrock model evaluation — and LLM observability such as LangFuse, Arize, or OpenTelemetry-based tracing.
- Amazon Bedrock, and AWS cost and usage data.
- Engineering productivity frameworks — DORA, DX Core 4, SPACE — and a working view of their limits.
- Program analysis, test generation, or developer tooling research.
- dbt, Airflow, Dagster, or equivalent transformation and orchestration; warehouse or lakehouse modeling.
- BI and visualization tooling, and the discipline of building on summary tables rather than raw events.
- Queueing and flow analysis: utilization, batch economics, and constraint identification.
The Opportunity
Most organizations deploying AI to engineering cannot say what they got for it and respond by buying more of it or by arguing. You will build the evidence instead and then use it to decide where the next capability goes.
The measurement program is the first project. The scope is applied AI across the delivery lifecycle, with real internal users, a platform team to build on, and leadership that will act on what you find.
Skills: gitlab,dora,design,python,jira,llm