Principal AI Engineer

Logic Hire Solutions LTD

United States

Hybrid

USD 180,000 - 260,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Logic Hire is seeking a Principal AI Engineer to lead delivery measurement and evaluation across the software lifecycle. You will design experiments, build classification and evaluation pipelines, and translate findings into actionable capabilities for delivery teams.

The role requires 10+ years in software and analytics, strong expertise in causal inference, experimental design, and experience with LLM-based evaluation, model monitoring, and data pipelines.

Qualifications

  • 10+ years in software engineering and quantitative analysis.
  • Production experience with LLM applications and evaluation.
  • Expertise in experimental design and causal inference.

Responsibilities

  • Design experiments across the software delivery lifecycle and define the causal questions.
  • Own analytical measurement, model traffic classification, and evaluation design.
  • Develop and validate LLM-as-judge pipelines and taxonomy drift handling.
  • Translate findings into deployable capabilities for delivery teams.
  • Instrument for evaluation from the start with robust ground-truth processes.
  • Publish rigorous results with clear uncertainty and limits.

Skills

Software engineering
Quantitative analysis
Experimental design
Causal inference
Python
SQL

Tools

Git
GitLab
Jira
Confluence
dbt
Airflow
Dagster
Tableau
Looker
LangGraph
LangChain
Bedrock
Arize
OpenTelemetry

Job description

Job Description: Principal AI Engineer — Delivery Measurement & Evaluation

Company: Logic Hire

Role Type: Full-Time

Required: 10+ years

Location: Remote / Hybrid (US) & Should be US Citizen or Green card only - H1B and other visa's will be rejected.

Reports To: Engineering Leadership & Finance

About The Role

Logic Hire is applying AI across the software delivery lifecycle — not only to writing code, but to testing, review, documentation, migration, incident response, and the validation stages where delivery is usually constrained. Measuring what that produces is the starting project, not the whole of it.

You will own the analytical and applied half of that work: designing the experiments that establish what actually helps, building the classification and evaluation pipelines the measurement program depends on, and then taking the findings back into the delivery lifecycle as capabilities teams can use.

This role pairs with a data engineer who owns extraction, identity, and the metric pipeline. You own what the numbers mean and what to do about them.

Key Responsibilities
  • Experimental Design & Causal Inference: Design and run controlled experiments across the software delivery lifecycle. Model tier routing, MCP coverage, permission configuration, repository context quality, and budget headroom — randomized across teams and reported with their limits stated. These are the cleanly causal questions available once a tool is deployed, and where the returns are.
  • Select and justify the appropriate experimental methodology for each question — randomized controlled trials, quasi-experimental designs, difference-in-differences, instrumental variables, regression discontinuity, or hierarchical models — and state clearly when a design does not support the claim being asked of it.
  • Own the analytical layer of the measurement program: work classification over model traffic, evaluation design, longitudinal within-unit analysis, and the staggered-adoption estimates that connect delivery outcomes to adoption timing.
  • Define the counterfactual. For every claimed improvement, establish what would have happened without the intervention — through holdout groups, matched controls, pre/post within-unit comparison, or adoption-timing natural experiments — and publish the assumptions that make the estimate valid.
  • Model uncertainty explicitly. Report confidence intervals, not point estimates. Distinguish statistical significance from practical significance. State what the experiment was and was not powered to detect.
  • Design for heterogeneity. Segment results by team, role, repository, language, service criticality, and tenure with the tool. Surface where an intervention helps one group and harms another rather than reporting a single blended average.
  • Evaluation & Classification Pipeline Development
  • Build and validate LLM-as-judge pipelines — sampling strategy, hand-labeled ground truth, precision and recall measured and published, and revalidation whenever the taxonomy or the model changes.
  • Design and maintain classification systems over model traffic: what work is being done, at what stage of the lifecycle, by whom, with what tool, and to what outcome. Own the taxonomy, its versioning, and its drift.
  • Build evaluations for internal AI capabilities: golden sets, regression suites, groundedness and answer-quality scoring, and the cost and latency telemetry beside them.
  • Establish ground truth rigorously. Define labeling protocols, measure inter-rater agreement, resolve disagreements, and maintain a labeled corpus that survives scrutiny from engineers who will dispute the results.
  • Instrument for evaluation from the start. Work with platform and data engineering to ensure the events, context, and metadata required for evaluation are captured at source rather than reconstructed after the fact.
  • Revalidate continuously. Any change to the model, prompt, taxonomy, or tool surface invalidates prior evaluation results. Own the revalidation schedule and publish when results are stale.
  • AI Capability Extension Across the Delivery Lifecycle
  • Extend AI beyond code authoring into the stages that constrain delivery: test authoring and maintenance, environment and data setup, migration and modernization, code review assistance, security remediation, and evidence assembly for certification.
  • Work directly with the constrained teams. Where validation or certification is the bottleneck, coding assistance produces little regardless of how well it works — find where the constraint actually sits and aim the capability at it.
  • Identify the real constraint. Apply queueing and flow analysis — utilization, batch economics, work-in-progress limits — to locate where delivery actually slows down, rather than assuming the bottleneck is where the tooling is easiest to deploy.
  • Prototype, measure, and either scale or kill. Build the minimum viable capability, evaluate it against the established baseline, and make an explicit recommendation to scale, iterate, or stop.
  • Translate findings into deployable capabilities. Work with platform and product engineering to turn a validated prototype into a supported capability teams can adopt without your involvement.
  • Turning Findings into Practice
  • Identify what the most effective practitioners do differently. Study the top decile of users, document their workflows, prompts, context strategies, and tool configurations, and publish practices rather than rankings.
  • Teach and enable. Run working sessions, write internal documentation, and build onboarding material so effective practices propagate without requiring your direct involvement.
  • Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry, so what you learn becomes the default rather than folklore.
  • Close the loop between measurement and configuration. When an experiment identifies a better configuration, drive it into the platform defaults rather than leaving it as a recommendation.
  • Publish practices, not scorecards. Avoid building artifacts that rank individuals. The output of this work is better delivery, not performance surveillance.
  • Data Collection, Aggregation & Instrumentation
  • Instrument and extract from operational systems and APIs. Design sampling strategies that survive scrutiny, and know when a data source cannot answer the question being asked of it.
  • Resolve identity across systems. Join sources never designed to be joined — GitLab, Jira, Confluence, gateway telemetry, model registries, CI/CD — and model the summary tables that reporting reads from.
  • Own the analytical data model. Define what a "unit of work" is, how it is attributed, and how it flows from raw events to summary tables. You need not own the pipeline, but you must be able to build one when the answer depends on it.
  • Perform exploratory analysis with distributions rather than averages, cohort and time-series work, and reports that state their own coverage and limits.
  • Guard against measurement artifacts. Recognize when a change in tooling, taxonomy, team composition, or reporting cadence is producing a change in the numbers rather than a change in reality.
  • Reporting, Governance & Communication
  • Report to engineering leadership and finance on what is working, what is not, where delivery is actually constrained, and what each claim does and does not establish.
  • State the limits of every claim. Publish coverage, sample sizes, known confounders, and the conditions under which a finding would not hold.
  • Communicate in both directions — to an executive audience that wants a number, and to engineers who will dispute it. Hold the line on rigor without becoming an obstacle to delivery.
  • Exercise care with personnel-adjacent data. Aggregate reporting by default. Maintain a clear sense of what should not be built even when it is technically easy, and elevate when asked to build it anyway.
  • Report to finance in financial terms. Connect AI investment to delivery cost, cycle time, rework, and capacity — with the uncertainty stated — so budget decisions rest on evidence rather than enthusiasm.
RequirementsExperience
  • 10+ years spanning software engineering and quantitative analysis. This role needs both; a strong background in one and a passing acquaintance with the other will not carry it.
  • Production experience with LLM applications: prompting, tool and function calling, context management, evaluation, and knowing where models fail in practice.
  • Experimental design and causal inference — randomized and quasi-experimental designs, difference-in-differences, instrumental variables, hierarchical models — and the judgment to say when a design does not support the claim being asked of it.
Technical Skills
  • Strong Python and SQL, with a statistical stack (pandas, statsmodels, scikit-learn, or R).
  • Data collection: instrumenting and extracting from operational systems and APIs, designing sampling that survives scrutiny, and knowing when a source cannot answer the question being asked of it.
  • Aggregation: resolving identity across systems, joining sources never designed to be joined, and modeling the summary tables reporting reads from. You need not own the pipeline, but you must be able to build one when the answer depends on it.
  • Analytics: exploratory analysis, distributions rather than averages, cohort and time-series work, and reports that state their own coverage and limits.
Domain Knowledge
  • Real familiarity with the software delivery lifecycle — code review, CI/CD, test strategy, release and change management — sufficient to hold a credible conversation with the teams you are measuring.
  • Git and GitLab at instrumentation depth: merge request and pipeline data models, diffs and SHAs, what merge, squash, rebase, and cherry-pick do to line-level analysis, and the API and hook surfaces available for capturing it.
  • Jira and Confluence integration experience — REST APIs, changelog and page version history, the GitLab–Jira development panel, and the field and label conventions that determine whether the resulting data means anything.
  • Care with personnel-adjacent data: aggregate reporting by default, and a clear sense of what should not be built even when it is technically easy.
Communication
  • Communication that works in both directions — an executive audience that wants a number, and engineers who will dispute it.
Tech StackCategoryTechnologies

Programming Languages Python, SQL, R (optional)

Statistical & ML Libraries pandas, statsmodels, scikit-learn, NumPy, SciPy

LLM & AI Frameworks LangGraph, LangChain, Bedrock Agents, Strands, OpenAI API, Anthropic Claude API

Evaluation & Observability Ragas, DeepEval, Bedrock Model Evaluation, LangFuse, Arize, OpenTelemetry-based tracing

Data Pipeline & Orchestration dbt, Airflow, Dagster, warehouse/lakehouse modeling

Version Control & DevOps Git, GitLab (merge requests, pipelines, diffs, SHAs, hooks), GitLab CI, server-side Git hooks

Project & Knowledge Management Jira (REST APIs, changelog, development panel), Confluence (REST APIs, page version history)

Cloud & Infrastructure Amazon Bedrock, AWS cost and usage data, self-managed GitLab instances

BI & Visualization BI tooling (Tableau, Looker, or equivalent), summary table modeling

Engineering Frameworks DORA, DX Core 4, SPACE

Connector Frameworks MCP servers and clients, webhook-driven capture

Preferred Qualifications
  • MCP servers and clients, or comparable connector frameworks.
  • Agent frameworks — LangGraph, LangChain, Bedrock Agents, Strands, or equivalent.
  • Enterprise deployment of coding assistants, and their telemetry.
  • Server-side Git hooks, GitLab CI, and system or webhook-driven capture on a self-managed instance.
  • Confluence and Jira as MCP-connected systems — permission propagation, scoped credentials, and audit logging.
  • Evaluation tooling — Ragas, DeepEval, Bedrock model evaluation — and LLM observability such as LangFuse, Arize, or OpenTelemetry-based tracing.
  • Amazon Bedrock, and AWS cost and usage data.
  • Engineering productivity frameworks — DORA, DX Core 4, SPACE — and a working view of their limits.
  • Program analysis, test generation, or developer tooling research.
  • dbt, Airflow, Dagster, or equivalent transformation and orchestration; warehouse or lakehouse modeling.
  • BI and visualization tooling, and the discipline of building on summary tables rather than raw events.
  • Queueing and flow analysis: utilization, batch economics, and constraint identification.
The Opportunity

Most organizations deploying AI to engineering cannot say what they got for it and respond by buying more of it or by arguing. You will build the evidence instead and then use it to decide where the next capability goes.

The measurement program is the first project. The scope is applied AI across the delivery lifecycle, with real internal users, a platform team to build on, and leadership that will act on what you find.

Skills: gitlab,dora,design,python,jira,llm

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied AI Engineer - Developer Experience
Applied AI Engineer - Developer Experience

Spectrum IT Recruitment • California (MO)

On-site
USD 150,000 - 210,000
Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

On-site
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
Full stack with Python AI/ML | Dallas TX
Full stack with Python AI/ML | Dallas TX

Programmers.io • Dallas (TX)

Hybrid
USD 150,000 - 190,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Cassi Home • United States

On-site
USD 180,000 - 260,000
AI Analyst (UA/RU Language speaking)
AI Analyst (UA/RU Language speaking)

Neurons Lab LTD. • Town of Poland (NY)

On-site
USD 120,000 - 155,000
Software Engineer, AI Platform
Software Engineer, AI Platform

Triwill Group • San Francisco (CA)

Hybrid
USD 140,000 - 180,000
AI Engineer
AI Engineer

Pinpoint Global Communications • United States

On-site
USD 120,000 - 180,000
AI Engineer
AI Engineer

Teserac, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Health Care Plan
Paid Time Off
Free Food
+2
Full Stack Developer / AI Focused
Full Stack Developer / AI Focused

Worky • Long Beach (CA)

On-site
USD 140,000 - 210,000
Staff Software Engineer, AI Systems
Staff Software Engineer, AI Systems

Dolby • Atlanta (GA)

On-site
USD 150,000 - 210,000