Senior / Staff Software AI Test Engineer, AI Engineering

TWG AI

Santa Monica (CA)

On-site

USD 190,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Full range of medical benefits
Bonus compensation

Job summary

TWG AI, located in Santa Monica, CA, is searching for a Senior or Staff AI Software Engineer in Test to build commercial-grade AI products. This role will focus on designing and building test automation frameworks and infrastructure, ensuring quality in AI-powered applications.

The ideal candidate will have 3-7 years of experience, with a strong command of Python, Java, and automation tools. A competitive salary of $190,000-$250,000 along with benefits will be provided.

Qualifications

  • 3-7 years of software engineering experience, focused on test automation.
  • Expert-level in Python, writing libraries and applying OOP practices.
  • Experience testing across multiple client surfaces like iOS apps and Chrome extensions.

Responsibilities

  • Design and build scalable test automation frameworks for AI agents.
  • Build evaluation infrastructure for benchmarking agent performance.
  • Integrate automated tests into CI/CD for validation before shipping.

Skills

Python
Test automation
Java
Data manipulation
Test data preparation

Education

Bachelor’s degree or higher in Computer Science, Engineering, or a related field

Tools

pytest
Selenium
GitHub Actions

Job description

At TWG Group Holdings, LLC ("TWG Global"), we drive innovation and business transformation across a range of industries—financial services, insurance, technology, media, and sports—by leveraging data and AI as core assets. Our AI‑first, cloud‑native approach delivers real‑time intelligence and interactive business applications, empowering informed decision‑making for both customers and employees. We prioritize responsible data and AI practices, ensuring ethical standards and regulatory compliance. Our decentralized structure enables each business unit to operate autonomously, supported by a central AI Solutions Group, while strategic partnerships with leading data and AI vendors fuel game‑changing efforts in marketing, operations, and product development. You will collaborate with management to advance our data and analytics transformation, enhance productivity, and enable agile, data‑driven decisions. By leveraging relationships with top tech startups and universities, you will help create competitive advantages and drive enterprise innovation. At TWG Global, your contributions will support our goal of sustained growth and superior returns, as we deliver rare value and impact across our businesses.

The Role

TWG Global is seeking a Senior or Staff AI Software Engineer in Test to join our AI Engineering team building commercial‑grade AI products. This is a software engineering role focused on test automation. You won’t just write test cases, you’ll design and build the frameworks, harnesses, evaluation infrastructure, and tooling that make testing AI agents and LLM‑powered applications possible at scale. Our agents are written in LangGraph and run on Azure on the TWG side, with a parallel Vercel‑based stack on the Palantir side. You’ll write eval sets against both, and you’ll validate the surfaces our users actually touch: iOS apps, plugins, and Chrome extensions, not just the model layer. You’ll work shoulder‑to‑shoulder with AI engineers and data scientists, contributing production‑quality code to shared repositories.

Key Responsibilities
Framework and harness engineering
  • Design and build scalable, reusable test automation frameworks for AI agents, LLM‑powered applications, and underlying APIs.
  • Write clean, maintainable Python for test harnesses, eval pipelines, synthetic data generation utilities, and internal tooling.
  • Treat test code as production code: code review, type hints, documentation, library design.
Evaluation infrastructure
  • Build evaluation infrastructure for benchmarking agent performance against SOTA LLMs, competitors, and internal baselines.
  • Own regression suites, golden datasets, rubric‑based evals, and metric dashboards.
  • Build tooling for synthetic test data generation, edge‑case discovery, and adversarial testing.
Resilience and load
  • Design and run release, system, performance, and load tests against streaming, stateful, and async systems.
  • Build chaos and fault injection tooling for token expiry, connection pool exhaustion, provider failover, and cache pressure scenarios.
  • Drive contract testing across LLM providers (Bedrock, Anthropic, OpenAI) to catch parity drift.
CI/CD and observability
  • Integrate automated tests into CI/CD so every model, prompt, and code change is validated before it ships.
  • Build trace‑based assertions on LangGraph state, tool calls, and agent decisions—debugging an agent failure means replaying graph state, not re‑running a prompt.
  • Make observability a first‑class testing surface (LangSmith, audit logs).
Human‑in‑the‑loop and partnership
  • Implement HIL review workflows where automation alone cannot validate quality, then push the automation boundary outward.
  • Partner with AI engineers and data scientists on model evaluation, training and eval data prep, and root‑cause debugging of complex end‑to‑end failures.
  • Champion quality engineering practices across the team: code review, coverage standards, observability, reproducibility.
  • Ensure user‑centric validation so AI outputs are accurate, reliable, and meet real‑world application needs.
Requirements
  • 3‑7 years of software engineering experience, with a meaningful portion focused on test automation, SDET, or software engineering in test roles.
  • Expert‑level Python. You write Python every day, design libraries other engineers use, and apply OOP and clean‑code practices.
  • Hands‑on Java experience, enough to read, write, and test Java services, not just touch them.
  • Working understanding of the LangGraph or Vercel frameworks: graph state, nodes, edges, tool calls, and how to write evals against agentic flows.
  • Demonstrated experience building eval sets for LLM models (this is critical to the role).
  • Experience testing across multiple client surfaces: iOS apps, plugins, and Chrome extensions.
  • Hands‑on experience building automated test suites with frameworks such as pytest, Selenium, Playwright, Cypress, or similar.
  • Proven experience integrating test automation into CI/CD systems (GitHub Actions, Jenkins, CircleCI, GitLab CI, or similar).
  • Strong skills in data manipulation, test data preparation, and SQL.
  • Bachelor’s degree or higher in Computer Science, Engineering, or a related field.
Strongly preferred
  • Experience with Azure (our primary cloud) and containerization (Docker).
  • Experience testing RAG pipelines, agentic workflows, or multi‑step tool‑calling systems.
Benefits
Position Location

This position is located in Santa Monica, CA (on‑site).

Compensation

The base pay for this position is $190,000‑250,000. A bonus will be provided as part of the compensation package, in addition to a full range of medical, financial, and/or other benefits.

TWG is an equal opportunity employer, and all qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior / Staff Software AI Test Engineer, AI Engineering
Senior / Staff Software AI Test Engineer, AI Engineering

TWG AI • New York (NY)

On-site
USD 190,000 - 250,000
Medical benefits
Financial benefits
Bonuses
Senior / Staff Software AI Test Engineer, AI Engineering
Senior / Staff Software AI Test Engineer, AI Engineering

TWG Global AI • New York (NY)

On-site
USD 150,000 - 250,000
Senior Engineer, AI Engineering
Senior Engineer, AI Engineering

TWG AI • New York (NY)

Hybrid
USD 190,000 - 200,000
Medical benefits
Financial benefits
Relocation support
Senior Engineer, AI Engineering
Senior Engineer, AI Engineering

TWG AI • Santa Monica (CA)

Hybrid
USD 190,000 - 200,000
Medical benefits
Financial benefits
Relocation support
Senior Full-Stack Engineer, AI Engineering
Senior Full-Stack Engineer, AI Engineering

TWG AI • Santa Monica (CA)

On-site
USD 190,000 - 200,000
Real-world AI projects
Top-tier collaboration
Flat, autonomous culture
+1
Senior Full-Stack Engineer, AI Engineering
Senior Full-Stack Engineer, AI Engineering

Socket.dev • New York (NY)

On-site
USD 190,000 - 200,000
Competitive salary
Performance-based incentives
Full range of medical benefits
Senior Full-Stack Engineer, AI Engineering
Senior Full-Stack Engineer, AI Engineering

TWG Global AI • Santa Monica (CA)

On-site
USD 190,000 - 200,000
Work on real-world AI applications
Collaborate with world-class data sci/
Senior AI Test Engineer for LLM-powered Agents
Senior AI Test Engineer for LLM-powered Agents

TWG AI • Santa Monica (CA)

On-site
USD 190,000 - 250,000
Full range of medical benefits
Bonus compensation
Senior AI Test Engineer — On-site Santa Monica
Senior AI Test Engineer — On-site Santa Monica

TWG Global AI • New York (NY)

On-site
USD 150,000 - 250,000
Staff / VP, Data Science - Marketing & Sales Focus
Staff / VP, Data Science - Marketing & Sales Focus

TWG AI • New York (NY)

Hybrid
USD 280,000 - 300,000
Medical benefits
Performance bonus
Hybrid NY office