Senior AI Agent Engineer

EPAM Systems Inc

United States

Remote

USD 150,000 - 210,000

Full time

44 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

EPAM Systems Inc. in the United States seeks an experienced engineer to design and deliver AI agent systems that remain reliable beyond demonstrations.

You will own technical decisions and delivery for a defined system or substantial component, from requirements and design through production operation. The role blends strong software engineering with practical agent development, requires independent learning of new languages and runtimes, and involves mentoring others while communicating

Qualifications

  • Substantial production software engineering experience, five+ years.
  • At least one year delivering LLM-based agents to production.
  • Ability to explain contributions and trade-offs with operational results.
  • Hands-on agent development using SDKs, frameworks, or direct model APIs.
  • Ability to learn unfamiliar languages and runtimes and validate with tests.

Responsibilities

  • Design, build, deploy, and operate agent systems across execution, memory, identity, tool integration, and observability.
  • Collaborate with customers to turn needs into functional requirements and acceptance criteria.
  • Build prototypes and production-grade slices from design through deployment.
  • Develop agent workflows using cloud managed services for AI hosting (Bedrock, Azure AI Foundry, Vertex AI).
  • Implement an LLM gateway and routing layer with model tiering and observability.
  • Design memory and context-management for multi-step workflows and isolation between users.

Skills

Production software engineering
LLM-based agents
Cloud AI hosting
Mentoring engineers
Strong programming
Testing and diagnostics
Observability
English proficiency

Job description

Design and deliver AI agent systems that remain reliable beyond a successful demonstration. In this role, you will own technical decisions and delivery for a defined system or substantial component, from requirements and design through production operation. You will combine strong software engineering with practical agent development, make trade-offs explicit, and help other engineers deliver high-quality work. We welcome engineers from different programming backgrounds. We value depth in your existing language and the ability to learn the languages and tools needed for the role.

Responsibilities
  • Design, build, deploy, and operate agent systems across execution, memory, identity, tool integration, and observability within your area of responsibility
  • Work directly with customers and stakeholders to turn needs into functional requirements and measurable acceptance criteria.
  • Personally build working prototypes and production‑grade slices of the solution, from design through implementation, testing, and deployment
  • Develop agent workflows and production deployments using cloud managed services for AI hosting (e.g., Bedrock, Azure AI Foundry, Vertex AI), selecting services and development tools against the system's requirements
  • Implement and operate an LLM gateway & routing layer covering model tiering, quota/cost control, and observability (e.g., LiteLLM, APIM AI Gateway)
  • Design memory and context‑management approaches for multi‑step workflows, including state persistence, retrieval quality, retention, and isolation between users or engagements
  • Build and maintain MCP servers and tool integrations for authorized security workflows, including reconnaissance, scanners, controlled exploit tooling, and internal services.
  • Define contracts, permissions, approval boundaries, and recovery behavior
  • Establish automated tests and agent evaluations for your components.
  • Distinguish model‑quality issues from software defects, and cover tool failures, interrupted execution, and unintended repeated actions
  • Diagnose production issues using logs, metrics, and traces.
  • Improve reliability, latency, and cost, and verify that mitigation and recovery work as intended
  • Contribute to CI/CD, release checks, rollback plans, code review, and technical documentation.
  • Mentor less experienced engineers and communicate decisions clearly to technical and non‑technical stakeholders
Requirements
  • Substantial production software engineering experience, typically five or more years, including at least one year delivering LLM‑based agents to production.
  • You can explain your contribution, the trade‑offs you made, and the operational results
  • Strong programming skills and production experience in at least one general‑purpose language, together with hands‑on agent development using SDKs, frameworks, or direct model APIs.
  • You can independently learn an unfamiliar language or runtime and validate your implementation through tests and diagnostics
  • Depth in at least one technical area and the ability to evaluate language, runtime, and framework trade‑offs against the system's requirements, including concurrency, performance, maintainability, and integration needs
  • Experience designing systems or substantial components, evaluating alternatives, and independently making technical decisions within an agreed scope
  • Practical experience with MCP, tool use, prompt and context engineering, memory, and agent evaluation, including testing behavior across multiple steps and failure conditions
  • Practical experience with cloud managed services for AI hosting (e.g., Bedrock, Azure AI Foundry, Vertex AI), including runtime, state, identity, integration, and observability concerns, and the ability to adopt an unfamiliar platform through working implementations
  • Production engineering skills in API design, asynchronous processing, persistence, authentication and authorization, observability, retries, and security boundaries.
  • You can investigate issues that span multiple components
  • Experience with automated testing, CI/CD, code review, and AI‑assisted development, together with a disciplined approach to validating generated output.
  • You can clarify requirements, communicate designs, and support the growth of other engineers
  • Proficiency in English at a B2+ level
  • Nice to have Experience in application security, web penetration testing, or red‑teaming within an authorized scope
  • Experience with sandboxed code execution, browser agents, or autonomous tool orchestration
  • Experience improving evaluation coverage, deployment safety, or operating costs for agent systems
  • Experience conducting technical interviews or contributing to engineering standards and knowledge sharing
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead AI Agent Engineer
Lead AI Agent Engineer

EPAM Systems Inc • United States

Remote
USD 180,000 - 240,000
Technical Architect - ML
Technical Architect - ML

Quantiphi, Inc. • Princeton (NJ)

On-site
USD 140,000 - 190,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

Perplexity AI • New York (NY)

On-site
USD 150,000 - 210,000
Senior Forward Deployed Engineer
Senior Forward Deployed Engineer

Visvero, Inc. • Piscataway Township (NJ)

On-site
USD 150,000 - 190,000
Software Engineer, AI Platform
Software Engineer, AI Platform

Triwill Group • San Francisco (CA)

On-site
USD 140,000 - 180,000
Senior AI Agentic Lead - On-Site & Automation Architect
Senior AI Agentic Lead - On-Site & Automation Architect

Vidorra Consulting Group • Mountain View (CA)

On-site
USD 120,000 - 150,000
Principal AI Engineer
Principal AI Engineer

Heitmeyer Consulting • Columbus (OH)

On-site
USD 180,000 - 230,000
AI Engineer
AI Engineer

Talentify • Town of Texas (WI)

Hybrid
USD 120,000 - 190,000
Hybrid Work Options
Award-Winning Culture
Generous PTO
+7
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 190,000 - 230,000
AI Engineer
AI Engineer

Heitmeyer Consulting • Columbus (OH)

On-site
USD 140,000 - 190,000