Generative AI Engineer- Inferencing

Photon

United States

On-site

USD 140,000 - 170,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Photon is seeking a Software Engineer - Gen AI Inferencing to design, build, and operate reusable toolkits for Gen AI RAG capabilities. The role emphasizes scalable AI inference and integration with data science workflows across multi-repo platforms.

You will collaborate with product teams, data scientists, and engineers to deploy models in containers, drive CI/CD, and advance automation. 5+ years of Python/Unix experience and GenAI tooling are preferred.

Qualifications

  • 5+ years of OOP in Python/Scala/Java
  • Experience with AI/ML/GenAI lifecycle management
  • Hands-on with MLOps, fine-tuning and inference
  • Deploying models in containers in production
  • CI/CD activities and automation

Responsibilities

  • Code solutions and unit tests per acceptance criteria
  • Design and modify architecture components and interfaces
  • Mentor engineers in CI/CD practices and tool automation
  • Refine stories, requirements, and estimate work
  • Perform spikes and PoC to mitigate risk
  • Automate release activities
  • Develop automated test suites (integration/performance)
  • Collaborate with product, data analysts, and data scientists

Skills

Python development
Scala/Java
GenAI/ML lifecycle
MLOps / CI-CD
Containerization
FastAPI / API
CI/CD automation
NoSQL (MongoDB/Redis)
Git tooling

Tools

Triton Inference Server
vLLM
Jenkins
SonarQube
pytest
Artifactory
Ansible
VS Code

Job description

For the past 20 years, we have powered many Digital Experiences for the Fortune 500. Since 1999, we have grown from a few people to more than 4000 team members across the globe that are engaged in various Digital Modernization. Our current focus and innovation in Digital Hyper expansion TM offers nearly limitless opportunities for career growth. For a brief 1-minute video about us, you can check out https://youtu.be/uJWBWQZEA6o.

Photon is an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We bring out the best in each other.

Software Engineer - Gen AI Inferencing

Required Qualifications:

  • 5+ years OOP in Python/Scala/Java programming experience with expert-level development skills.
  • Experience with AI/ML/GenAI Lifecycle Management and Development and its ecosystem.
  • Hands-on experience building frameworks using MLOps, Fine-Tuning techniques, and Inference Frameworks.
  • Experience with deploying models using vLLM/Triton Inference Server in containers in production with automation.
  • Performs Continuous Integration and Continuous Development (CI-CD) activities.
  • Performance tuning those models and deployment to provide higher throughput.
  • Track record of maintaining large-scale Python/Unix-based systems.
  • Hands-on experience and knowledge of generative AI RAG processes for various use cases, including chunking, embedding, retrieval, reranking, and summarization.
  • Hands-on experience in application development in one or more areas: MongoDB, Redis, Angular/React Frameworks, Containerization, Building API-based applications leveraging FastAPI services, JWT Integration, API Gateway.
  • Develop efficient utilities, automation frameworks, and data science platforms that can be utilized across multiple Data Science teams for AI/ML and GenAI work.
  • Working in large-sized teams that collaboratively develop on a shared multi-repo codebase using IDEs (e.g., VS Code rather than Jupyter Notebooks), Continuous Integration (CI), Continuous Deployment (CD), and Continuous Testing.
  • Strong automation, scripting, and Python development skills. Hands-on DevOps experience with one or more of the following enterprise development tools: Version Control (GIT/Bitbucket), Build Orchestration (Jenkins), Code Quality (SonarQube and pytest Unit Testing), Artifact Management (Artifactory), and Deployment (Ansible).

Position Summary

  • This position is focused on the design, build, and operation of reusable toolkits for Gen AI RAG capabilities.
  • This job is responsible for developing and delivering complex requirements to accomplish business goals. Key responsibilities of the job include ensuring that software is developed to meet functional, non-functional, and compliance requirements, and that solutions are well designed with maintainability/ease of integration and testing built in from the outset. Job expectations include a strong knowledge of development and testing practices common to the industry, as well as design and architectural patterns.

Responsibilities:

  • Codes solutions and unit tests to deliver a requirement/story per the defined acceptance criteria and compliance requirements.
  • Designs, develops, and modifies architecture components, application interfaces, and solution enablers while ensuring principal architecture integrity is maintained.
  • Mentors other software engineers and coaches the team on Continuous Integration and Continuous Development (CI-CD) practices and automating the tool stack.
  • Executes story refinement, definition of requirements, and estimating work necessary to realize a story through the delivery lifecycle.
  • Performs spikes/proof of concept as necessary to mitigate risk or implement new ideas.
  • Automates manual release activities.
  • Designs, develops, and maintains automated test suites (integration, regression, performance).
  • Utilizes multiple architectural components (across data, application, business) in the design and development of client requirements.
  • Manages multiple priorities and simultaneously engages with multiple teams.
  • Participates in estimating work necessary to realize a story/requirement through the delivery lifecycle.
  • Is vocal and actively participates in all sessions with business stakeholders and agile teams.
  • Collaborates with product teams, data analysts, and data scientists to design and build solutions.

Desired Qualifications

  • Experience building & deploying Gen AI inferencing platforms with open-source toolsets, building inferencing & servicing capabilities (AI Gateway, Policy Store, Observability) for RAG/MCP use cases, etc.
  • Hands-on experience driving and maintaining a culture of quality, innovation, and experimentation.
  • Research on new tools and capabilities for better UI and UX for advanced analytics platforms, quick prototyping and demonstration of features and capabilities, and participation in various user forums.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Generative AI Engineer- AI/ RAG
Generative AI Engineer- AI/ RAG

Photon • Dallas (TX)

On-site
USD 140,000 - 190,000
Gen AI Inferencing Engineer
Gen AI Inferencing Engineer

Photon • United States

On-site
USD 150,000 - 230,000
Software Engineer - AI/RAG
Software Engineer - AI/RAG

Photon • Dallas (TX)

On-site
USD 120,000 - 180,000
Gen AI Platform Engineer — RAG & MLOps
Gen AI Platform Engineer — RAG & MLOps

Photon • United States

On-site
USD 140,000 - 170,000
AI Engineer
AI Engineer

Kaleidoscope Innovation • Fort Worth (TX)

On-site
USD 140,000 - 190,000
AI & RAG Software Engineer (GenAI, Data Pipelines)
AI & RAG Software Engineer (GenAI, Data Pipelines)

Photon • Dallas (TX)

On-site
USD 120,000 - 180,000
Gen AI Data Engineer
Gen AI Data Engineer

Tiger Analytics • United States

Hybrid
USD 120,000 - 150,000
Significant career development opportunities
Work in a fast-growing entrepreneurial environment
Java Gen AI Engineer
Java Gen AI Engineer

Photon • Dallas (TX)

On-site
USD 140,000 - 180,000
Manager of Artificial Intelligence
Manager of Artificial Intelligence

Photon • Dallas (TX)

On-site
USD 150,000 - 180,000
Gen AI Java Lead
Gen AI Java Lead

Photon • Dallas (TX)

On-site
USD 140,000 - 190,000