Prometa.ai is a rapidly growing technology company that develops deep learning–based recommendation systems and agentic AI solutions powered by Large Language Models (LLMs). We operate active projects across various industries and continue to strengthen our pioneering position in the field by offering AI-powered products and consultancy services tailored to sectors such as finance, retail, healthcare, insurance, and telecommunications. At Prometa.ai, we design next-generation AI systems by combining the strengths of LLMs, RAG pipelines, and autonomous agent workflows.
We are currently looking for an AI Architect who can lead the design and deployment of scalable, production-ready AI infrastructures—someone who not only understands the latest research but can also architect robust solutions that bridge experimentation and enterprise-level deployment!
KEY RESPONSIBILITIES
- Solution Discovery & Design: Collaborate with customer business and IT teams to translate requirements into solution designs with clearly defined inputs, outputs, success metrics, and failure scenarios.
- Agentic AI Architecture: Design end-to-end LLM solutions using RAG, tool-calling, and multi-agent workflows, including guardrails, fallback and timeout strategies, and human-in-the-loop checkpoints.
- On-Premise AI Architecture: Recommend model selection and serving architectures for private and closed-network environments, considering vLLM, quantization, GPU and memory capacity, performance, and cost.
- POC-to-Production Delivery: Rapidly build and code proof-of-concept solutions, then define the technical roadmap, effort estimate, and architecture required for production deployment.
- Enterprise Integration & Data Infrastructure: Integrate AI solutions with legacy, cloud-native, and enterprise systems while designing data ingestion, retrieval, reranking, and vector or graph database architectures.
- Evaluation & Observability: Establish evaluation and monitoring frameworks using golden datasets, LLM-as-a-judge, relevancy, completeness and safety metrics, regression testing, and observability tools.
- Security & Compliance: Embed PII protection, prompt injection defence, access control, logging, and auditability into solution architectures in line with KVKK, BDDK, GDPR, MASAK, and PCI-DSS requirements.
- Scalable AI Deployments: Oversee containerized and GPU-optimized deployments using Docker, Kubernetes, or OpenShift, ensuring performance, fault tolerance, and cost efficiency.
- Product Architecture: Contribute to Prometa’s Orkestra product architecture, including telemetry schemas, SDK integration patterns, scoring pipelines, and monitoring capabilities.
- Customer Engagement & PreSales: Lead customer meetings, workshops, demos, and technical presentations; prepare architecture documents, solution proposals, and RFP responses while supporting sales teams.
- Collaboration & Mentorship: Work closely with business, engineering, data, security, and compliance teams while guiding engineers through architectural standards, code reviews, and technical mentorship.
WHAT WE ARE LOOKING FOR
- Bachelor’s degree in Computer Science, Software Engineering, Electrical and Electronics Engineering, Mathematics, Artificial Intelligence, or a related field.
- 6+ years of experience in software, data, or machine learning, including at least 2 years working with LLM or ML systems in production.
- Proven experience designing and deploying at least one end-to-end GenAI system—such as a RAG solution, AI assistant, or agent that serves real production traffic.
- Advanced Python skills, with hands-on experience in FastAPI or similar frameworks, REST APIs, microservices, and enterprise integrations.
- Strong applied knowledge of LLM engineering, including prompt design, embeddings, chunking, vector databases, reranking, tool/function calling, and hallucination mitigation.
- Experience with at least one agent orchestration framework, such as LangGraph, LlamaIndex, OpenAI Agents SDK, or Semantic Kernel, and the ability to determine when an agent-based approach is unnecessary.
- Experience serving open-weight models in private infrastructure using Docker and Kubernetes, together with practical knowledge of LLM evaluation and observability.
- Ability to explain technical decisions and architectural trade-offs clearly to both technical and non-technical stakeholders, with strong customer-facing presentation skills and fluency in written and spoken Turkish and English.
- Experience delivering AI solutions in banking or financial services under KVKK and BDDK requirements, particularly involving core banking systems.
- Experience deploying and optimizing LLMs in air-gapped, closed-network, or on-premise environments using technologies such as vLLM, TGI, Ollama, or NVIDIA Triton.
- Practical experience with model quantization, inference optimization, and GPU and memory capacity planning.
- Experience in Turkish NLP, including Turkish model and embedding selection, dataset preparation, and quality evaluation.
- Experience developing AI platform or observability products using technologies such as OpenTelemetry, telemetry schemas, scoring pipelines, or SDKs.
- Practical knowledge of AI security, including red teaming, prompt injection testing, and the OWASP Top 10 for LLM Applications.
- Cloud or enterprise architecture certifications such as Azure AI Engineer, AWS Certified Machine Learning, Google Professional Machine Learning Engineer, or TOGAF.