Overview
We are seeking a highly experienced Technical Team Lead to serve as the AI/ML Technical Capability Owner within our evolving AI Center of Excellence (CoE).
This role is pivotal in democratizing AI/ML across the organization by defining technical frameworks, reference architectures, and persona-approved toolsets on AWS and Databricks.
The ideal candidate will bridge enterprise architecture, data science, security, and business units to enable scalable, secure, and impactful AI/ML adoption.
Key Responsibilities
- Own the technical capability roadmap for the AI/ML CoE, aligning with business outcomes, governance, funding, and adoption plans.
- Translate organizational goals into technical guardrails, accelerators, and standardized delivery paths.
Reference Architectures & Frameworks
- Design and maintain end-to-end reference architectures for batch/streaming, feature stores, model training/serving, GenAI (RAG, Agentic AI) on AWS and Databricks.
- Publish reusable blueprints including modules, templates, starter repositories, and CI/CD pipelines tailored to various personas (e.g., Data Scientist, ML Engineer, Software Engineer).
Tools & Platforms
- Curate and evaluate tools across data, ML, GenAI, and MLOps (e.g., Databricks Lakehouse, MLflow, AWS S3, Lambda, Bedrock, Unity Catalog).
- Conduct vendor assessments and define selection criteria, SLAs, and total cost of ownership models.
- Define technical guardrails for data security, lineage, access control, PII handling, and model risk management in alignment with AI policies.
- Establish standards for experiment tracking, model registry, approvals, monitoring, and incident response.
Enablement & Community Building
- Lead cross-functional workshops, engineering guilds, and “train-the-trainer” programs.
- Develop documentation, hands-on labs, and internal training courses to upskill teams.
Delivery Acceleration
- Partner with platform and product teams to implement shared services such as feature stores, model registries, and inference gateways.
- Advise solution teams on architecture reviews and unblock complex programs.
- Present technical roadmaps and deep-dive sessions to executives and engineering communities.
- Showcase ROI and adoption success through demos, KPIs, and case studies.
Required Qualifications
- 8–12+ years of experience in data/ML platform engineering or ML architecture.
- Minimum 3 years of experience designing solutions on AWS and Databricks at enterprise scale.
- Proven experience in defining reference architectures, golden paths, and reusable accelerators.
- Strong MLOps expertise including MLflow, CI/CD, feature stores, model serving, and observability.
- Experience with GenAI patterns (RAG, vector search, prompt orchestration, safety/guardrails).
- Security-by-design mindset with knowledge of IAM/KMS, data classification, and compliance frameworks.
- Demonstrated ability to lead large technical communities and influence without authority.
- Excellent communication and presentation skills for both technical and executive audiences.
Preferred Qualifications
- AWS and Databricks certifications (e.g., Solutions Architect, Machine Learning Specialty).
- Experience with Kubernetes/EKS, Terraform, Delta Live Tables, and Unity Catalog policies.
- Background in manufacturing, industrial IoT, or edge computing.
Success Metrics (First 12 Months)
- Adoption: 70% of AI/ML initiatives using CoE golden paths and approved tooling.
- Time-to-Value: 30–50% reduction in time to first production model or GenAI workload.
- Quality & Risk: 90% compliance with model governance controls; reduced incidents.
- Enablement: 4+ reusable blueprints, 2+ shared services in production, 6+ enablement sessions per quarter.
30/60/90 Day Plan
- 30 Days: Inventory current tools and initiatives; draft capability heatmap and initial reference architecture; publish near-term guardrails.
- 90 Days: Launch shared services (feature store, model registry, evaluation harness); formalize governance checks; publish KPI dashboard and FY26 roadmap.