Job Type: Full-Time (Hybrid Mode)
Sprouts.AI is an AI-native “Generative Demand Platform” disrupting the $1.1T B2B GTM category. We help demand-gen and sales teams generate qualified pipelines through real-time intelligence signals and agentic AI execution. Customers are already seeing:
- 50%+ increase in SDR/BDR productivity
- 30%+ reduction in martech spend
As we scale, we’re building the next-generation infrastructure for massive data ingestion, vector search, multi-agent orchestration, real-time LLM workloads, and GPU-powered inference. We’re hiring a Foundational AI Infrastructure Engineer who can architect and operate this system end-to-end.
This is not a typical DevOps role.
You will not “wait for requirements.”
You will read the code, understand the system, and shape the infra proactively.
If you want to build the backbone of a high-velocity, agentic AI platform, this is your role.
What You’ll Own (Outcomes, Not Tasks)1. AI-Native Infrastructure Ownership
- Architect and own the compute, storage, networking, vector search, orchestration, and inference layers.
- Derive infra needs directly from backend + AI pipelines.
- Build infra that supports massive concurrency, low latency, and unpredictable agent workflows.
2. GPU & LLM Inference Readiness
- Deploy and optimize LLM runtimes (vLLM, TGI, TensorRT-LLM).
- Configure GPU autoscaling, batching, quantization, caching.
- Ensure multi-cloud portability and avoid vendor lock-in.
3. Event-Driven Architecture for Agentic Systems
- Implement backpressure, retries, DLQs, and idempotent pipelines.
- Support parallel agent execution and high-throughput ingestion.
4. High-Performance API Gateway & Traffic Control
- Implement rate limits, quotas, throttling, and tenant-isolated traffic shaping.
- Prevent token floods and agent-triggered burst loads.
5. Internal Developer Platform (IDP)
- Build reusable infra templates, CI/CD scaffolds, local dev reproducibility, canary deployments, rollback mechanisms.
- Enable engineers to ship in hours, not weeks.
6. Observability “Everywhere”
- LLM-level observability (token usage, eval metrics, latency maps).
- Build dashboards for agent behavior, vector search performance, API latency, GPU utilization.
- Build transparent cost dashboards for compute, storage, token usage, and inference.
- Optimize infra for cost-performance (batching, caching, autoscaling, routing).
- Keep infra costs predictable as LLM usage scales.
8. Data Lifecycle & Index Management
- Define TTL, retention, compaction, cold storage rules.
- Ensure metadata consistency at scale.
9. Security, Governance & Multi-Tenancy
- Ensure safe tool execution for agents (sandboxing, ACLs, rate limits).
- Zero Trust patterns for internal services.
10. Code-Level Performance Insights
- Read Python/Node code to identify bottlenecks, parallelism issues, inefficient queries, unbounded concurrency, missing caching.
- Suggest backend and AI pipeline optimizations that reduce infra load.
Tech You’ll Work WithCloud, Compute & Orchestration
Terraform (must)
Docker
Load balancers, service mesh
AI/LLM Infrastructure
vLLM, TGI, TensorRT-LLM
Batching, quantization, caching
Token routing + observability
Data & Vector Search
Monitoring & Reliability
OpenTelemetry traces
Loki / ELK stacks
P95/P99 latency optimization
GitHub Actions / GitLab
GitOps (ArgoCD optional)
Infra templates, one-click deploys
Must-Have Skills (Hard Filters)
- Strong production experience with Kubernetes + Terraform
- Cloud infra ownership (AWS/GCP/Azure)
- Experience with GPU or model-serving infra (even small-scale)
- Hands-on with vector DBs or high-performance search
- Ability to read Python/Node code and infer infra impact
- Proven history of optimizing complex systems for cost + latency
- Real observability implementation experience (metrics, traces, logs)
Nice-to-Haves (Signal Boosters)
- vLLM or TGI deployment experience
- Knowledge of LangGraph/LangChain internals
- Experience building IDP / self-serve platforms
- Advanced GPU optimization experience
- Built multi-region or multi-cloud systems
- Experience with OpenTelemetry for multi-agent architectures
- Hands-on RAG or AI pipeline infra work
- Optimization of P95/P99 latencies
You’ll Thrive Here If You…
- Prefer proactive problem discovery over reactive “fix requests.”
- Can zoom between 5,000-foot architecture and 5-line code-level insights.
- Love squeezing milliseconds out of systems.
- Believe infra should accelerate product velocity, not slow it down.
- Enjoy the chaos, creativity, and pace of a high-speed AI-native company.
Please Don’t Apply If…
- You want a traditional DevOps job focusing only on deployments and clusters.
- You cannot read backend code or understand LLM/agent workflows.
- You rely too heavily on vendor-managed AI services.
- You need strict processes and long planning cycles.
- You’ve owned infra for a high-scale or AI-heavy system.
- You’ve debugged distributed systems in production.
- You can explain cost vs. latency trade-offs clearly.
- You’ve tuned vector DB search, async workers, or GPU workloads.
- You’ve built observability or IDP tooling used by other engineers.
- You have a “this is inefficient” instinct just by scanning a diagram or code snippet.
What You’ll Get
- A mission-critical role shaping the foundation of Sprouts.ai’s AI-native platform.
- Ownership of infra decisions across compute, data, storage, AI inference, observability, and cost.
- The opportunity to architect GPU-ready infra for future private LLM deployments.
- Fast track to Head of Infra / Platform for strong performers.
- A culture of rapid iteration, autonomy, and engineering excellence.
- Deep exposure to frontier AI systems, multi-agent orchestration, and large-scale vector search.