An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Tykhe Inc. is seeking a producer of VLM and multimodal AI systems for on-site work in San Francisco. You will ship production-grade visual-language models, orchestrate end-to-end workflows, and contribute to tool-use driven agentic AI projects.
Ideal candidates have 0–3 years of experience, strong Python/PyTorch skills, and a background in vision or multimodal research. A startup or high-intensity environment is highly valued.
Shipped agentic VLM / multimodal vision systems to real users in production — not demos, not pure research. Owned the model + orchestration layer end- to- end. Must be vision- language multimodality (images/video + VLM), not sensor- fusion or audio- only.
Applied VLM engineering: shipped and hardened vision- language models in production — visual reasoning, detection/segmentation as needed. Depth in applied fine- tuning (SFT, RLHF), eval design, and model orchestration, not pretraining from scratch.
Must work on- site 5 days/week in San Francisco. Ideally open to living in a hacker house.
0 - 3 years of experience shipping production VLM / multimodal AI systems, using Python and PyTorch
Startup or founding- engineer experience (founder, CTO, eng #1–5) OR high- intensity company (Meta Reality Labs, Anduril, SpaceX, top YC). Long single big- tech tenure with no builder signal and no relevant- team work is a negative.
BS or MS in CS, ML, or Engineering from a strong program, OR demonstrated equivalent via shipped production AI systems.
Agentic AI with real tool- use: built multi- step agent loops (not RAG wrappers or single- turn chatbots). Experience with eval harnesses and data flywheels.
Genuinely mission- aligned with computer vision, wearables, and industrial AI — not just chasing LLM trends. Evidenced by project choices, side work, or career trajectory.
Production AR / wearable AI (Meta Reality Labs, Snap Spectacles, Apple Vision Pro) or autonomous driving CV (Waymo, Cruise, Aurora).
Master's with vision or multimodal research component (published work or thesis).
Edge inference / on- prem model deployment (vLLM, Triton, TensorRT, quantization). Serving open- weight models on constrained hardware.