Case studySolo-built product (e-commerce)
The challenge
Viewers constantly want to buy fashion they see on screen but screenshots lead nowhere. Calling multimodal LLMs (Gemini/GPT-4V) per frame cost $50+/episode.
Built with
- Next.js
- FastAPI
- PyTorch
- YOLO Fine-tuning
- SLMs (Llama, Mistral)
- LLM Pre-annotation
- Human Annotation Guidelines
- Gemini Vision
- GCP Cloud Run