Powered by AISKOOL
Vision-Language Models: CLIP, SAM & Multimodal Apps
The new CV stack: CLIP-style embeddings for zero-shot search and classification, SAM for promptable segmentation, multimodal LLMs for visual reasoning — and how to compose them into products classic CV can't touch.
3 sections·6 lessons
What you'll learn
1. CLIP: Images and Text, One Space
- The embedding trick that changed everything
- Check: CLIP
2. SAM: Segment Anything, on Demand
- Promptable segmentation as a building block
- Check: SAM
3. Multimodal LLMs & Composing the New Stack
- Visual reasoning, and when to use which tool
- Final check: the new stack
Share this course