Building Vision AI with Foundation and Generative Models

Building Vision AI with Embodied Vision Models

A code-first course on building vision systems with foundation and generative models. Part of the Hands-On AI Science series, designed around Innovation-First Learning principles.
Companion Online Textbook

Vision AI Tasks

Vision AI has moved from narrow classifiers to foundation models that understand, describe, and generate images. Every product that touches cameras, documents, or visual content now depends on these capabilities. This course prepares students to build vision systems that see, reason, and create.

Foundation & Generative Models

Core concepts, models, and ideas behind modern computer vision: convolutional and vision transformer architectures, contrastive learning, diffusion processes, latent space geometry, and multimodal alignment between images and text.

Tools & Platforms

PyTorch, OpenCV, Hugging Face Diffusers, Stable Diffusion, YOLO, SAM, CLIP, Roboflow, Weights & Biases, and Google Colab.

Modular Syllabus

A specific course syllabus is built for each audience: graduate or undergraduate, across engineering, digital health, or computer science.

Innovation Through Tools Mastery

As AI and mature libraries handle standard tasks, professional developers must focus on innovation. Student projects tackle new use cases by generating unique data and fine-tuning task-specific vision models.

Guided Student Projects

Students begin their projects while learning the material and enrich them as new concepts arrive. Each team gives several in-class presentations for discussion and feedback.

Typical Weekly ScheduleSample Syllabus (PDF)Poster (HIT)

Week 1

Image Processing Fundamentals

OpenCV, NumPy, filtering, edges

Week 2

CNNs & Image Classification

PyTorch, ResNet, transfer learning

Week 3

Object Detection

YOLO, Roboflow, bounding boxes

Week 4

Semantic Segmentation

SAM, U-Net, pixel-level masks

Week 5

Project Proposal Presentations

Student proposals, peer feedback

Week 6

Vision Transformers

ViT, DINOv2, Hugging Face

Week 7

Multimodal Models & CLIP

CLIP, BLIP-2, visual QA

Week 8

Interim Project Presentations

Progress demos, instructor feedback

Week 9

Generative Models: GANs

StyleGAN, image-to-image translation

Week 10

Diffusion Models

Stable Diffusion, HF Diffusers

Week 11

Image Editing & ControlNet

Inpainting, ControlNet, img2img

Week 12

Video Understanding & 3D Vision

Optical flow, depth estimation, NeRF

Week 13

Final Project Presentations

Live demos, peer evaluation

This modular series is taught as concrete courses. Full syllabi for the specific Vision AI offerings:

2026 Graduate
Generative AI: VAEs to World Models HIT

Generative AI: From Variational Autoencoders to World Models

A 13-week graduate course on generative modeling, spanning variational autoencoders, diffusion models, and world models.

2026 Graduate
Deep Generative Models: Audio-Visual BIU

Deep Generative Models for Audio-Visual Data

A 13-week graduate course on deep generative models for audio-visual data, covering VAEs, GANs, diffusion and the Stable Diffusion family, and synthetic data generation.

2026 Undergraduate
Deep Generative Models: Visual HIT

Deep Generative Models for Visual Data

A 13-week undergraduate course on deep generative models for visual data, covering VAEs, GANs, diffusion and the Stable Diffusion family, and synthetic data generation.

2026 Undergraduate
Images & Vision: Pixels to Deep Learning BIU

Images and Vision: From Pixels to Deep Learning

A 13-week undergraduate course tracing computer vision from image formation and classical image processing through classical vision and 3D reconstruction to deep learning for detection and segmentation, taught with a code-first approach.

Building Language AI

Language AI

LLMs and Agents

Building Scalable AI

Scalable AI

Big Data and Distributed Intelligence

Building Temporal AI

Temporal AI

Sequential Intelligence and RL