
Visual & Multimedia Generative AI
A hands-on course on generating, editing, and controlling images and content with AI.
Online or on-site
8 hours
Professionals, Researchers
Access the course
An open preview of the course, covering concepts, exercises, and responsible-use principles that remain useful as tools change. The live course adds the latest platforms, workflows, and practical case studies.
Read the notesAbout This Course
A practical, beginner-friendly course on visual and multimedia generative AI. It is designed to help you understand what these models can and cannot do, how to interact with them, how to generate and edit images, how to preserve styles and characters, and how to use them responsibly in professional and administrative settings.
The specific tools may change over time, in fact, I had to restructure the course several times since 2023. For that reason, we prioritize a broader, more transferable approach that will keep you prepared for future changes.
The openly published notes focus on that lasting foundation and offer a preview of the course's approach. The live training goes further: we work with specific tools, compare their capabilities, and build workflows suited to the current landscape and each group's needs. That material is updated between editions and reserved for the sessions.
Course Content
Day 1: Introduction and image generation
- • Generative AI versus traditional AI
- • Multimodal models
- • Prompt engineering
- • Iteration: the visual conversation
- • Comparing ChatGPT, Gemini, and Copilot
Day 2: Understanding, editing, and controlling images
- • Image-to-text and image analysis
- • Content, subject, style, and composition references
- • Generative editing: add, remove, and replace
- • Inpainting and outpainting
- • Character, product, and style consistency
Day 3: From image generation to content production
- • Text inside images: posters and labels
- • Infographics and diagrams
- • Case study: an institutional campaign
- • Consistency across formats
- • Model, service, and application
- • Hugging Face and open models
Day 4: Video, other modalities, and responsible use
- • Image-to-video and text-to-video
- • Voice, voice cloning, and music
- • What AI still gets wrong
- • Privacy, intellectual property, and image rights
- • Deepfakes, transparency, and the AI Act
- • Content provenance: C2PA and SynthID
- • Text watermarking: Claude, Gemini, and other models
Includes
- • 4 sessions of 2 hours each
- • Open notes and up-to-date, course-specific materials
- • Hands-on exercises in every session
- • Individual mentoring
- • Certificate of completion