About the Multimodal & Generative category

Vision-language models, diffusion, image/video generation and anything crossing modalities.