Computing Short Films using Language-guided Diffusion and Vocoding through Virtual Timelines of Summaries

DOI https://doi.org/10.51191/issn.2637-1898.2023.6.10.71

Luís Arandas (https://orcid.org/0000-0001-5413-2832) University of Porto – INESC-TEC, Porto, Portugal

Author’s contact information: luis.arandas@inesctec.pt

Miguel Carvalhais (https://orcid.org/0000-0002-4880-2542)

University of Porto – i2ADS, Porto, Portugal

Author’s contact information: mcarvalhais@fba.up.pt

Mick Grierson (https://orcid.org/0000-0002-6981-5414)

University of the Arts London – CCI, London, United Kingdom

Author’s contact information: m.grierson@arts.ac.uk  

INSAM Journal of Contemporary Music, Art and Technology Main Theme of the Issue: Technological Aspects of Contemporary Artistic and Scientific Research

Publisher: INSAM Institute for Contemporary Artistic Music, Sarajevo, Bosnia and Herzegovina  

Section: THE MAIN THEME

Abstract: Language-guided generative models are increasingly used in audiovisual production. Image diffusion allows for the development of video sequences and some of its coordination can be established by text prompts. This research automates a video production pipeline leveraging CLIP-guidance with longform text inputs and a separate text-to-speech system. We introduce a method for producing frame-accurate video and audio summaries using a virtual timeline and document a set of video outputs with diverging parameters. Our approach was applied in the production of the film Irreplaceable Biography and contributes to a future where multimodal generative architectures are set as underlying mechanisms to establish visual sequences in time. We contribute to a practice where language modelling is part of a shared and learned representation which can support professional video production, specifically used as a vehicle throughout the composition process as potential videography in physical space.

Keywords: artificial filmmaking, deep generative models, language-guided diffusion, short film computing, audiovisual composition, multimodal sequencing.

PDF:

6. INSAM Journal 10, Arandas et al

ISSN 2637 – 1898
On the cover: Photo from the premiere of the opera Third Bullet by Vojislav Vučković, generated with DALL-E, initiated by Milan Milojković
Design and layout: Milan Šuput, Bojana Radovanović