Turning Music into Images

A weekend experiment exploring multimodal AI by transforming live piano performances into dynamic AI-generated images. The system successfully mapped musical features to emotional and visual concepts, creating a novel form of audio-visual expression. I developed the complete pipeline from MIDI capture to image generation, exploring underrepresented modalities in AI research.

projectentertainment
Year: 2024

Timeline Details

About This Item

Inspired by the rise of multimodal AI models like Gemini, this project explored an underexplored modality: translating music into images. The workflow captures MIDI output from a keyboard, extracts features such as speed, register, and complexity, and uses an LLM to translate these into emotions, styles, and color palettes. These outputs condition a Scribble ControlNet model, turning a turtle sketch into dynamic, mood-driven visuals. Areas for improvement include more accurate handling of rhythm, faster sequential image generation with potential latent-space journeys, and richer audio-to-visual translation using CLAP for full audio feature extraction. The project was influenced by Pymidifile for MIDI parsing inspiration, Kasper Jordaens' demo at the first Generative AI Belgium meetup, and discussions with peers like Stijn Spanhove and Sebastiaan Van den Branden on creative AI in music.

Year

2024

Skills & Technologies

projectentertainment