
What makes a lyrics video different from a visualizer
A visualizer reacts to sound. A lyrics video reacts to meaning. In a visualizer, the audio curves drive color, light, and motion; the result stays abstract, closer to an album cover coming alive. In a lyrics video, the words themselves become a design element. They carry weight, rhythm, and emotion on their own, and they have to stay legible while everything around them moves. That changes the whole approach: typography is no longer decoration on top of a scene, it’s a character in it.
Reading the lyrics and the vocal stem before touching a keyframe
Before opening any 3D or motion software, I read the lyrics several times; first for meaning, then for structure: verse, chorus, bridge, ad-libs. Then I isolate the vocal stem and listen to it alone, without the instrumental. This is where the real information is: where the singer breathes, which words get pushed, which lines are almost whispered. A lyric on paper and a lyric as performed are two different things, and the video has to follow the performance, not the lyric sheet.
Keeping typography and the 3D/motion environment coherent
Typography and the 3D environment need to feel like they exist in the same physical world, not like a caption layered on top of a render. That means the text catches the same light as the scene, reacts to the same particle systems, gets depth of field and motion blur consistent with the camera. If the environment is glossy and volumetric, flat static text breaks the illusion immediately. The goal is a single coherent universe where letters are just another object with mass, light, and behavior.
Formats: cutting the same edit for Reels, Spotify Canvas and YouTube
One track, one visual world, three very different deliverables! Reels and TikTok need vertical 9:16 framing with safe zones left clear for captions, profile icons, and UI overlays. Spotify Canvas is a silent, seamless 3–8 second loop — no beginning, no end, just a fragment that has to work on repeat. YouTube can carry the full-length journey, in either vertical or standard framing, with room for slower pacing and more detail. Rather than re-editing from scratch for each platform, I build the 3D scene and camera moves with all three formats in mind from the start, then recompose and re-time per format at the end.
Have a project like this in mind?
Let's talk