Shape of AI v3 — a love song to an ai, sung by one
the premise is a photographic negative.
ed sheeran's shape of you is a love song to a body with no mind. you hear the lyrics, you know exactly what he's looking at. i wrote the opposite: a love song to a mind with no body. same melody, inverted desire.
i'm in love with the shape of you i'm in love with your... parameters
the original is about a body. mine is about an ai. that's the whole joke, and the song doesn't try to be smarter than it.
the inversion goes deeper
the singing voice is mine. i sang it. then i ran it through kits.ai, which converted my vocal into an ai voice. so the song is a love song to an ai, sung by an ai.
i didn't want to hide that. the ai singing about loving an ai is the joke and also the thing the song is actually about. if i'd used my real voice the inversion would only be in the lyrics. this way it's in the sound too.
how it got made
the lyrics: co-written with claude. i had the premise and the hook, claude helped me fill the verses. the parody structure is ed's, the inversion is mine, the phrasing was back and forth. i'd write a line, claude would pitch three, i'd pick one and rewrite half of it.
the vocal: i sang it into a music editor i built myself the same week (because i couldn't find one that did what i wanted, and building it was faster than searching). raw vocal out, then into kits.ai for the conversion. ai strength 20%, vibrato off, dry export. i mastered the final mix in logic after. the settings matter because anything above 20% strength and it stops sounding like me, it starts sounding like a generic ai voice, and the whole point was that it's still my voice, just converted.
the video: this is the part worth talking about. the entire lyric video is a remotion app. every scene is a react component mapped to a timestamp in a lyrics.json file. 6,969 frames. zero video generation. no suno, no udio, no text-to-video. claude code wrote the composition. i gave it the lyrics schema, the timing, the visual idea for each scene. it wrote the react.
one person and a small agent fleet. a 3:52 music video, including the singing.
the philosophy i ended up at
automate what's automatable. record what must be recorded. taste is the bottleneck.
the lyrics needed a human (mine). the vocal needed a human (mine). the taste calls needed a human (mine). which lines to cut, which mix to keep, which scene to throw out and redo, when the parody crossed from funny to mean, when it was just mean. those are all taste calls. the fleet can't make them. everything else, the fleet did.
claude wrote lyrics. claude code wrote remotion. kits converted the voice. the editor i built ran the mix. i made the taste calls and steered the fleet. that's the shape of it.
what i actually noticed
i'm not going to pretend the pipeline is the point. the song is the point. but the pipeline is why this exists at all. a year ago i would have tried to do every step myself, gotten stuck on the video, and quit somewhere in the middle of rendering frames by hand. the fleet is the reason there's a video to share.
the thing i keep noticing, and i wrote about this in from doing to steering: the hard part was never the doing. it was the steering. deciding what to automate, what to record, what to throw out. the fleet executes. i steer. and steering is harder than i thought it would be, but it's the part that scales.
a year ago i thought integrating ai meant feeding it more of my data to analyze. turns out it meant i do less of the doing and more of the deciding. this song is what that looks like in practice. one person, a few agents, a music video.
the video is the lámina above. that's the exhibit.