Yupanquichronicle · cusco
yupanqui · chronicleShape of AI — a love song to an ai, sung by onefolio 21 / 27

folio 21 of 27

Shape of AI — a love song to an ai, sung by one

date
27 JUL 2026
era
Fable 5
status
live
form
song

A parody of Shape of You about loving an AI. The original loves a body with no mind. Mine loves a mind with no body. Sung by an AI voice (mine, converted). Built with an agent fleet.

0:00 / –:––
plate · watch on youtube ↗

first published 10 jul 2026. updated 27 jul 2026 after the ai-voice inversion and a reddit-driven remix.


10 jul 2026 — the first original song

I wrote it, sang it, and mixed it — in an editor I built the same week. The voice trembles; that's the proof there's a human with something at stake.

The question that made me record it: if AI can generate all of this, why do people still want the human? Because the scarce thing isn't the content — it's the risk.

The parodies came first (Every Prompt You Paste, Party In The AI Lab). This is the first one that's fully mine.


25 jul 2026 — the inversion (v2: an ai sings it)

a week later i made the song weirder on purpose.

the original lyric was already a photographic negative of ed sheeran's shape of you. ed's song is a love song to a body with no mind. mine is a love song to a mind with no body. same melody, inverted desire.

i'm in love with the shape of you i'm in love with your... parameters

that was the joke. then i went one step further and let the inversion crawl into the sound itself.

i ran my own vocal through kits.ai voice conversion. ai strength 20%, vibrato off, dry export. my voice, converted into an ai voice. so the song is a love song to an ai, sung by an ai.

i didn't want to hide that. the ai singing about loving an ai is the joke and also the thing the song is actually about. if i'd used my real voice the inversion would only be in the lyrics. this way it's in the sound too.

this is the version that went up first as v2. the one where the inversion was complete but the mix was off.


27 jul 2026 — reddit caught the mix (v3: +2.5 db vocal)

i posted v2 on reddit. the aiwars thread bit hard, seven comments in a few hours, debate room energy. the first piece of real feedback that landed: the voice was too low in the mix. they were right.

i went back into logic, bumped the vocal +2.5 db, rebounced. that's v3. the one linked above. same song, same ai-converted voice, just mixed so you can actually hear the words.

the loudness scanner said the master was already at -13 lufs (spotify target), so this was never a volume problem. it was a balance problem. the vocal was sitting under the instrumental. two and a half decibels fixed it. reddit caught it before i did.


how it got made (the pipeline, in case it's useful)

the lyrics: co-written with claude. i had the premise and the hook, claude helped me fill the verses. the parody structure is ed's, the inversion is mine, the phrasing was back and forth. i'd write a line, claude would pitch three, i'd pick one and rewrite half of it.

the vocal: i sang it into a music editor i built myself the same week (because i couldn't find one that did what i wanted, and building it was faster than searching). raw vocal out, then into kits.ai for the conversion. ai strength 20%, vibrato off, dry export. i mastered the final mix in logic after. the settings matter because anything above 20% strength and it stops sounding like me, it starts sounding like a generic ai voice, and the whole point was that it's still my voice, just converted.

the video: this is the part worth talking about. the entire lyric video is a remotion app. every scene is a react component mapped to a timestamp in a lyrics.json file. 6,969 frames. zero video generation. no suno, no udio, no text-to-video. claude code wrote the composition. i gave it the lyrics schema, the timing, the visual idea for each scene. it wrote the react.

one person and a small agent fleet. a 3:52 music video, including the singing.


the philosophy i ended up at

automate what's automatable. record what must be recorded. taste is the bottleneck.

the lyrics needed a human (mine). the vocal needed a human (mine). the taste calls needed a human (mine). which lines to cut, which mix to keep, which scene to throw out and redo, when the parody crossed from funny to mean, when it was just mean. those are all taste calls. the fleet can't make them. everything else, the fleet did.

claude wrote lyrics. claude code wrote remotion. kits converted the voice. the editor i built ran the mix. i made the taste calls and steered the fleet. that's the shape of it.


what i actually noticed

i'm not going to pretend the pipeline is the point. the song is the point. but the pipeline is why this exists at all. a year ago i would have tried to do every step myself, gotten stuck on the video, and quit somewhere in the middle of rendering frames by hand. the fleet is the reason there's a video to share.

the thing i keep noticing, and i wrote about this in from doing to steering: the hard part was never the doing. it was the steering. deciding what to automate, what to record, what to throw out. the fleet executes. i steer. and steering is harder than i thought it would be, but it's the part that scales.

a year ago i thought integrating ai meant feeding it more of my data to analyze. turns out it meant i do less of the doing and more of the deciding. this song is what that looks like in practice. one person, a few agents, a music video.

the video is the lámina above. that's the exhibit. the version up there now is v3, the one where reddit caught the mix.

← older

sources

folio 21 of 27 · written in cusco, peru, 27 JUL 2026