Faking human music with computers (improved).
Automated piano covers from any song? Basically, yes.
This is a follow-up to my previous post on pop2piano. The results back then were not too bad, but they only worked on a fairly narrow range of songs, and still required a lot of manual work to get something good.
SheetSage2 is the audio-to-score transcription model released alongside the YuE2 music generation model. It simply improves on everything, basically.
Here is an 18-song playlist of covers, spanning rock, electro, orchestral music, and some weird picks.
The playlist is not random: it is the same set of songs I used to evaluate pop2piano, so you can compare them directly with the old playlist. Back then, I picked songs I liked, but only kept the ones where the result was listenable. On a lot of songs, the output was just banging on a single note over and over again, even on very melodic songs, as soon as instruments outside the pop palette showed up (e.g. distorted guitars).
With SheetSage2, even the songs I had rejected came out at least quite good. The range of music it handles seems essentially generalist.
I reused the same pipeline as last time; the whole pipeline is in this gist. The only change was swapping pop2piano for SheetSage2, which conveniently already comes with a CLI.
It runs locally; the weights are on Hugging Face, under a non-commercial license. There is no official UI, only a few community demos, but the command line is all you need:
python infer.py song.mp3 --output output --render-audio
This gives you an editable score (ABC), MIDI files for the melody and chords, and a piano rendering of the result. All my tests, generative and transcription, ran in under 2 minutes, so it’s a surprisingly efficient model.
The project can also generate songs from prompts with YuE2, and you can steer the generation by feeding it a score, such as the ones SheetSage2 produces. I’ll certainly try that out.
To finish, here’s the transcription of Cosmoose and Calla Soiled’s Heroic: