The voice comes in slightly wrong. It has the shape of a human singing — vibrato, breath, the small catch at the top of a phrase — but the edges are soft in a way no throat produces. Stack four of them and you get a choir that never stood in a room together.
That's Holly+, and the voice belongs to Holly Herndon, except when it doesn't.
Herndon has been building machine-learning instruments and singing with them since well before the current wave. She is a composer with a doctorate from Stanford's computer music program, and her work keeps circling one question: what happens to a voice when it can be copied, and who should own the copy.
Her answer has been to build the copy herself, and then hand out the keys.
That's her 2022 cover of Dolly Parton's "Jolene," sung by her own digital twin. The instrumentation is by Ryan Norris and the video is by Sam Rolfes, who motion-captured a 3D model of Herndon to make the thing on screen move. The lead vocal was generated by feeding a modified score into Holly+ and letting the model sing it in her voice.
Spawn took six months to stop being boring
Before Holly+ there was Spawn, an AI she and Mat Dryhurst built for the 2019 album PROTO. Spawn was not an API call. It was hardware.
"We bought Spawn's parts and built her in our studio," Herndon told The FADER in 2019. "Jules installed an operating system and some software, we created our own training sets, and we started listening to the outcome."
Jules is Jules LaPlace, the developer who worked on the ensemble alongside them. And the outcome, for a long stretch, was nothing much.
"We had about six months of boring results before we started to get interesting results," she said.
The turn came on a specific piece of audio. "The spoken part of 'Birth,' which is trained on my voice, was the first time we were like, 'You can hear the logic of the neural network at work,'" Herndon told the same interview.
The part that gets lost when PROTO is described as an AI album is how much of it is people. Herndon assembled a vocal ensemble in Berlin and ran live training sessions, where singers performed and Spawn listened.
"After touring Platform for years, we were really missing communal music-making," she said. "We put together this motley crew of individuals in Berlin and started experimenting."
She has put a number on the machine's share of the record. "It really only makes up 20% of the audio."
Training on herself was the ethical shortcut
The path from Spawn to Holly+ came out of a decision about training data, and she has been direct about why.
"We realized we could create a naturalistic likeness and the only approach that we felt comfortable with, was training on ourselves directly," she told Ars Electronica in 2022. "That's when I started just using my own training data and Holly+ was born."
The instruments were built with Herndon Dryhurst Studio, Never Before Heard Sounds, and Voctro Labs. The result is a model anyone can sing through, hosted at the Holly+ project site.
Giving it away was the point, and she said so when she announced it.
"I am releasing Holly+ in collaboration with Never Before Heard Sounds, the first tool of many to allow for others to make artwork with my voice," she wrote, in remarks quoted by MusicTech. "My voice is precious to me! It is 1 of 1."
Both things at once. Precious, and handed out.
Dryhurst described the mechanism plainly: "The novel idea was, what if we gave that model to everybody to be able to use and wrapped it in a protocol that would share profits from any media created with her voice 50-50 back to Holly?"
She calls the model the artwork
The framing Herndon uses for all of this puts the weight on the training set rather than the output.
"We focus so much on the models and their outputs, but we often forget to talk about all the training data that goes into the models," she told Art Basel in October 2024.
In the same conversation she said something that reframes what she thinks she's making: "The model is the artwork … It's the model that can generate infinite artworks, in any kind of medium."
And on why ownership gets complicated: "AI models require large amounts of data to work well. They are collective accomplishments that require experiments in collective ownership and compensation."
She has also described the training data itself in generational terms — "we like to think about the training data as children that we're sending into the future because it will be training models for decades and decades to come."
If you want the adjacent territory, imoliver's songwriting process sits at the opposite end of the same problem — he writes every word himself and generates the performance, where Herndon built the performer and let other people write.
Herndon and Dryhurst are still working the data question rather than the output question. The voice is out there, the protocol splits the money, and the model keeps being the thing she points at when someone asks what she made.