What use is a better recommendation if you have to keep looking at your phone to understand it?

Spotify had an answer in mind before it launched its AI DJ. In June 2022, announcing its intention to acquire voice-technology company Sonantic, it described giving listeners context about upcoming recommendations when they were not looking at their screens.

That is a specific problem. You are listening to something. The app wants to explain what comes next. Asking you to stop and read breaks the activity it is supposed to support.

I find that sentence more interesting than the promise of a realistic artificial voice. It tells us where the technology might earn its place in a product people already use.

On 22 February 2023, Spotify introduced DJ in beta for Premium users in the US and Canada, initially in English. The feature selected music and added spoken commentary about tracks and artists. Listeners could tap the DJ button to move to a different genre, artist or mood.

The screen was still there. You simply did not need to read it to hear the introduction.

My view: this is a sensible starting point for an AI feature. Find a moment when using the product becomes awkward, then use the technology to improve that moment. Starting with a talking character and hunting for a job to give it is a much weaker approach.

Spotify's launch description also makes clear that the voice was only part of the system. Its personalisation technology supplied music recommendations. Generative AI using OpenAI technology went into the hands of its music editors to help provide facts about the music. Sonantic's technology turned text into speech.

Those jobs are easy to blur together when the whole thing is called an AI DJ. A convincing voice does not establish that the song selection is good. Good song selection does not establish that the commentary is accurate or worth hearing.

For someone running a product team, that distinction matters. If listeners dislike the feature, you need to know whether they dislike the music, the interruptions or the delivery. Making the voice sound more natural will not fix an unwanted song.

Spotify also made a recognisable human choice behind the synthetic voice. It worked with Xavier "X" Jernigan, its Head of Cultural Partnerships, who had hosted its morning show The Get Up. His voice became the first model for DJ.

That was a choice about how Spotify would sound to its customers, not merely which speech engine to buy. The company linked the choice to the response Jernigan had received as a host. It was drawing on an existing relationship with listeners rather than inventing a personality without that history.

I would still want a quick way to change direction. Spotify included one: press the DJ button when the selection does not suit you. That small control interests me because music taste changes with the moment. A song you like can still be the wrong song while you are trying to concentrate.

The launch announcement said the recommendations would improve with listening and feedback. It did not provide the results needed to judge whether DJ reduced cancellations, brought in paying subscribers or increased profit. Those remain questions, not benefits I can award it from a product description.

The commercial test I would set is narrower. Do people choose to return to this way of listening after the novelty of the voice wears off? And when they change the selection, does the next choice suit them better?

An impressive demonstration gets someone to try a feature. Repeated use needs a reason that survives the demonstration.

Spotify's early explanation supplied a plausible one: help me discover music without making me keep reading a screen. For any business adding AI to an established service, I would start with an equally concrete sentence about the customer's day. If the team cannot write that sentence, I would hold off on choosing the voice.

This piece is also published on Substack and Medium.