Teaching LLMs to Speak Spotify — Yves Raimond & Jacqueline Wood, Spotify

summarized

TLDR

Spotify's Large Taste Model, used daily by a quarter of US premium subscribers, shifts personalization from ranking to reasoning by embedding catalog items as semantic IDs into an LLM's vocabulary. The NEO training paradigm preserves the LLM's language ability while adding domain knowledge through a frozen backbone stage, enabling steerable and explainable recommendations across surfaces like DJ sessions and prompted playlists. The system demonstrates that generative personalization can be deployed at industrial scale with gains in autoplay, podcast discovery, and user engagement.

Key points

Spotify's Large Taste Model is used daily by about one in four US premium subscribers.

The NEO training paradigm includes a domain grounding stage where the LLM backbone is frozen to preserve language ability.

Ablation studies showed that continuous pre-training degrades the LLM's natural language capabilities to near zero.

Grounded LLM judges using textual user profiles achieved 75% alignment with human preferences for podcast recommendations.

Techniques

  • Semantic ID quantization
  • NEO training paradigm
  • Domain grounding with frozen backbone
  • Multitask instruction tuning
  • Grounded LLM judges
Transcript (captions)

0:01 [music] Welcome everyone. Um really glad and thank you Deven for inviting us as well to talk today. Uh so um uh Jackie and I are going to present today about making

0:23 LLMs speak Spotify and how we turned our recommendation system on its head to be LLM native. Uh just a few quick words about myself. Uh I joined Spotify a year ago. Before that I was at Google working

0:37 on personalization for Google search and prior to that I was uh working on personalization at Netflix. Um all right so let's dive in. Um first a couple of numbers to describe

0:51 the scale of the problem that we have to solve. Um Spotify has about I was a little bit sad not to see it on the DAU by chart because it should be up there but it has about 760 million uh active

1:04 uh users monthly uh in about 184 markets. But one thing that makes uh the Spotify personalization problem particularly challenging is the size of its catalog. Uh so of course you know

1:16 Spotify has basically all music ever published a bit more than 100 million uh music tracks but it also has a range of videos, podcasts and audio books as well. So the matching problem uh is uh

1:29 actually surprisingly uh complicated. So today what we're going to talk about Jackie and I is basically the extent of this matching problem and how we uh how we are solving it. I'm going to talk

1:40 about a little bit of history, the new phase that we're entering and then Jackie is going to help us uh go into the guts of the modes as well to understand how those things are trained

1:49 as well. Uh so a little bit on history. Uh initially Spotify personalization was really built around curation. So uh people basically manually assembling playlists that target specific taste.

2:02 And that's still a very common uh use case on Spotify. Right now we have about 10 billion playlists with many many uh created every every hour. Uh then you know Spotify moved into

2:14 taking these creation signals and order signals and moving into recommendations. So basically being able to turn these uh creation signals to something that can be applied at scale. Uh and a great

2:26 example of that would be discover weekly for example that was launched in 2014 as one of the one of the early like recommendation use case on on Spotify. But the phase that we entering now which

2:38 uh uh which we're going to talk about in more details is something that we call generative personalization where we are not only solving a matching problem from the user to the content. We're also

2:48 solving the ability to generate an experience that is interactively and dynamically shaped around each user. And that transition from recommendations to generative personalization involves a

3:03 couple of different uh a couple of different big shifts. One is move from moving moving from personalization as guessing uh where basically you have a ranking algorithms that spits out a

3:15 brand ranked list of entities in your catalog to personalization as reasoning that can introspect these results and really try to understand whether that's indeed the right match for this user in

3:25 this particular context. The other aspect as well is moving from blackbox algorithms. uh it's very difficult to fully introspect a multi-stage ranking system for example

3:38 to transparent and steerable personalization where the user is always fully in control. So uh basically giving the ability to these mods to speak and understand English.

3:50 The other thing that these these systems can do is uh uh not stopping just at recommending but also generation of experiences and explaining as well. uh and we're going to show a couple of

4:02 examples of that. So one example that uh we launched a couple of years ago uh so pretty early in that journey was the Spotify DJ. Uh so the Spotify DJ is something that you

4:14 can spin up that will start playing music for you of course personalized uh and but one interesting thing about it is that since last year you can tap that button on the bottom uh right and

4:26 steer it in whatever direction you see fit. So at any point you can chime in and let the algorithm know what you want and it's going to steer the session in the in in that in that direction.

4:37 Uh another example of what we call generative personalization is uh showcased in a prompted playlist here that Deanch showed a little bit earlier as well. Uh and here you can you

4:49 basically have full unfettered access to the recommendation algorithm that Spotify has. uh and you can uh prompt it with very high level prompts or very detailed prompts. On the left hand side

5:01 here, I have a prompt that tells me uh create me a playlist of bands that are playing in San Francisco tonight. Uh so turns out there's a bunch of good shows if you're excited to check them out. On

5:11 the right hand side, you see a prompt that is asking for a playlist to accompany me on my run. And what's interesting with the right hand side as well is that you will see that the

5:19 experience itself gets dynamically shaped as a function of the request on the user to be able to introduce itself and different phases in my run as well. Uh, another one that I'm really excited

5:31 about, we launched it in New Zealand a couple of months ago and uh, it's coming soon in more markets is something called the taste profile. And that basically gives you the ability to introspect in

5:42 natural language what the Spotify algorithm has understood about you in a way that you can edit and refine. Uh, so if you see something that's missing or something that's wrong. Uh so for

5:52 example one of my edit is that all Disney music are my kids because we have a bunch of shared devices at home but please don't recommend that to me that's not my taste please. Uh and the

6:03 algorithm would then take that into account uh making sure that we never recommend this this in the wrong context. Uh and similarly you can also use the taste profile to share some more

6:12 aspiration uh goals as well. Uh so getting into a new genre, getting into a new topic, uh learning about a new language for example, all of these things can be can be done.

6:23 Uh another one that's coming uh soon uh which we uh announced very recently is something called personal podcast where the the the generative personalization system doesn't stop at just recommending

6:33 and ascending experiences but also generating content as well. In this particular example, I'm generating a daily brief that's uh that's generated on a cadence daily. Uh and I'm going to

6:45 uh let it uh tell me about what's happening in my community. All right. So now to go into the guts of it. So there's one big system that controls like all of these different

6:57 applications I mentioned and we call that internally some the large taste model. We're not great at naming these internal things. Um so it has a couple of properties. One is that it

7:06 understands every historical interaction piece of content on Spotify. Uh it combines prediction and reasoning to the point that I mentioned earlier. So not only guessing but also reasoning layered

7:17 on top and it gives users the ability to shape and generate experiences in real time. So it's fully steable and promptable by uh by users. Um and uh as of today about one in four US premium

7:29 subscribers interact with that system on a daily basis as well. So that's pretty exciting. Um, one thing that's exciting as well to uh is that deploying this system across existing recommendation

7:40 surfaces as well also led to some gains. Uh, we saw gains on autoplay, we saw gains on podcast discoveries. Uh, we saw gains on users interacting with DJ messages as well. So, we saw pretty

7:51 pretty sizable gains across the board as well by deploying the system. On that note, I'm going to hand it over to Jackie to talk to us about what's one of the core component that underpins this

8:01 whole system. Hi everyone, I'm Jackie or Jacqueline, a staff machine learning engineer at Spotify. So let's dive a little bit deeper and talk about how these models

8:13 are actually trained at least at Spotify. So semantic IDs were presented in the previous talk, but that is how we are embedding these openweight LLMs with knowledge of Spotify's catalog. They are

8:28 created by taking existing content embeddings such as podcast episode embeddings and applying a quantization algorithm to convert them to a set of discrete tokens. We then take a

8:43 openweight LLM such as Quen and we modify its vocabulary to add these new special tokens and then we fine-tune the model to be able to understand both natural language as well as these new

8:58 special tokens semantic IDs that represent Spotify catalog entities. So we can power experiences such as this where the user can ask in natural language for a podcast on morality. And

9:11 that is passed to the prompt along with their listening history represented as semantic IDs. And the model responds both with a relevant semantic ID podcast episode

9:24 as well as a natural language description of why they recommended that to this user. So how is this model actually trained? Um we published a paper linked here

9:39 uh describing our training paradigm called NEO which consists of four distinct stages. The first I already covered which is the semantic foundation stage where we construct meaningful

9:55 semantic ID tokens and then add them to an openway LLM's vocabulary. The second stage we call domain grounding in which we align these new semantic ID token embeddings in the

10:10 original language embedding space. We do this by learning a birectional mapping between semantic ids to text, text to semantic ids and any combination.

10:26 And we actually freeze the LLM backbone at this stage and only train the new semantic ID embeddings. So the original model weights and embeddings are frozen and we just learn those new semantic ID

10:43 tokens and this helps us to mitigate catastrophic forgetting of the pre-trained LLM's core language abilities. The third stage we call capability

10:54 induction which is multitask instruction tuning on tasks that Spotify cares about such as the ones shown here. next item recommendation retrieval etc. This is done by unfreezing the whole

11:11 model all of its weights and embeddings and running either full parameter fine-tuning or Laura fine-tuning on the multiple Spotify tasks and then there is an optional fourth stage to do post-

11:25 training such as RL fine-tuning etc. So, how much of a difference does this four-stage training paradigm actually make? I'll dive into a few of the abilations we've done to investigate

11:40 this. First, we assessed whether multitask training is actually hurting performance by comparing the multitask model against single task variance. And we consistently saw that across our

11:54 tasks, the multitask model can match or actually beat the single task performance indicating that there is some positive cross-learning happening across the tasks. This is particularly

12:09 noticeable for audiobook recommendations. If you look here, which is a new newer content type at Spotify, demonstrating that these multitask models can help with cold start entities

12:22 by learning from other items in the catalog such as podcast recommendations, how to make meaningful audiobook recommendations. Then we did some abilations on the

12:34 actual training recipe. We evaluated both dropping the frozen backbone domain grounding stage altogether as well as combining the domain grounding stage with the capability

12:48 induction stage in a multitask instruction tuning stage that those are rows A and B here in the middle and we saw for both of those that it degraded performance but actually the biggest

13:01 drop in performance was from using a randomly initialized backbone instead of the pre-trained openweight LLM that we are using. We also evaluated using continuous pre-training for the domain

13:15 grounding stage. And as you can see, it's it's minimal actual degradation on the task specific performance. But where continuous pre-training really hits us is on the natural language and

13:28 world knowledge capabilities of the pre-trained backbone LLM we are using. after we do continuous pre-training, it goes to essentially zero versus if we do the frozen backbone domain grounding. We

13:41 retain all of that core language ability and are still able to learn the semantic ids. I want to call out that these abilations were done with Quen, but we also

13:51 validated that these findings hold with Llama. So, it is not specific to the model backbone, but actually the training paradigm itself. And then lastly, we did some

14:05 investigation on different inference strategies and their effect on accuracy versus latency. We tested beam search with both constrained decoding and not. And we saw that even without constrained

14:20 decoding, we can generate valid semantic IDs 98% of the time. Constrained decoding does add a little latency overhead, but it's also helpful for specific cases where you want to

14:34 target specific types of content, such as only make new content recommendations, for example. We also compared top P sampling to beam search and saw that top P sampling

14:47 pretty significantly hurts our accuracy. So although beam search is a little more latency intensive, we decided that trade-off worked for us. So what is meaningful about this? There

15:05 have been lots of work in the industry in the space on generative semantic ID retrieval, toolbased LLM recommenders, the plum paper, etc. But NEO is actually the first example of combining all these

15:19 capabilities into one system that understands grounded catalog items, is naturally language steerable, can do search, recommendation, explanation use cases, as well as

15:34 low latency tool-free inference at an industrial scale. And we are using this in production today. Um there's a paper linked as well here for how we're using this to power

15:50 podcast discovery. What we saw is that using a model train like this, we can break users out of their habitual patterns and get them to listen to more unfamiliar content. And

16:04 we saw huge wins online with this. So none of this works without meaningful evaluations. So I want to talk about that a little bit as our recommendations are becoming more generative and

16:21 explanatory. We our original maybe traditional offline eval metrics are not sufficient. Yes, they can tell us whether or not the user interacted with that content, but

16:37 they can't tell us whether or not the user or that recommendation makes sense for the user, whether the explanation is accurate, whether it aligns with the user's intent, etc. So, that's where LLM

16:50 judges really shine. But in our work at Spotify, we really find that you need to invest in grounding your LLM judges in meaningful data. so that they can be reliable

17:04 evaluators that align with human preferences. So, a few examples um for evaluating the podcast recommendations that I just mentioned, we create textual user profiles summarizing the users's

17:20 listening history and that is passed to the LLM as a judge and we saw that this corresponded with a 75% alignment between the LLM judge and human preferences.

17:32 Similarly, you can use actual behavioral signals to ground these LLM judges. So, for example, for a search task, you can for a given query, you can take similar queries and how the user has interacted

17:47 with them in the past and pass that to the model. And we saw that overall it increased alignment by 5% but on ambiguous queries it actually increased alignment by 91%.

18:01 Showcasing the value these this grounding of the LLM judge plays especially in ambiguous cases where LLM judges tend to struggle. Lastly, we used grounded LM judges to

18:16 scale up our Cranfield style collections. These are evaluation sets that are constructed by taking candidates from multiple different sources, creating a

18:29 pool, and then using a human to rank that pool. However, that human ranking stage is expensive. So, we invested in significant grounding for our LM as judge and are able to have an LLM judge

18:44 that aligns with uh human system rankings with an agreement value of 0.87. In summary, um like we're going to hear a lot about today in all the talks,

19:01 there is a new era of personalization among us. this generative personalization. And if you want to power language steerable personalized recommendations

19:14 for your user, this is how we taught openweight LLMs to speak Spotify. Thank you. [applause] >> [music]

Frontier News · by Hyperjump Technology