Gemini 3.8 Flash TTS with Voice Cloning

summarized

TLDR

Google's Gemini 3.8 Flash TTS models finally ship voice cloning, but with heavy guardrails: a spoken consent clip must match the reference speaker, and generation is watermarked with SynthID and C2PA. The headline feature is voice design—describing a voice in plain language across 100+ languages—but the benchmark claims are muddied because the lead author was formerly CEO of Hume AI, which runs the benchmark they cite. Independent tests from Artificial Analysis put Flash TTS second and Flash Light sixth in preference ELO, and the pricing sits mid-pack, not cheap enough to ignore the alternatives.

Key points

Two models: Flash TTS for creative use and Flash Light TTS for bulk generation.

Voice cloning requires a spoken consent clip matching the reference speaker.

Benchmark claims from Hume AI are questioned because the blog post author was formerly Hume's CEO.

Independent Artificial Analysis benchmarks rank Flash TTS second and Flash Light sixth in preference ELO.

Voice design lets users describe a voice in plain language across over 100 languages.

Tools mentioned

Techniques

  • voice design
  • voice cloning
  • stage directions per line
  • inline tags for nonverbal sounds
  • multi-speaker scenes from a single script
  • consent-based voice cloning
  • watermarking with SynthID and C2PA
Transcript (captions)

0:00 Okay, so I think it's fair to say that Google has kind of dropped the ball on LLMs lately. The proposed Gemini 3.5 Pro never came and hopefully all of that's going to change very soon with the

0:11 Gemini 4 release. But on the multimodal side, it's a very different story. The Deep Mind multimodal team have been doing really well with the multimodal releases that have been based on the

0:22 Gemini Flash models. So we've had Omni 1.1 Flash update on the voice side. We've seen the live models with things like live translate and subscribe. And now there are two new texttospech

0:34 models, Gemini 3.8 flash TTS and Gemini 3.8 Flash Light TTS. Now Google's pitch is that these are the most expressive voice models yet. And there are number one on a bunch of the benchmarks. And

0:47 certainly some of that holds up, but when you look at some of the independent numbers, it's perhaps a bit more mixed than the blog post makes it sound. Okay, so first off, there are two models.

0:56 They're aimed at different jobs. FlashTS is the creative one. This is for things like character voices, audio books, podcasts, games, all those kinds of things. Flashlight TTS, on the other

1:09 hand, is the volume one. It's built for dubbing, bulk audio generation, voice agents, and anything where you're generating a lot of speech, and cost matters more than getting every single

1:20 nuance perfect in there. So, the headline feature here is all about voice design. Instead of just picking from a list of preset voices, you describe the voice that you want in plain language,

1:32 you give it a role, an accent, the characteristics of the voice, and it builds that voice for you across over a 100 different languages. And this is something that we've seen both from

1:42 other TTS providers, but also from open models, but certainly does seem that Google's improved the whole concept of this here. This is one of the things that I found often to be the most

1:52 frustrating with TTS models is that while things have been getting fantastically better and we're starting to see really good proity in the models, we're starting to see really good human

2:03 likeness, etc. often the challenge becomes is that when you go through the test voices that they've got, none of them really fit what you're trying to do. And so you end up having to try and

2:13 do some kind of voice cloning. And obviously the challenges there can be a lot of legal issues. So, we have seen some really nice attempts at this from the Quent TTS people, but some of the

2:22 more recent ones that I've looked at, no matter how hard I try to describe the voice, it ends up going in a slightly different direction. In fact, what I've often ended up doing is basically

2:34 generating a lot of different voices with slightly different descriptions, going through and curating the best one, and then actually cloning that artificially made voice. Now, alongside

2:45 voice design in here, there's a library of over 2,000 different voices in here. They've got a whole bunch of different regional accents, and you can actually sort of take a voice, manipulate it the

2:54 way you want, and then actually save that voice so that you can use it going forward. And I guess this is really no surprise. Google's had this feature in testing for quite a while, but they just

3:04 haven't rolled it out. And that's the whole feature of voice cloning. So, you give it a 30-second sample and it creates that voice. Now, obviously, compared to other startups, when Google

3:14 does something like this, they get in way more trouble if it goes wrong. So, they clearly have taken their time with doing this and they've really focused on putting some guard rails on this. So,

3:23 the person whose voice it's going to actually create has to go through a spoken cassend clip. And that clip has to match the reference speaker before the clone ever gets created. On top of

3:33 this, of course, they've also got the whole sort of safety features in there that every clip these models generate is watermarked with synth ID and there are C2PA credentials on top of that. The

3:44 other challenge that a lot of people are going to face is that because of legal requirements that voice cloning in AI Studio isn't available in a lot of different countries and even in some

3:55 states in the United States. So, it does show just how touchy of a subject that can be, especially for a big company like Google. Now once you've actually got your voice, the other big thing you

4:04 can do here is you can give stage directions line by line. So you can say that this one should be whispered, this one should be rushed, this one should be sarcastic, etc. And you can drop in sort

4:14 of nonverbal tags like I've shown you in some of the other TTS systems for things like sigh, laughs, gasps, all those kinds of things. Another nice feature that you've got in here is you can

4:26 actually do two speaker scenes from a single script. And Google claims with that that it can actually hold the voices for hours of long form audio without the voice sort of actually

4:37 drifting. That's one of the issues that I've run into with some of the open TTS systems recently is that once you get that voice right and you start generating anything of any length, it

4:47 starts to end up drifting off. All right, so let's talk a little bit about the benchmarks because this is where it gets interesting but also a bit questionable in some of the things that

4:55 they've done. So, Google's numbers mostly come from Hume AI. They talk about that the Flash TTS is number one on Hume's voice design benchmark and it leads on accents. And you can see people

5:06 like Logan actually tweeting this out. But here's the thing. One of the two people that wrote the blog post is Alan Cohen. And Allan actually was the CEO and chief scientist at Hume until

5:18 January this year when Google DeepMind licensed their tech and hired a bunch of the people from that startup over. So I'm certainly not saying that their numbers are wrong, but I'm not sure if

5:27 quoting benchmark numbers from a benchmark that you designed yourself is the best way to judge this. Now the good thing here is that artificial analysis have also done their own tests on that.

5:37 And I think those are reasonably independent. If we look at the sort of preference ELO, which is the common way that people actually test these kind of models, you can see that the Gemini 3.8

5:47 Flash TTS is coming second in here and the Flash Light TTS is coming in sixth in here. And can we see that the flash model is just ahead of the new Quen Audio 3 TTS Plus? Now, at the time of me

6:01 recording this, that one is still not an open model, but if Quen goes on to release the weights for that as an open model, like they did with the Quinton TTS before, that's going to be really

6:09 interesting to see. As well, where the Gemini models really do seem to take the lead are some of the benchmarks around pronunciation, etc. That's especially useful if you're using the flashlight

6:19 model to bulk generate audio of where you're saying things like numbers, IDs, things that people really want to be able to hear correct the first time. That's something that really does matter

6:29 if you're building a voice agent with this kind of model. If we look at how the models actually do on other languages, we can see this is and I would argue that this is Google's

6:38 strength, right? Google has a whole bunch of different multilingual data and you can see this clearly in here of looking at how they compare to other providers. That said, Cartisia and 11

6:49 Labs here are really making a strong move at the multilingual TTS systems. Finally, the other big demo here is actually price. So, Google's kind of sitting in the middle of the pack here.

7:00 They're obviously a lot better than some of the offerings from Cartisia and 11 Labs and certainly from Miniax, etc., but they're by far from being the cheapest in here. And I think in some

7:10 ways more importantly, the flashlight model isn't hugely cheaper than the flash model itself, which is definitely a difference than when we compare the text versions of these models. All

7:19 right, so let's go out to AI Studio, have a play with this, and then we'll try out the two models, and then we can try all this out in Collab as well. All right, so if you want to check out the

7:28 voices, you can actually come into AI Studio and just pick one of the voice models. You can see here I've picked the Gemini 3.8 8 flash TTS and then you'll be asked to basically hook up an API key

7:40 to it. But from there you can basically go through and you can explore a whole bunch of different models in here. So if I just come in here I can basically take the model and then I can make a

7:50 description in here. And the cool thing is if you want to they've actually got voice rewriting in here. So you can see that if I click on that that will basically just use Gemini to flesh out

8:00 the actual voice prompt in here of what I want the voice to be like. And then you'll see from there it will give you three voice options that you can use. Everything from describing the person

8:12 themselves through to the position that they are on a microphone and that kind of thing. So that's really kind of cool. Unfortunately, as I'm recording it now, it seems like the voices are failing in

8:23 here. But normally what you would do is you would get basically three voices and those voices then you can customize it. You can then refine it, etc. All right. After multiple attempts to try and

8:34 generate this, it doesn't seem like this part is working. Now, let's just jump into the code and have a play with the actual collab for this. All right, so let's take a listen to it.

8:42 >> Hi, I'm Sam and today we're looking at the new Gemini 3.8 texttospech models. >> So, as you can see that the basic voice just works pretty well. And the cool thing is you've got lots of choices if

8:54 you want to do just sort of standard voices and stuff in here. Let's come down and listen to the comparison between the flash model and the flashlight model.

9:04 >> The quarterly numbers are in. Revenue is up 11% and churn is the lowest it has been in 2 years. >> Okay, so that was the normal flash TTS. Now let's listen to the flashlight.

9:15 >> The quarterly numbers are in. Revenue is up 11%. And churn is the lowest it has been in 2 years. Okay, so the common thing I notice with the flashlight TTS is it

9:29 tends to be a bit more lazy. It might go on for longer, etc. And honestly, I feel that the price difference is not huge. So, you've got the flash TTS being 50% more that if I was doing something that

9:41 wasn't too long, I'm probably just going to go for the good voices. Okay, so you can see now we've got the style control. This is basically where we can give it direction on how we want what we're

9:53 asking it to speak, how we want it pronounced, etc. >> We just got the results back and the prototype actually works. >> Okay, so that one certainly is flat and

10:04 bored. Let's try thrilled and breathless. >> We just got the results back and the prototype actually works. >> Okay, and they're now nervous, almost

10:13 whispering. We just got the results back and the prototype actually works. >> Okay, so you can hear that all three of them definitely sound very different even though they're the exact same

10:24 words. So, one of the things that I do like about this is that this affects the length of the audio. So with some of the open models, one of the challenges that you have is it will try and say

10:37 something in the style that you're asking, but it's also trying to hit a certain time length in the way that it returns back. Here we're basically letting the time length be dictated by

10:48 the style that it's saying. All right. The other thing that you can do in here is that in the actual script that you want them to say, you can actually put in these inline tags. And so these are

10:58 things like whisper, laugh, sigh, that kind of thing. >> Wait, shh, did you hear that in the hallway? Never mind. It was just the cat. You should have seen your face.

11:08 >> Okay, so you could hear there that the inline tags, it certainly does what they're basically there for. In the case of whisper, sigh, laugh, the laughs will change also depending on the style

11:20 directions you've given it as well. All right, so next up is the whole idea of voice design. So now what we're doing is we're not just passing in a script and sort of a guidance of how to speak it.

11:33 We're actually giving a description of the voice itself. So the idea here is that this voice should then carry over to everything that it's going to speak even if we give it other sort of

11:46 descriptions of how to speak it going through. So, this is where you can put in things like where the person comes from, the accent that you want, if you want to define something about age, that

11:56 kind of thing in here. So, you can see in here I've got a character called the astronomer that's in his late 60s with a gentle British accent. Look out past the rings of Saturn. That

12:09 faint light left its source long before humans walked the Earth. Okay. So, you can see there it's got some vibes of sort of a David Atenburgh voice there, but without actually

12:22 mentioning him. Personally, I feel the voice sounds a lot younger than late60s, but let me know in the comments what you think if you try that out. All right, next up is the fun one where you can

12:32 basically have multipe dialogues going on in here. So, here we're doing one API call with two voices. And so we've got this idea of per turn style, meaning that we can basically set what each

12:46 character says in here. And you can see we can define the actual speakers, who they are, but we can also define the style for each line of dialogue that they've actually got. So in order to

12:55 play the whole thing here, I've put the notebook up. You can have a play with this yourself and try it out and just see how it goes. But let's take a little listen to this.

13:05 >> Welcome back to the show. Today we're testing the new Gemini voices. >> Thanks for having me. The turn taking is a lot more natural than the last version.

13:13 >> So is it good enough for real-time agents? >> For a lot of use cases, I think it is >> a lot of use cases. That's doing okay. You can hear there. This is definitely

13:24 got the vibes of sort of a notebook LM podcast kind of thing that gets generated. Now for me, I don't know if it's just that I'm so used to hearing these voices now that they kind of sound

13:35 fake to me. But the cool thing is you can swap out other voices. You could make your own etc as you go through this. All right. Lastly, we've got the voice cloning in here. And the voice

13:44 cloning can be done just by code. Right. So you can see the code here is basically just making an HTML interface where you can record the reference audio. So which is going to be what you

13:57 want the audio to sound like. And then second where you can record the consent audio. I do think that Google's in a very hard spot here. We know that this is something that their TTS systems have

14:08 been able to do for quite a number of years now. And this goes back to a paper called Soundstorm which at the time was absolutely amazing. This is more than 3 years ago now. And even then they were

14:20 able to basically clone voices so well. And at the time I was actually given a demo of this and it was amazing just how well it could clone someone's voice back then. Jump forward to today. Many of the

14:32 open models now have got this worked out. And yet, this is the first time that we've seen Google actually make this available to people. And I think it's been very clear that Google's very

14:42 conscious of the ramifications of this and that people can use this kind of tech for doing a lot of nefarious kind of things. But I got to say in some ways I wonder whether Google's version here

14:53 has actually been dumbed down a little bit to be not quite as good as some of the open versions that are out there. So basically what you do here with this HTML interface is that you can record in

15:06 your reference audio. So this is the audio to replicate and then you can record in the consent audio. And this is the really important bit where you're basically giving Google consent to make

15:18 a synthetic voice based on your voice. And if we come down and actually listen to these first, I'll play you the original bit that I recorded in here. The cosmos is within us. We are made of

15:29 star stuff. We are a way for the cosmos to know itself. Okay, so that was the sort of original recording I did. And now let's listen to the clone of my voice done with this. This was generated

15:41 from a 30- secondond sample of my voice. Now, I'll leave it to you guys to decide. For me, it sounds quite different. I think possibly had I uploaded a longer reference audio, it

15:52 might be able to get a better version of my voice. Probably the best model that I've seen doing this openwise was the Quen TTS. I made a whole video about that when it came out, and I still think

16:02 that one is really amazing at how well it can clone voices. And in some way, I would say that was their strength, at least back then. And yet, they were probably weaker with the descriptions

16:13 for creating a voice. And that's the Gemini's strength in here now. So, just to finish up, I'd say this is definitely worth checking out. It is really cool how easy it is nowadays to basically do

16:24 voice things. Things that just used to be so difficult in the past are now a very simple API call. And I think for people who perhaps don't have hardware to run their own TTS system to run

16:37 something locally, this is going to be quite nice for just giving you an API that you can ping and get highquality voices out with really nice proid and with a custom style that you can define

16:49 all by yourself. So anyway, let me know in the comments. I know there are TTS models coming out all the time. I try to keep up with them. I've missed some. If you've got any recommendations that you

16:59 would like me to cover, please put them in the comments below. I'd love to hear that from everyone. Anyway, as always, if you found the video useful, please click like and subscribe, and I will

17:08 talk to you in the next video. Bye for now.

Frontier News · by Hyperjump Technology