Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Google's Gemini 3.8 Flash TTS models finally ship voice cloning, but with heavy guardrails: a spoken consent clip must match the reference speaker, and generation is watermarked with SynthID and C2PA. The headline feature is voice design—describing a voice in plain language across 100+ languages—but the benchmark claims are muddied because the lead author was formerly CEO of Hume AI, which runs the benchmark they cite. Independent tests from Artificial Analysis put Flash TTS second and Flash Light sixth in preference ELO, and the pricing sits mid-pack, not cheap enough to ignore the alternatives.
Key points
Two models: Flash TTS for creative use and Flash Light TTS for bulk generation.
Voice cloning requires a spoken consent clip matching the reference speaker.
Benchmark claims from Hume AI are questioned because the blog post author was formerly Hume's CEO.
Independent Artificial Analysis benchmarks rank Flash TTS second and Flash Light sixth in preference ELO.
Voice design lets users describe a voice in plain language across over 100 languages.
Tools mentioned
Techniques
- voice design
- voice cloning
- stage directions per line
- inline tags for nonverbal sounds
- multi-speaker scenes from a single script
- consent-based voice cloning
- watermarking with SynthID and C2PA
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Okay, so I think it's fair to say that Google has kind of dropped the ball on LLMs lately. The proposed Gemini 3.5 Pro never came and hopefully all of that's going to change very soon with the
Gemini 4 release. But on the multimodal side, it's a very different story. The Deep Mind multimodal team have been doing really well with the multimodal releases that have been based on the
Gemini Flash models. So we've had Omni 1.1 Flash update on the voice side. We've seen the live models with things like live translate and subscribe. And now there are two new texttospech
models, Gemini 3.8 flash TTS and Gemini 3.8 Flash Light TTS. Now Google's pitch is that these are the most expressive voice models yet. And there are number one on a bunch of the benchmarks. And
certainly some of that holds up, but when you look at some of the independent numbers, it's perhaps a bit more mixed than the blog post makes it sound. Okay, so first off, there are two models.
They're aimed at different jobs. FlashTS is the creative one. This is for things like character voices, audio books, podcasts, games, all those kinds of things. Flashlight TTS, on the other
hand, is the volume one. It's built for dubbing, bulk audio generation, voice agents, and anything where you're generating a lot of speech, and cost matters more than getting every single
nuance perfect in there. So, the headline feature here is all about voice design. Instead of just picking from a list of preset voices, you describe the voice that you want in plain language,
you give it a role, an accent, the characteristics of the voice, and it builds that voice for you across over a 100 different languages. And this is something that we've seen both from
other TTS providers, but also from open models, but certainly does seem that Google's improved the whole concept of this here. This is one of the things that I found often to be the most
frustrating with TTS models is that while things have been getting fantastically better and we're starting to see really good proity in the models, we're starting to see really good human
likeness, etc. often the challenge becomes is that when you go through the test voices that they've got, none of them really fit what you're trying to do. And so you end up having to try and
do some kind of voice cloning. And obviously the challenges there can be a lot of legal issues. So, we have seen some really nice attempts at this from the Quent TTS people, but some of the
more recent ones that I've looked at, no matter how hard I try to describe the voice, it ends up going in a slightly different direction. In fact, what I've often ended up doing is basically
generating a lot of different voices with slightly different descriptions, going through and curating the best one, and then actually cloning that artificially made voice. Now, alongside
voice design in here, there's a library of over 2,000 different voices in here. They've got a whole bunch of different regional accents, and you can actually sort of take a voice, manipulate it the
way you want, and then actually save that voice so that you can use it going forward. And I guess this is really no surprise. Google's had this feature in testing for quite a while, but they just
haven't rolled it out. And that's the whole feature of voice cloning. So, you give it a 30-second sample and it creates that voice. Now, obviously, compared to other startups, when Google
does something like this, they get in way more trouble if it goes wrong. So, they clearly have taken their time with doing this and they've really focused on putting some guard rails on this. So,
the person whose voice it's going to actually create has to go through a spoken cassend clip. And that clip has to match the reference speaker before the clone ever gets created. On top of
this, of course, they've also got the whole sort of safety features in there that every clip these models generate is watermarked with synth ID and there are C2PA credentials on top of that. The
other challenge that a lot of people are going to face is that because of legal requirements that voice cloning in AI Studio isn't available in a lot of different countries and even in some
states in the United States. So, it does show just how touchy of a subject that can be, especially for a big company like Google. Now once you've actually got your voice, the other big thing you
can do here is you can give stage directions line by line. So you can say that this one should be whispered, this one should be rushed, this one should be sarcastic, etc. And you can drop in sort
of nonverbal tags like I've shown you in some of the other TTS systems for things like sigh, laughs, gasps, all those kinds of things. Another nice feature that you've got in here is you can
actually do two speaker scenes from a single script. And Google claims with that that it can actually hold the voices for hours of long form audio without the voice sort of actually
drifting. That's one of the issues that I've run into with some of the open TTS systems recently is that once you get that voice right and you start generating anything of any length, it
starts to end up drifting off. All right, so let's talk a little bit about the benchmarks because this is where it gets interesting but also a bit questionable in some of the things that
they've done. So, Google's numbers mostly come from Hume AI. They talk about that the Flash TTS is number one on Hume's voice design benchmark and it leads on accents. And you can see people
like Logan actually tweeting this out. But here's the thing. One of the two people that wrote the blog post is Alan Cohen. And Allan actually was the CEO and chief scientist at Hume until
January this year when Google DeepMind licensed their tech and hired a bunch of the people from that startup over. So I'm certainly not saying that their numbers are wrong, but I'm not sure if
quoting benchmark numbers from a benchmark that you designed yourself is the best way to judge this. Now the good thing here is that artificial analysis have also done their own tests on that.
And I think those are reasonably independent. If we look at the sort of preference ELO, which is the common way that people actually test these kind of models, you can see that the Gemini 3.8
Flash TTS is coming second in here and the Flash Light TTS is coming in sixth in here. And can we see that the flash model is just ahead of the new Quen Audio 3 TTS Plus? Now, at the time of me
recording this, that one is still not an open model, but if Quen goes on to release the weights for that as an open model, like they did with the Quinton TTS before, that's going to be really
interesting to see. As well, where the Gemini models really do seem to take the lead are some of the benchmarks around pronunciation, etc. That's especially useful if you're using the flashlight
model to bulk generate audio of where you're saying things like numbers, IDs, things that people really want to be able to hear correct the first time. That's something that really does matter
if you're building a voice agent with this kind of model. If we look at how the models actually do on other languages, we can see this is and I would argue that this is Google's
strength, right? Google has a whole bunch of different multilingual data and you can see this clearly in here of looking at how they compare to other providers. That said, Cartisia and 11
Labs here are really making a strong move at the multilingual TTS systems. Finally, the other big demo here is actually price. So, Google's kind of sitting in the middle of the pack here.
They're obviously a lot better than some of the offerings from Cartisia and 11 Labs and certainly from Miniax, etc., but they're by far from being the cheapest in here. And I think in some
ways more importantly, the flashlight model isn't hugely cheaper than the flash model itself, which is definitely a difference than when we compare the text versions of these models. All
right, so let's go out to AI Studio, have a play with this, and then we'll try out the two models, and then we can try all this out in Collab as well. All right, so if you want to check out the
voices, you can actually come into AI Studio and just pick one of the voice models. You can see here I've picked the Gemini 3.8 8 flash TTS and then you'll be asked to basically hook up an API key
to it. But from there you can basically go through and you can explore a whole bunch of different models in here. So if I just come in here I can basically take the model and then I can make a
description in here. And the cool thing is if you want to they've actually got voice rewriting in here. So you can see that if I click on that that will basically just use Gemini to flesh out
the actual voice prompt in here of what I want the voice to be like. And then you'll see from there it will give you three voice options that you can use. Everything from describing the person
themselves through to the position that they are on a microphone and that kind of thing. So that's really kind of cool. Unfortunately, as I'm recording it now, it seems like the voices are failing in
here. But normally what you would do is you would get basically three voices and those voices then you can customize it. You can then refine it, etc. All right. After multiple attempts to try and
generate this, it doesn't seem like this part is working. Now, let's just jump into the code and have a play with the actual collab for this. All right, so let's take a listen to it.
>> Hi, I'm Sam and today we're looking at the new Gemini 3.8 texttospech models. >> So, as you can see that the basic voice just works pretty well. And the cool thing is you've got lots of choices if
you want to do just sort of standard voices and stuff in here. Let's come down and listen to the comparison between the flash model and the flashlight model.
>> The quarterly numbers are in. Revenue is up 11% and churn is the lowest it has been in 2 years. >> Okay, so that was the normal flash TTS. Now let's listen to the flashlight.
>> The quarterly numbers are in. Revenue is up 11%. And churn is the lowest it has been in 2 years. Okay, so the common thing I notice with the flashlight TTS is it
tends to be a bit more lazy. It might go on for longer, etc. And honestly, I feel that the price difference is not huge. So, you've got the flash TTS being 50% more that if I was doing something that
wasn't too long, I'm probably just going to go for the good voices. Okay, so you can see now we've got the style control. This is basically where we can give it direction on how we want what we're
asking it to speak, how we want it pronounced, etc. >> We just got the results back and the prototype actually works. >> Okay, so that one certainly is flat and
bored. Let's try thrilled and breathless. >> We just got the results back and the prototype actually works. >> Okay, and they're now nervous, almost
whispering. We just got the results back and the prototype actually works. >> Okay, so you can hear that all three of them definitely sound very different even though they're the exact same
words. So, one of the things that I do like about this is that this affects the length of the audio. So with some of the open models, one of the challenges that you have is it will try and say
something in the style that you're asking, but it's also trying to hit a certain time length in the way that it returns back. Here we're basically letting the time length be dictated by
the style that it's saying. All right. The other thing that you can do in here is that in the actual script that you want them to say, you can actually put in these inline tags. And so these are
things like whisper, laugh, sigh, that kind of thing. >> Wait, shh, did you hear that in the hallway? Never mind. It was just the cat. You should have seen your face.
>> Okay, so you could hear there that the inline tags, it certainly does what they're basically there for. In the case of whisper, sigh, laugh, the laughs will change also depending on the style
directions you've given it as well. All right, so next up is the whole idea of voice design. So now what we're doing is we're not just passing in a script and sort of a guidance of how to speak it.
We're actually giving a description of the voice itself. So the idea here is that this voice should then carry over to everything that it's going to speak even if we give it other sort of
descriptions of how to speak it going through. So, this is where you can put in things like where the person comes from, the accent that you want, if you want to define something about age, that
kind of thing in here. So, you can see in here I've got a character called the astronomer that's in his late 60s with a gentle British accent. Look out past the rings of Saturn. That
faint light left its source long before humans walked the Earth. Okay. So, you can see there it's got some vibes of sort of a David Atenburgh voice there, but without actually
mentioning him. Personally, I feel the voice sounds a lot younger than late60s, but let me know in the comments what you think if you try that out. All right, next up is the fun one where you can
basically have multipe dialogues going on in here. So, here we're doing one API call with two voices. And so we've got this idea of per turn style, meaning that we can basically set what each
character says in here. And you can see we can define the actual speakers, who they are, but we can also define the style for each line of dialogue that they've actually got. So in order to
play the whole thing here, I've put the notebook up. You can have a play with this yourself and try it out and just see how it goes. But let's take a little listen to this.
>> Welcome back to the show. Today we're testing the new Gemini voices. >> Thanks for having me. The turn taking is a lot more natural than the last version.
>> So is it good enough for real-time agents? >> For a lot of use cases, I think it is >> a lot of use cases. That's doing okay. You can hear there. This is definitely
got the vibes of sort of a notebook LM podcast kind of thing that gets generated. Now for me, I don't know if it's just that I'm so used to hearing these voices now that they kind of sound
fake to me. But the cool thing is you can swap out other voices. You could make your own etc as you go through this. All right. Lastly, we've got the voice cloning in here. And the voice
cloning can be done just by code. Right. So you can see the code here is basically just making an HTML interface where you can record the reference audio. So which is going to be what you
want the audio to sound like. And then second where you can record the consent audio. I do think that Google's in a very hard spot here. We know that this is something that their TTS systems have
been able to do for quite a number of years now. And this goes back to a paper called Soundstorm which at the time was absolutely amazing. This is more than 3 years ago now. And even then they were
able to basically clone voices so well. And at the time I was actually given a demo of this and it was amazing just how well it could clone someone's voice back then. Jump forward to today. Many of the
open models now have got this worked out. And yet, this is the first time that we've seen Google actually make this available to people. And I think it's been very clear that Google's very
conscious of the ramifications of this and that people can use this kind of tech for doing a lot of nefarious kind of things. But I got to say in some ways I wonder whether Google's version here
has actually been dumbed down a little bit to be not quite as good as some of the open versions that are out there. So basically what you do here with this HTML interface is that you can record in
your reference audio. So this is the audio to replicate and then you can record in the consent audio. And this is the really important bit where you're basically giving Google consent to make
a synthetic voice based on your voice. And if we come down and actually listen to these first, I'll play you the original bit that I recorded in here. The cosmos is within us. We are made of
star stuff. We are a way for the cosmos to know itself. Okay, so that was the sort of original recording I did. And now let's listen to the clone of my voice done with this. This was generated
from a 30- secondond sample of my voice. Now, I'll leave it to you guys to decide. For me, it sounds quite different. I think possibly had I uploaded a longer reference audio, it
might be able to get a better version of my voice. Probably the best model that I've seen doing this openwise was the Quen TTS. I made a whole video about that when it came out, and I still think
that one is really amazing at how well it can clone voices. And in some way, I would say that was their strength, at least back then. And yet, they were probably weaker with the descriptions
for creating a voice. And that's the Gemini's strength in here now. So, just to finish up, I'd say this is definitely worth checking out. It is really cool how easy it is nowadays to basically do
voice things. Things that just used to be so difficult in the past are now a very simple API call. And I think for people who perhaps don't have hardware to run their own TTS system to run
something locally, this is going to be quite nice for just giving you an API that you can ping and get highquality voices out with really nice proid and with a custom style that you can define
all by yourself. So anyway, let me know in the comments. I know there are TTS models coming out all the time. I try to keep up with them. I've missed some. If you've got any recommendations that you
would like me to cover, please put them in the comments below. I'd love to hear that from everyone. Anyway, as always, if you found the video useful, please click like and subscribe, and I will
talk to you in the next video. Bye for now.