Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Todd Fisher demonstrates a guitar plugin that speaks typed or spoken text using AI tools, including text-to-speech, word segmentation, pitch detection, and a local LLM for conversational responses. The plugin, built with JUCE and running in Logic Pro, slices words via energy gap and syllabifier methods, then uses a vocoder to make the guitar 'sing' by mapping pitch to synthesized notes. A more advanced version employs Whisper for speech-to-text and a local LLM to answer questions through the guitar, though the demo showed choppy results.
Key points
- The speaker built a guitar plugin that plays typed text through text-to-speech, sliced per word using energy gap segmentation and sonority peak syllabifier, though manual editing was still needed for accuracy.
- Pitch detection via the Yin algorithm and a vocoder were used to synthesize notes that follow the guitar's pitch, enabling the guitar to 'sing' by blending synthesized notes with voice clips.
- A conversational mode was implemented: the user speaks into a microphone, Whisper transcribes it, a local LLM generates a response, and the guitar plays the response with word-by-word slicing.
- For singing-like output, the speaker used the World library to pitch-shift pre-recorded vowel samples from the vocal set dataset, mapping each guitar note to a shifted sample, though this required pre-baking and was not live.
- The plugin was built using the JUCE framework and runs as a standard effect plugin in Logic Pro (a DAW), demonstrating integration with existing music production tools.
- The speaker showcased the project live, with some technical difficulties (e.g., word slicing glitches, mixer issues), but overall the system functioned as a proof-of-concept.
- The talk emphasized that AI lowers the barrier for creative side projects, encouraging the audience to build their own passion projects.
Tools mentioned
Techniques
- energy gap segmentation
- sonority peak syllabifier
- Yin pitch algorithm
- vocoder
- pitch shifting
- ADSR envelope
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
So I am Todd Fisher. I love the guitar. That's one of my passions in life. Uh today I want to talk about uh this project I've been working on for a while here, effectively making my guitar
speak. Uh but of course I want to start out with kind of framing it under this awesome premise. You know, we've all been to live performances where like your mind was just blown and it was
awesome. I you know, the the first one I remember way back I was in high school. I went to a Slipknot concert, so a little bit heavier music uh here in the Bay Area. And I remember at some point
the drummer was drumming a cool drum solo and his drum set started raising up and I was like, "Whoa, that's cool." And everyone got kind of excited, right? And then halfway through the solo, his drum
set started to go like this and tipping. And his whole drum set was like on the wall while he was upside down playing his drum solo. And everyone was just cheering like, "Whoa, this is so crazy,
mind-blowing." Like, "Whoa." It was so so impactful, right? I still remember it today. Um fast forward probably about a year ago, uh I went to New York to see the
Stranger Things on Broadway, uh the the first shadow. And it was really cool because the whole time I was there, I was watching it and it felt like it was actually like a a Netflix episode but in
real life. And it was just like the effects that they did there was so so amazing. I remember even one time at near the end there was a scene where somebody somebody was falling backwards
and it was like a slow-motion scene and it probably about a minute long the whole scene. But just like the fact that like they they were able to produce that in real life but it looked like it was a
post-production slow-motion scene. Just just mind-blowing. So it's like, "Whoa, that's cool." So given you uh everyone here has probably had some experiences uh with live performances in the past
and I think it's really awesome just to like see the creativity uh that people can do with leveraging technology in some creative way. And so, you know, I'm thinking about all
the projects that we have as engineers over the years where it's like, "Hey, that's a cool idea. I'm going to write this down." I have like a big giant list. Anyone have a list of projects
that maybe they'll get to? Yeah, so it all happens, right? Uh and it's really awesome cuz the last 6 months or so with AI uh being able to, you know, push us forward with some of these projects,
it's just really cool to see where we can take all these projects. So so so part of uh my goal today is to inspire you guys to find whatever project you're passionate about, go start building it
because it's super easy now with AI it's easier with AI. I know sometimes there's caveats, but but in general uh I I want to have you guys leave this session being inspired to go build whatever that
cool thing is for you personally that's going to, you know, help you to learn stuff or maybe make a difference in the things that you're you're trying to figure out.
Uh so I want to start off with framing this kind of in the realm of the guitar. So the guitar because the guitar guitar's been around for several years several thousands years or hundreds of
years, sorry. Um and at some point somebody said, "Hey, I'm going to put a pickup on a guitar. I'm going to plug it into a speaker and now we have rock and roll. We're going to make it super
loud." That was pretty awesome. Uh and then of course we have a lot of people getting these stomp boxes these these effects pedals. We put them all together. We have really awesome sounds.
You know, really famous people have a lot of these. Uh and then at some point Peter Peter Frampton uh came along and said, "Hey, what if we get the sound of the guitar, put it through it an actual
hose like a physical hose, put it in my mouth, and then basically play the guitar sound in my mouth, and inform some sort of words." So that's pretty awesome that that's known as the talk
box. Uh and then in the last couple decades we have a lot of progress made with software emulation. Uh you think of, you know, Pro Tools, you think of Logic uh
uh Fruity Loops. There's a lot of software even to the point where all the effects are in software now. There is the argument made hot take uh that maybe you don't need all those physical
effects pedals anymore. So so there's that question, right? And then of course looking forward it's what is that next evolution of the guitar uh with AI in the picture?
So with that said, I want to take you back several years ago uh Halloween time I was passing out candy, bunch of trick-or-treaters, you know, it got kind of boring. So I was like, "Hey, what if
I go and actually bring my guitar out with my amp and just play?" Uh so for the past what 10 or plus years or so, I've actually been playing guitar on Halloween passing out candy and it's
kind of fun. Uh and then fast forward probably 3 years ago, I decided to dress up as Eddie Munson. Anyone know Eddie Munson? Stranger Things guitar player guy.
So pretty awesome guy. So I dressed up as him and I was like, "Yeah, this is going to be so fun. I can play the music with him, whatever." And you kind of go the Oh, hold on.
Hold on, wait a minute. Technical difficulties. All right, there it goes. So a bunch of little Stranger Things awesomeness, right? Uh but figured out,
"Hey, what if there was like some more stuff I could put in this whole experience?" And so I decided to go and build a little app that would actually paint the Stranger Things alphabet on my
garage door just because it's fun, right? Uh and so let me just show you a quick example of how this worked. So effectively, make sure I have the mic permission. Whenever I play a note,
it would go and communicate whatever lights. And I even had it to where I could actually uh set custom messages such as Happy Halloween, all that fun stuff. It was it was overall pretty fun
uh just to kind of mess with that, right? So once again, you know, whatever I type in there, I could actually spell whatever and and very much uh a cool nod to Stranger Things. But it kind of got
me thinking like what is the next evolution of this project? And I settled on this idea of like, "Hey, how hard would it be to make my guitar speak?" It sounds easy, maybe, maybe not. So, so
today I want to share kind of my journey in this process and where kind of I'm at today. So, looking around the different tools, there's a framework out there called
JUCE. Really awesome for anyone building audio software, look into JUCE. It's it's pretty good. And then of course there's a a number of plugins
sorry, plugin formats out there for your your digital audio workstation. And for those that are not aware, your DAW is effectively your IDE but for musicians and and music producers. And
then I started with some text to speech stuff with Piper and some built-in Apple stuff. And then some other really fun digital signal processing.
So, with that said, my first stab here was I want to get raw text so I could just type in whatever text I want push it through the text to speech, get the audio clip, and then whenever I play a
note on the guitar, I want to go and play that back. So, let's go see how that works. So, switching over to my Logic Pro here. And this is the plugin I made. So, it's
just like any other plugin in Logic where you just plop it in there. It's just, you know, chaining all the effects together. And so, this is what I came up with.
>> Developers. >> So, it's playing >> Developers. >> Pretty awesome, right? >> Developers.
>> Yeah, kind of kind of reminds me of something, right? >> Developers. >> Start clapping. Everyone, start clapping. Developers. Developers.
Developers. Developers. Developers. Awesome, thank you. That's pretty awesome. You guys are great. Um So, I got it to where it's playing an actual audio file. That's pretty
awesome. But it turns out, you know, in English or in any languages for that matter, there's more than one word. So, the next kind of evolution Oh, that's a little bit chatty. The next evolution
there is let's actually go and slice it per word now. So, I got it to the point where it's playing. Now it's going to slice per word.
>> Look at me. I can speak. Look >> Oh. >> at >> Oh.
>> me. I can speak. >> So, now it's speaking words, and that's pretty awesome, right? Uh but it turns out that uh
as we speak, there's a number of challenges in how we automatically slice words. And so, I looked into this thing called energy gap segmentation. The general idea here is if you look in a
any waveform over here, we see that here's a bunch of of words that we're speaking, right? Uh the idea there is there's typically silence in between words. So, let's just cut it whenever
the decibels are are very much close to zero, right? Uh but the issue with that is there's actually sometimes when you know as I'm speaking right now, for example, uh there's actually no silence
in between some of my words. So, it gets a little bit challenging to where it's not 100% foolproof, right? So, beyond that, I looked into this thing called sonority peak syllaba
syllabifier. That's a hard word to say. I effectively identifying the syllables of the audio signal and identifying that there's vowels in here. Vowels typically lead to syllables.
That's kind of the idea. So, I said, "Okay, let's take the sonority peak, add it to the the energy gap, and figure out if we could just make that work automatically."
So, with that said, uh made it kind of work. So, let's just play this one. >> Thank you for letting me be >> Nope. Oh.
>> here. It feels so good
>> Oh. >> to get out of my big tiny
cage once in a while. >> So, there you go. It's working pretty
well. Not quite as good as I want to. So, long story short, uh I settled on just the ability to go and actually I could drag this and manually edit some of these
uh uh segments in here. So, worked okay, right? Um but moving on, you know, we with the spirit of evolving the thought, evolving
the project, right? It's like, "Okay, we're we're having the AI say stuff, that's great. But what if we could actually make it sing?" Let's take it to the next step cuz this is music, why
not, right? Uh so, looked into pitch detection. And for those that are not aware, uh as you hear any any noise out there, there's typically multiple frequencies going on at any given time.
You know, you think of when you play the the C key on the piano, uh there is definitely a fundamental frequency or the the kind of the the one that we identify as the note, but there's a
bunch of other frequencies. So, we needed a way to go and figure out like, "How do we detect that fundamental frequency? So, when I press something on the guitar, how do I translate that into
an actual note?" Uh and then I found this uh Yin pitch algorithm. Basically, what it does is it detects the pitch. Uh I won't get into all the details, but look it up, it's kind of a fun uh really
cool way of detecting the pitch. Uh but effectively, what I do is I play the guitar, uh I detect the the pitch, I I make what's effectively a pitch sawtooth or aka a synthesized note. Uh so, if for
those that are not familiar with how audio works on the the computer, you think of all the electronic music, all that stuff, uh that is basically uh a synthesized note. We have ADSR, which
effectively are the levers to figure out how to actually make the note sound in different ways. Uh and then so, basically, we get the pitch notes or the synthesized note. Uh we then push it
through uh the the voice clip effectively. So, think of the talk box. Uh we're kind of filling up the cavity of the talk of the voice, that is. And we push it through a vocoder, and it
should sing. So, that's kind of the idea, right? So, with that said, let's go ahead and jam out a little bit because I have a guitar, and it's fun. So.
Let me get back to Logic here. Uh so, I'm going to play some chords. Uh was going to uh play a song that is very much related to the talk or the the title of my talk, uh while my guitar
gently speaks, but because it's going to be posted online, I don't want to muddy up the waters with any copyright things. So, I will play some chords that may or may not sound similar to a famous song.
So, So, that's the backing track. So, let's go ahead and have some AI speaking on top of it.
>> I'm here. This >> No. All right, technical difficulties. Let's
try that again. Live demos always the best. >> I'm here. This can't
Thanks. >> Uh technical difficulties again, so sorry for that, but let's just run through it and see what happens. >> can't
Thanks. That's peak. I'm the guitar that can't speak. So. >> Awesome. So, there you go. So, some bugs to work out, but overall we are able to
to say words on top of this and for what it's worth I'm trying to mix the uh the the synthesized note uh with this clarity lever right here. So, balance or mixing the the synthesized note with the
actual voice that the AI is giving. Uh so so, that worked pretty good. Um But, there's also this other question, you know, given that uh when you speak it's typically
conversational, what if I could take this to the next level? And what if we actually had this microphone right here where I could speak into the mic, it would then respond on the guitar. Uh so,
the way that I I accomplished this is speak in the mic, use uh speech to text so, Whisper, uh put that into raw text, run a local model on my computer. Uh it doesn't really matter what LLM, just any
local model. I have a conversation coming and then from that output go and plop it on the guitar. So, let's see how this works. Anyone have a question you want to ask
my guitar? Anybody? >> What is reality? >> Let's do it. What is reality?
So, it's thinking, going through the LLM things. And let's see what it says. >> That's quite a question to
start with. How does your music help you understand and that elusive concept of reality?
What do you write in that? >> So, yeah, very existential like that. So, thank you for that suggestion there. So, so there you go. We have a choppy
version, but but it is working. So, so that's a win, right? Uh Awesome, and thanks for the clap, sir. Um But really like it's not quite singing
yet, right? So, moving on in this project is like how do I actually make it sing, right? Uh I found a lot of open source options out there as far as recordings or samples, if you will. Uh
there's a thing called vocal set out there. They had basically recorded a bunch of singers, and I could use those audio files to go and do some fun stuff with them. And effectively I took those
samples, I used a project called World, which helps with things like pitch shifting and some other things, and I'm able to then shift the the pitches of everyone singing on those audio clips
and map it to the guitar. Uh this is a very heavy process, and so I can't really do that live. I have to effectively pre-bake that. And then once it's already pre-baked, I could then go
and jam out with it. So, with that said, uh let me go ahead and show that example here. Um. So, uh because it takes so long, I
actually just started with the vowel sounds. So, there's somebody singing all each of the five vowels. And so, we'll see how that sounds, right? Kind of fun, kind of weird, but overall
it's working. It's it's closer to singing, right? Uh So, let's go ahead and throw that on top of all the chords over here. And let's see if we can make it sound
decently well. Uh So, we'll just play these chords again. So, not quite your opera singer, but getting closer and and so I I think
there's something awesome there. Uh so, effectively once again, that is uh going through uh the whole synthesized process, uh putting a sample, shifting the the pitch shift of the sample, and
then effectively mapping each fret or each note on the guitar to one of those samples that are pre-baked effectively, right? Uh uh So, with that said,
you know, where do I want to take this next? Uh very much want to get into more of of the, you know, AI heavier options. Uh if anyone has any other ideas of how to make the guitar sing, feel free to
come up to me after. I think it's pretty awesome. Uh but more importantly, you know, everyone here, uh going back to the the charge to go and and build some awesome, you know, side projects, some
passion projects of yours, like what is that for you? Go and build awesome things because nowadays with AI, we could build so many really cool things, and time is typically not the the big
time suck that that it once was, right? So, with that said, uh be awesome uh or or be good to each other, and thank you very much. >> Woo!