Claude Can Actually Watch and Edit Videos. Here's How.

summarized

TLDR

Claude Fable 5, combined with Higgsfield's MCP connector and Frame.io, can watch a video transcript, flag missing or mismatched B-roll, generate replacement clips, and attach them back to the timeline, turning a manual review loop into an automated production pipeline. The system uses a four-step workflow—watch, marks, render, drop back—and a custom Claude skill to train the model on client-specific taste for emotional visual matching.

Key points

  • Claude Fable 5 reads video transcripts line by line against visuals to detect missing or emotionally mismatched B-roll.
  • The workflow uses Frame.io for transcript generation and comment placement, with an MCP connector to automate the loop.
  • Higgsfield's MCP connector allows Claude to directly generate B-roll clips using models like SeeDance, GPT Image, and Soul without manual routing.
  • Step one: Claude watches the video by pulling the transcript via Frame.io MCP or Chrome extension.
  • Step two: Claude flags moments where B-roll is missing or wrong and posts frame-accurate briefs back to Frame.io.
  • Step three: Claude reads its own comments, routes each brief to the appropriate Higgsfield model, and renders the clip.
  • Step four: Claude uploads each generated clip as an attachment on the corresponding Frame.io comment for the editor to drop in.
  • The creator built a custom Claude skill that teaches the model producer-grade taste for specific client channels, improving B-roll relevance.

Tools mentioned

Techniques

  • Transcript-based video analysis
  • Automated B-roll generation via MCP
  • Custom Claude skill for emotional visual matching
  • Four-step automated review-to-production workflow
Transcript (captions)
I taught Claude Fable 5 to watch and edit my videos and the videos of my clients. So, today I'm going to show you exactly how. I'm going to show you the prompts, the tools, and where each one plugs in. So, if you run a content-based business, a channel, a course, a podcast, a launch, I know that the bottleneck in what you ship is not the script and it's not the camera. It's the thing in the middle. The one where you watch your own cut, mark what's off, hunt for the clip that matches the emotion of the line, maybe give up, use something close enough, and then publish a video that's a little flatter than it should have been, but that's the best you could do. Well, that single loop is where most creators lose their weekends and where most launches end up slipping past their deadlines. So, I built a system that closes that loop. I'm going to walk you through it using a real example, a YouTube video that I shipped last month for a client in the relationship coaching space on a deadline that on paper was not going to happen. And before you ask, the tech stack we used was actually pretty simple and straightforward. >> [music] >> Two tools, Claude Fable 5, which is Anthropic's newest model, recently launched. So, here's the distinction. You don't want to treat Fable 5 like a chatbot [music] because that's like treating a chief of staff like a virtual assistant. They're the same species, but a very different type of employee on a totally different level. Fable 5 reasons through complex problems. It judges [music] its own output. It iterates when the first pass is not right. So, if you give it the right tools, it'll [music] stop answering questions like any other chatbot and it will become your [music] work associate. The second tool we use is Higgsfield, which is the fastest growing GenAI creative platform. Every serious visual model in one place. See dance, Soul, GPT image, Cinema Studio, Marketing Studio. Their MCP connector plugs Claude, including Fable 5, directly into that entire toolkit. By the way, I'm partnering with Higgsfield on this video because that connector is what turns the whole workflow from a review checklist into a full production studio. >> [music] >> Without it, Fable 5 can say what the visual should be. With it, Fable actually creates the visual. [music] And thank you, Higgsfield, by the way, for partnering with us on today's video. So, let me show you how the entire workflow runs. So, look, the story most creators are told is that AI can help with the front half of the content, [music] with the scripts, with the outlines, with the hooks, maybe with the thumbnails, but editing is off-limits. It has to be the human part. That is a limiting belief that basically leaves an entire half of your production schedule on the table because editing, the real work of it, >> [music] >> is not the drag-and-drop part. It's the decisions. It's the little details. Where does the line need a visual? Is the visual we cut to actually matching the emotion the line is carrying? Is the pacing right? Are we losing the viewer at the 4-minute mark because the B-roll went flat? Well, those are all reading >> [music] >> decisions. Reading the transcript against the cut. And reading is what Fable 5 is best in the world at. [music] Because missing B-roll is a real problem. Most independent creators have whole sections of their videos with nothing visually happening, >> [music] >> just A-roll, which is basically what you see here called, in production terms, the talking head >> [music] >> and nothing else, dead air. But in any cut that actually has B-roll, there is a worse problem, and almost no one actually realizes it. And that's [music] the wrong B-roll. Because, you see, maybe the clip is there, but it just does not carry the emotion of the line being spoken over it. >> [music] >> The script is about, let's say, a couple rebuilding trust, which was our case. The B-roll is a stock shot of two strangers walking through a park. >> [music] >> Clearly, the viewer doesn't consciously notice. They just feel the cut go flat. So, if they're on YouTube, they just [music] click off or move to the next video and you lose them. And actually studies on viewer retention consistently show that the steepest drop-offs happen at points where the script and visual go out of [music] sync, even when the visual is technically fine. Generic isn't a feature, specific is. That is what changed for me. The system that I'm about to show you catches the wrong fills, generates the right ones, and closes the whole loop in a matter of minutes instead [music] of days. The whole workflow has four steps. Every video that we ship, whether mine or our clients, >> [music] >> goes through these four steps. Step number one is the watch. Fable 5 needs to go through the video and watch it line [music] by line. I'll explain in a second. Number two is the marks. Fable 5 needs to flag every moment where the visual is missing or where it doesn't match the line >> [music] >> and then post a frame-accurate brief back to the tool that we use to review videos, which is called Frame.io. Step number three is the render. So, Higgsfield MCP will take over from here because Fable 5 hands over the brief to Higgsfield and Higgsfield generates the clip. >> [music] >> An MP4 lands in your folder and then that step turns the whole workflow from a review checklist into an actual production studio. We'll come back to this and I'll show you exactly step by step how to do it. The fourth step is the drop back. Okay, so Fable 5 needs to upload each clip back to Frame.io as an attachment on the comment that it fills. So, when the editor opens Frame.io, they see the clip already to the exact time code and can drop it directly into the cut. Okay, so I'll walk you through in the way that I actually run it using >> [music] >> actual footage but not of my client. We'll use my own just for >> [music] >> privacy reasons. So, like I said, step number one is the watch. So, basically when my editor finishes the first cut, they upload it directly to Frame.io and drop the link into Slack. >> [music] >> And that's the trigger, okay? The instruction to Claude is not something that I type. It's automated. I have a small automation watching Slack, so whenever an editor drops a fresh Frame.io link into the channel, Claude knows exactly what to do with it. Opens the link and watches >> [music] >> the video. Now, when I say watches the video, what that actually means is slightly different because it's not what you would imagine, okay? Claude doesn't only look at pixels. So, it can see, but it cannot hear. [music] So, for that to happen, Frame.io generates a transcript from the audio automatically, every line timestamped down to the frame, and then Fable 5 pulls that transcript through the Frame.io MCP that we built for this workflow, which, by the way, I'm going to share with you in the community. And then, it reads the transcript. That's how it listens to what is being said, okay? It goes through the transcript, reads it line by line against the visual moments that goes with [music] each one. So, that's basically the watch. It's not a pixel-by-pixel, sound-by-sound analysis, just the reading and the watching the images, the visuals [music] in the video. And that actually happens to be what Fable 5 is best in the world at. Now, if you're allergic to MCPs or you feel uh intimidated by this, if you don't want to install a connector, you can just run the same step with the Claude Chrome extension, okay? >> [music] >> Claude can open the Frame.io link in the Chrome extension, can see in the browser the transcript panel, [music] scrape it. It's a little slower, but it works. I ran that in the beginning before we built the MCP. I think the Chrome extension can be your how do you call those? Help wheels when you learn to ride a bike, >> [music] >> and then the MCP can be your end game. Now, the transcript is the source of truth for everything that comes next, every timecode, every motion, every place a visual should have landed and didn't. So, we'll come back to it. Now, step number two is the marks. Okay? Because now Fable 5 does the work that no one else is doing yet. It reads the transcript line by line and for each line, the question is what should the viewer be seeing right now? >> [music] >> Okay? That is how Claude is trained. And I need to flag two possible situations that Claude might run into. Number one, if the visual is missing. So, let's say there's just A-roll. There's nothing else, just the dare, like I said. Number two, maybe there is a visual, but it's not the right one. >> [music] >> It could be wrong. Maybe the clip does not match the emotional register of the line that's being said. A line about reconciliation playing over a stock shot of somebody scrolling through their phone. >> [music] >> You know, you can get the gist. Those are the kind of situations that Claude is trained to catch. [music] Okay? So, Fable 5 does not have producer grade taste for this out of the box. The taste is trained, [music] like I said. So, what we did was we built a specific Claude skill that teaches the taste for each specific client. [music] So, then Claude knows exactly what it should be watching for and what kind of emotions the ideal viewer of the channel is actually expecting. So, basically, the Claude skill that I'm talking about teaches Claude how to spot the two failure modes and how to write a B-roll brief that actually fixes those. >> [music] >> Which shots go with which emotions, which framings land, which shot length skill, which registers drift. The skill is what turns a competent flag, this line needs a visual, into a producer grade one. Something like this line needs a 3-second close-up of hands, warm tones, hesitation before closeness. You see the difference, right? That skill is the second thing that I'm actually going to share with you in our community. So, stay tuned. I'm going to show you the details at the end. [music] Now, for the relationship coaching client, this is where the whole project was saved. The first cut had B-rolls in most of the places, but they were generic. For example, the script was specific, particular kind of intimacy, a particular kind of conflict. But the B-roll was just people [music] existing in the same room. So, in the end, Fable 5 ended up flagging something like 60 moments where the B-roll was there, but was not pulling its weight. >> [music] >> We hadn't noticed half of them on our own pass, I have to be honest. Because they were the kind of thing that you only catch when you read against the line. [music] But Fable didn't stop there. Fable posted the marks straight back into Frame.io >> [music] >> with the right timestamps, with specific briefs as well, okay? Not just need B-roll here. More along the lines of need a 3-second close-up of two hands clasping across a kitchen counter, soft window light, warm tones. The line is about rebuilding trust after a hard year. You can see how big of a difference the B-roll would make when right. Now, if you're wondering how does Fable get the time codes accurately? Well, it gets them frame by frame by clicking individual words in the transcript. So then Frame.io jumps the playhead to the exact moment when a word is clicked. And Fable 5 reads the precise time code back. >> [music] >> The editor can trust the placement exactly. Now, step number three is the render. Okay, so here's where our story completely changed. [music] Because before Higgsfield MCP existed, this workflow ran halfway, okay? Claude would still write the brief and post the marks, beautifully specific and exactly what each line needed, but the generation step was still manual. So, someone in our team would actually end up manually [music] generating every single one of those B-rolls. Think about generating 60 of them. It's a lot of work, right? They would take each brief, [music] open Higgs Field in another tab, generate the clip, wait, download it, come back to Frame.io, upload it as an attachment on the comment, and then the next one, and then the next one. Going through the same loop, every clip was a separate context switch. Claude could design the visual, but we were the ones running the render. The MCP is what closed that gap. Because with Fable 5 driving Higgs Field through one connection, this basically became the step that turns the workflow from a good review pass into a full production studio. >> [music] >> So, let me walk you through it in detail, because it is where most of the value lives when it comes to this whole process. [music] So, the first part within this step takes about the same time as making a cup of coffee, okay? And it's the only setup that you'll ever do for this workflow. Ever, I promise. [music] It's five clicks, totally not intimidating. Walk with me step-by-step. Go to settings, click connectors, and the plus icon, and then paste the [music] Higgs Field MCP URL. mcp.higsfield.ai/mcp. That's it. You can just copy it from the screen here. And then, click add. So, now Fable 5 has Higgs Field's full generation stack available inside your chat. >> [music] >> Okay, you don't need to install anything. You don't need to configure model routing. Fable 5 handles that automatically per [music] click. Now, the next step, what Fable 5 is actually plugging into. Because before we run a single render, I need to show you what's on the other end of the connector, okay? The toolkit is bigger than the surface suggests. [music] So, this is Higgs Field AI. When Fable 5 calls the MCP, this is the entire toolkit that it has access to. >> [music] >> SeeDance 2.0 for cinematic video, GPT Image 2 for stills, Soul for character consistency. If you're building a series with the same face across every clip, that is going to live here. >> [music] >> Cinema studio, marketing studio, the trending database, Fable 5 [music] picks the right one for each clip in your brief. You don't need to pick. Okay, you don't need to know even. You just describe what the line needs and then Fable 5 does the [music] routing. And honestly, not having to make that decision changed the game for me and hopefully it will for you as well. Because you don't have to >> [music] >> think about seed ants versus GPT image versus soul versus others. You just write [music] the brief for the motion. Then Fable 5 handles the tool selection. Now, the third step is the actual render. Because this is where Fable 5's chief of staff site does the work [music] that no chatbot can. Fable 5 does not need to paste the comments back into a new chat. [music] Okay? It can see what it wrote through the same Frame.io MCP that pulled the transcript in the first place. Or if you're using the Chrome extension, it will see them in your browser. And it has one instruction that I shared with it. I asked it to walk the comments and for each one take the brief, route it to the right Higgs field model, drop the edit into a specific folder, and name it using the time code. If you want the exact prompt, that will also be available in the community. >> [music] >> So now, Fable 5 opens Frame.io, reads its own comments back, and starts working through them. For each one, it picks the model, seed ants for motion, GPT image for stills, soul when the shot needs character consistency, and calls Higgs field through the MCP. Like I said, I don't pick the model. Fable does. [music] That's the difference between a tool and a partner. Okay? So, that's the first one. It took 3 seconds. It has soft light, warm tones, exactly the brief. Now, compare that to what you'd get from a stock library. Stranger's hand, the wrong lighting, the wrong emotion, and the different lens the second that you see them side by side. >> [music] >> Now, step number four is the drop back, okay? Once the clips are in the folder, there's one [music] more step. Fabel uploads each MCP back to Frame.io as an attachment on the comment that it belongs to. The same Frame.io MCP that pulled the transcript in step one, [music] okay? So, now the editor doesn't have to go hunting for a folder full of files with timecode names. >> [music] >> They open Frame.io, see the comment, see the clip attached to right there, and drop it into the cut. Okay? Now, there is one thing that I don't automate. >> [music] >> There are now MCPs like Jumper that go one step further and place the clips directly onto the timeline in Premiere or Final Cut or whatever you're using. You can automate the drag in, but we [music] don't, okay? I personally want my editor doing the final touches, the pacing, the cut length, the last read before it delivers. That's where there is still a need for taste and human touch. So, Fabel can build the pipeline, but the editor still runs the final cut. Now, I want to make another clarification because I usually get asked about this. Maybe [music] you don't want to run this from Cloud at all or you don't have Cloud. Well, you can do everything >> [music] >> within Higgs Field because Higgs Field has a second entry point called Supercomputer where you basically pick Fabel 5 as the brain and run everything inside Higgs Field AI. It's a different workflow, but it uses the same model. I can cover that one in another video if you want me to, just leave a comment down below and let me know. For today, the MCP inside of Cloud path is what I actually use, so that's what I'm showing you. Now, there's one more layer that I think it's important to be aware of. This is not just a workflow, [music] okay? It's also a service that you can provide. Every solo YouTuber that you know is one editing pass away from doubling their output. And no, I am not promoting full AI workflows. I don't believe that YouTube is a place where mass producing videos without a human touch is the right call. What I'm saying is that using this workflow, you would be able to >> [music] >> support YouTubers in making better videos by also adding your own human touch throughout the workflow process. And YouTubers are not your only ideal customer. Every course creator has a content library that they cannot do anything with because the post-production tax is too high for them. Every faceless channel needs a constant feed of clips as well. All of them can be your ICP. >> [music] >> But don't sell B-roll generation, okay? That's the commodity. Sell senior producer review. The past that reads the transcript, watches the cut, marks what's off, generates the fix, and uploads it back. End to end, [music] okay? Even the full edit. And you can take half the time that it usually would. Now, when it comes to monetizing that, a very clean way would be to go for a monthly retainer per creator if they have a constant flow of videos that they need to produce every month. This way, you get predictable revenue, they get predictable delivery velocity, >> [music] >> and the trade is undefeated. Now, a creator who can go live with one extra video per week will make back any retainer you charge within a couple of months from the compounded views. So, you can frame the offer that way, and you no longer compete with stock subscriptions, but you actually offer added value. Obviously, as always, results depend on your outreach volume, the quality of your prompts, how you dialed in your first 10 case studies, the ones that come after the first 10 come through referral, usually. Now, throughout the video, I promised you three [music] things, okay? First, the Frame.io MCP that we built for this workflow. Second, the Claude skill that teaches the taste needed to identify and brief B-roll properly. And third, [music] the raw prompts that we use across the whole loop. All of these live inside of our Founders [music] Hive. The link is in the description. You can also use the QR code here, and you can join us. There is a lot more than just this in the community. You get my entire Cloud Academy. You get the entire business training on how to turn an idea into an actual business, how to frame your offers, how to sell, [music] how to close deals, all of that within the Founders community. Not only that, but we also speak every single week, so we can keep you accountable, and you can help us keep everyone else accountable. We're such a tight community. We really enjoy spending that time together, and we look forward to seeing you there as well. If you're not ready for the Founders community, we also have the Trailblazers Hive. That's for beginners who are looking to dip their toes in the world of AI. We have loads of challenges that you are always welcome to take. They're all free, and you will get so much value just from those alone. You can also come to our Hive Holes every 2 weeks, and there's a lot of [music] activity and community feeling in there, so you can ask questions, you can support everyone else. There's over 15,000 other people on the [music] exact same path as you are. So, I hope to see you in one of our two communities. In the meantime, thank you so so much for watching. Like this video if you did. Be sure to subscribe if you haven't done so. Share it with anyone in your circle of friends or family or coworkers who [music] are interested in making their video production workflow a little bit more fluid, a little bit less handheld. And [music] until next time, I suggest you go ahead and watch this video here, and I'll see you soon. Bye.

Frontier News · by Hyperjump Technology