My NEW FAVORITE Skill - Claude Code Drives My Whole Computer (Better Computer Use)

summarized

TLDR

A lightweight screen control skill for coding agents, built entirely with CLI commands and no additional tooling, shows that modern LLMs like Fable 5.1 and GPT-6 Astra can reliably drive a computer screen without the need for heavy computer use harnesses. The skill, under 400 lines, handles failure modes and is customizable per OS, making it a practical alternative for tasks like morning setup and app testing.

Key points

The skill uses only command-line commands (PowerShell on Windows, AppleScript on Mac) to control the screen, no extra tooling.

The skill is under 400 lines and includes scripts for deterministic window discovery, focusing, typing, and pasting.

The creator tested the skill on Mac, Linux, and Windows with multiple monitor setups.

The creator claims that modern LLMs (Fable 5.1, GPT-6 Astra) are very resistant to prompt injection attacks, reducing security concerns.

The skill is available as a plugin from the creator's AI skills GitHub repo.

Tools mentioned

Techniques

  • CLI-based screen control
  • deterministic window management scripts
Transcript (captions)

0:00 I have a new favorite skill that I'm using in Cloud Code and Codeex every single day for everything under the sun. I call it drive screen. And I want to show you how this works today because I

0:10 built it to be a super lightweight alternative to the really bloated and hard to manage computer use tools like what Claude has built into their platform. We have Codeex computer use.

0:19 There's a ton of open- source tools out there as well. And these are all powerful, but I've never been a fan of using them. That's why I built my own skill. And the beauty of what I built

0:28 here is you just need your coding agent. You don't have to install anything or build any kind of harness. Large language models, especially with Fable 5.1 and GPT6 Astra, now they're good

0:37 enough where they can drive your entire screen for pretty much anything with just the command line. I didn't realize that until I just tried it out a couple of weeks ago. And that's eventually led

0:48 to building this skill, which I love cuz it works no matter the operating system. And now I don't have to install anything else or rely on a massive out-of-the-box computer use tool. Like I said, I've

0:58 never been a fan of these. I want to control my tooling. And with these tools, it's really hard to customize things and tweak it as little issues come up as you're incorporating it in

1:07 your workflow. And so because of that, I've just always strayed away from computer use as a whole. The problem is computer use is incredibly useful and I'm realizing that more and more the

1:17 last couple of weeks. There's a reason I'm obsessed with using the skill now. So, if you're not using computer use at all right now, I would encourage you to at least start dabbling with it with a

1:27 tool like, you know, codeex computer use. Of course, I would encourage you to try out the skill that I'll link to in the description. There are just endless possibilities, ways to boost your

1:35 productivity using this kind of capability with your coding agent. Now, of course, there are some security concerns, especially when you're running just a skill instead of a full harness.

1:44 We'll definitely talk about that, but there are so many quick wins I've started relying on every single day that I want to show you here. I mean, the most obvious one is to use your coding

1:54 agent to help you get set up in the morning. So, you turn on your computer, you open up your coding agent, and you say, "Get my morning set up, right? Like, here are the things I'm working on

2:02 for today." A it could even go look at your task management software with the CLI or MCP and then it can get things open for you in your browser tabs, opening up tabs in Obsidian, getting

2:11 desktop app started, Docker containers, whatever you need. It can set up everything. And so like this alone saves me like a solid like 10 15 minutes every single morning. And that adds up,

2:21 especially with all the other things I'll show you here. And so it just uses the skill and simple commands to drive everything. There is no additional tooling installed here. And I've tested

2:32 this skill on each operating system, Mac, Linux, and Windows. I've tested with multiple monitor setups and different display settings. It's really good and really adaptable. And here's

2:41 the thing, even if you're not super interested in this specific skill I'm showing you right now, the big takeaway here, what really just blows my mind is how adaptable large language models have

2:51 become. There's so many tasks that I've just written off before as the LLM can't do it, or I think you need some massive tool like for computer use where now you just don't. Like you probably have a

3:01 couple ideas in your head right now for something you think an LLM can't do like computer use or editing a video. You think you need some fancy tool, but you actually don't. Just like try it out

3:11 right now. Give it an MP4 and say edit this video. Take out the ums, add in B-roll and transitions. And especially with like Fable and Astra, like it'll knock it out. And to be totally honest

3:20 with you, this drivesreen skill has been simply born out of necessity because recently I've just been kind of overwhelmed with all the different things happening in the AI space, things

3:29 I got going on, applications that I'm trying out, especially a lot of open source projects like Kestre for example. I'm trying out all of these and a lot of them they have a UI or there's something

3:38 where I can't just use my browser automation. So I needed some kind of computer use to test things out, but also I wanted something super lightweight and easy for me to manage. I

3:48 didn't want to bring in a full computer use tool. And so at the time I just took this GitHub URL and all the other ones I was testing. I mean this is just one example here. I gave it to Claude and I

3:57 said research the web, dive into the repo to understand the features, then start the desktop app and then control my screen to test out different features and then leave it open for me at the end

4:07 to continue my own testing. And I had no idea if that would work. I thought maybe it would go for some larger computer use tool and install it or maybe it would just say it can't do it and tell me to

4:16 do things manually. But to my pleasant surprise, it was able to go through it end to end simply by running commands in the terminal. So this is about what my session logs looked like, especially

4:27 before I had the skill. So it runs command after command to click around and take screenshots and analyze results, but it did the whole thing end to end. It it took a while. It's not the

4:36 most efficient. I mean, computer use isn't in general, but I just let it go, take control of my computer. I worked on another device. I came back and everything was done and everything was

4:46 still up for me to continue the testing as well. It was a beautiful thing. And with all these commands for me on Windows, it's just simply PowerShell, right? Like there's no additional tool

4:56 brought in. For Mac, you can use Apple script. Again, I've tested this for every OS. And so, I was hooked on this right away. Testing a bunch of different apps, doing my morning setup like I

5:04 showed earlier, having it set things up for my presentations and recordings every single day. Now, it didn't start as a skill like I just say like use PowerShell to drive my screen and do

5:13 XYZ. But it became obvious very quickly that we need a full skill because sometimes the computer use wouldn't even work through PowerShell. Like the agent would get totally tripped up or there

5:22 would be certain failure modes that I'd see come out as patterns. And so I've documented all the failure modes. I've laid out the workflow, even built some scripts to make things more

5:30 deterministic. And that is what I have packaged up as the skill for you. The sponsor of today's video is Plural, and they built something genuinely cool. DevOps agents that live in your own

5:41 Kubernetes clusters. Figure out what's broken when something is and then open up pull requests to fix it. They call it workbenches. So, let me show you. Right now, in my own cluster, I have one of my

5:51 services, the payments API, that is crash looping. But all I have to do here is go to the workbench where I can send off any job. And I simply said the payments API service is crash looping.

6:01 Find the root cause and open up a poll request to fix it. And here, Plural spins up sub aents that each have their own context to dive really deep into the pods, the logs, the GitOps repository. A

6:11 couple minutes later, it identifies the exact problem. It gives me the diff and the evidence and even builds me a dashboard to help me understand the problem that's happening here in my

6:19 cluster. And then my favorite part is it spins up a cloud code instance in my cluster to make the fix and then open up the pull request. And so the agent runs in my environment with my model keys and

6:30 I still get to be the one to merge the pull request in the end. And you can wire up that same workbench to your alerts, run it on a cron, or even trigger it by saying plural fix this in

6:39 a Jira ticket. You can self-host Plural or use their cloud and it's free to get started. I'll have a link to them in the description. So, of course, I want to get into how this skill works, how I've

6:49 kept it simple, security considerations. We'll talk about that in a second. But what I want to do for you really quickly here is show you a live demo of this skill working. And so, I'm obviously

6:57 going to speed it up a lot, but I'm going to send off this prompt so we can see it in action. loading the skill and then super easily going through any kind of setup that I have for it. And so I'm

7:06 going to turn off my video and then fast forward here so you can watch it happen. So there we go. We completed everything within just a minute here. I know it's a very simple example I'm showing you

7:20 here. I don't want to do anything super comprehensive right now, but the point is that it's fast. It didn't take a lot of commands either because the skill guides it and even gives it some tooling

7:29 to move faster. And by the way, what it did here even includes validating the environment to make sure it has all the tools it needs and that it has the permission it needs on the machine as

7:38 well. And so that's just a super quick example. Now, let's talk more about how the skill works. So this skill as a whole is a little under 400 lines long. So not the longest, but also not the

7:48 shortest. I I wanted to have a full workflow laid out for the agent here and also make it work for every operating system. And there's a lot of nuances I wanted to build in little failure modes

7:58 that I've seen as I've tested it out. And so the workflow starts by figuring out, do we even need screen control at all? I want the agent to have the opportunity to push it back when you use

8:08 this skill because sometimes it really doesn't require computer use when you think there's some other kind of tool like browser automation to get the job done. Like it says right here, screen

8:18 control is the slowest and least reliable way to make a computer do something. So it's the most adaptable and flexible, but it's the slowest. And there's the security issues where you

8:28 can have prompt injection attacks come on your screen so the agent views that and does something malicious. Large language models really don't fall for that anymore which is also a big reason

8:37 why I'm trusting it with this kind of capability now. But just something to keep in mind. That's also why I stick to really simple things when I'm using this tool. If I want any kind of production

8:47 grade computer use, I'm still going to be using some kind of larger harness or tool. And then going up to the top of the skill really quickly, I want to call out the different scripts that ship with

8:56 the skill. This provides more determinism. So instead of the agent having to figure out the exact commands for certain things in real time like I had to do initially, I have all the best

9:07 practices packaged up into these scripts now for window discovery and focusing and typing and pasting. There's many different ways to do these things. I figured out the best from all the

9:16 testing I've done. And let me tell you, I've done hours and hours and hours of testing and I built it into essentially this little CLI for the agent to call upon with the different commands and

9:26 arguments. And so it's my own little custom CLI. We don't have to get into the details of all this right now. There's quite a bit here because there's so many different operations the agent

9:35 can perform, but it's a big part of what makes this more reliable. And it's also very easy to customize things. If you're working on some other kind of distribution of Linux for example that

9:44 might need things to operate a bit differently, you can just go in and edit the skill, right? So like when you install the skill from my skills repository, everything comes with the

9:53 reference, of course, the skill.md and all the scripts for the CLI. And then scrolling back down again, we have all the hard rules. These are things I learned the hard way just as I was

10:02 building up the skill and using it over and over again. So that makes it more reliable. We have the control loop. Like whenever you're operating in a window, here's what you should do, right? like

10:10 first discover the window with a command screenshot it so you have the lay of the land focus in on it then you can act and here are the ways that you act and the CLI that you have to help you with that

10:20 right so you can see everything coming together here I don't even have to go and read through literally everything but there's also like a whole CLI for watching a session like if you have it

10:29 drive another coding agent session it needs to watch for approvals or things like that I've done a lot of that as well auto driving we have uh different traps that we needed to watch out for.

10:40 And so some of these things are specific to what I've been automating with like browser tabs and desktop applications and then other coding agent sessions. Maybe you're going to be doing the same

10:49 things, maybe not. But again, that you can take this skill and customize it to your liking. So if you want to take out traps cuz they don't apply to you or add some in, you can feel free to do that,

10:58 right? Like that's the kind of thing that you can't really do with a larger computer use tool. Just going back to that reason again and again why I love having this flexibility so much. Now,

11:07 like I mentioned a bit ago, you also do have to be aware of prompt injection attacks for just computer use in general. I'll link to this article in the description. It's a really good

11:16 read, but also I got to say, and this might bite me in the butt a little bit, I don't think you have to worry about prompt injection attacks that much anymore, as long as you're using the new

11:26 best models like Fable 5.1 and GBT6 Astra. There are a lot of studies coming out that are showing that even without fancy guard rails or a fancy harness, these LLMs are very very hard to prompt

11:38 injected. And so I'm I'm just not concerned about it these days. At least nearly not nearly as much as I was before. I don't think you have to be either. And especially if we're using

11:47 this skill for simpler things like what I showed in this video, like morning setup and testing desktop apps, things like that, you really don't have to worry about it. And so yeah, I would say

11:56 just make sure you're using the most powerful LLMs when you're using computer use makes it the most reliable anyway. And this skill, it is pretty token efficient. It's not the fastest, but

12:05 it's because it's waiting on different inputs and things, not because it's burning through tens of thousands of tokens. And so use at your own risk, and I would encourage you to use better

12:14 models. But I will say, at least for me, I don't consider it a risk for computer use anymore, especially when I have something that I can control so much like this. So, I would encourage you to

12:23 install the skill, try it out. You don't have to bring in any additional tooling. It's just right here in my AI skills GitHub repo. So, you can install it as a cloud code plugin if you want to bring

12:32 in all my skills or you can just take that skills folder that I have linked in the description and bring it into any coding agent. And so, I hope that you find this useful and even just exploring

12:42 with me in general the new capabilities of LLMs, how they're able to do these incredible things without an additional harness anymore. And so if you appreciated this video, you're looking

12:51 forward to more things on coding agent skills and AI coding in general, I would really appreciate a like and a subscribe. And with that, I will see you in the next video.

Frontier News · by Hyperjump Technology