Can You Replace Your AI Coding Subscriptions With Open Source?

summarized

TLDR

Open source AI coding agents with local models can handle routine changes like adding a CSV export, but the lack of measured local performance means paid subscriptions remain worth keeping for unresolved or difficult work. The example shows that careful acceptance testing and tool permissions are essential, and the cost arithmetic depends heavily on hardware and time assumptions.

Key points

Open code agent uses tool calling to let a model read files, edit code, and run commands with configurable permissions.

The CSV export example exposed a bug where commas inside labels break column parsing, which was fixed by escaping special fields.

Context window size (64k tokens recommended) limits what the model can act on, and automatic compaction may lose important details.

Fully local inference removes hosted inference costs but adds hardware, electricity, and time, making the savings uncertain.

The speaker recommends first moving routine well-tested changes to an open workflow before cancelling any paid subscriptions.

Tools mentioned

Techniques

  • Tool calling (function calling)
  • Plan mode for analysis
  • CSV escaping rules (double quoting special characters)
  • Context window management and automatic compaction
  • Permissions controls (allow/ask/deny)
Transcript (captions)

0:00 Can open source replace the coding subscriptions you pay for? I'd judge it by a finished change, including the broken bits. A free chat window hasn't earned your cancel button yet. The

0:09 workflow is to understand a repository, add a feature, run its tests, and fix what fails. Open code and all provide the pieces. The example here was built and tested during research without a

0:21 local model generating the code. It demonstrates the work a replacement needs to finish. It doesn't establish a model's ability to finish it. Which subscription would I still keep? One

0:31 that earns its place on work this setup hasn't proved it can handle. Take a small expense tracker. It already stores labels and amounts and it can add up the bill. Now it needs an export that

0:41 another program can read. The change has to fit existing code and preserve the data including labels with a comma, a quote or a line break. Those ordinary inputs give us something concrete to

0:52 inspect. At the end, there are two different jobs inside an AI coding setup. The model produces a response from the information it receives. The agent manages the conversation in the

1:01 tools that can read files, change code, and run commands. Open code supplies that agent. A lama serves the model, which means another program can send it a request and get a response. Keeping

1:12 those jobs separate lets you change where the model runs without designing another file editor. It also explains why downloading an agent doesn't give you the computing power behind a paid

1:21 service. The open code and repositories each carry an MIT license. Model weights have their own terms. Quen's current 3.8 27 billion parameter model card list Apache 2.0. That's a concrete model you

1:35 can inspect, not a promise that it's the right size for your laptop. Olama's catalog lists an 18 GB download for its default version. The download size isn't the total memory requirement once the

1:46 model is working. You also need room for the active conversation and the runtime. A smaller file on disk doesn't settle the quality question either. Alama's documented setup command is Alama launch

1:57 opin code. Its guide also shows a manual connection to a local server. The address on screen points back to your own computer and the model name tells the server what to use. Both have to

2:07 match the model you actually loaded. An open source client can also connect to a hosted model, so the client's license doesn't tell you where your code is going. Alama documents a local only

2:17 setting that disables its cloud features. Agent web tools are a separate choice. Check the whole route before describing the workflow as offline. Before asking for the feature, give the

2:26 agent a way to understand the repository. Start with the readme, the file that stores expenses and the existing tests. Ask where the change belongs and which command verifies it.

2:36 Open code documents a plan mode for analysis with edits and shell commands subject to its permissions. That gives you a place to check the proposed approach before the code changes. For

2:46 this repository, the important detail is that amounts are integer cents. 1,200 means $12. The export should preserve that representation. The baseline has two tests and both pass. Save that

2:58 result before making the change because a later failure only tells you something useful if you know what worked beforehand. Then write down the acceptance conditions. Keep the totals

3:07 working. Include the column names and preserve awkward labels. Open Code's rules documentation describes putting project guidance in an agents.m MD file. Keep the commands and conventions there.

3:18 The tests themselves remain executable evidence. A sentence telling an assistant to be careful can't tell you whether a comma moved a price into the wrong column. Now follow one request

3:28 through the system. You ask for the export and provide the relevant files and rules. The agent sends that context along with descriptions of its available tools to the model server. The model can

3:39 respond with a request to use a tool. The agent checks its permissions, executes the allowed action, and sends the result back into the conversation. Alama's tool calling example shows this

3:49 exchange explicitly. The model's request to read a file, and the files actual contents are separate messages. That distinction is what turns a suggestion into an action with evidence behind it.

3:59 The same exchange can produce a patch and then run the test command. If it fails, the next request can include the failure, the relevant source, and the requirement that was violated. The model

4:09 can propose a repair based on what happened. It might still fix one case and break another. Verify the code left on disk, especially when the conversation has been more confident

4:18 than the compiler. The CSV example makes that feedback visible. The first implementation deliberately joins the label and amount with a comma, then joins records with line breaks. For a

4:29 plain label, such as hosting, that works. for hosting EU. The comma inside the label looks like another column separator. A parser can put EU where the amount should be. The export still

4:39 returns a string, but the data has changed meaning. Checking that the function returns something would miss the problem. The test run catches three failures, the comma, the quote, and the

4:49 embedded line break. Four other tests pass, including the two tests for totals. That narrows the problem to how fields are written. The repair encloses fields containing those special

4:59 characters in double quotes and doubles any quote inside a field. Those are the escaping rules described in the CSV format document. Now the comma in hosting EU stays inside the label while

5:10 the separator after the closing quote introduces the amount. One character can have two jobs. The surrounding quotes tell the parser which job it's doing. After that repair, all seven tests pass.

5:21 A separate check reads the actual file with PowerShell's CSV parser and gets back the original labels and amounts, including the line break inside a label that exercises a different reader beyond

5:32 the strings expected by the test author. The versions, commands, and results are saved with the research. This deliberately naive version isn't evidence that open code made a mistake.

5:41 A production export needs more requirements, including its target import tool and how it handles invalid amounts. Keep those conditions the same across agent comparisons. Otherwise, an

5:52 assistant that skips difficult cases can appear faster because it delivered less. Count your manual corrections as part of the work. As the work continues, the conversation grows. It now contains

6:02 instructions, code, tool descriptions, patches, and test output. The model can only use what fits into the request it receives. That working space is called its context window. Alama's current

6:13 documentation recommends at least 64,000 tokens for coding tools and agents and warns that a larger context needs more memory. Its documented default below 24 GB of graphics memory is much smaller.

6:25 So a model's advertised maximum, the server's configured capacity, and the information an agent sends are three different things to check. Think of that context as the working desk for this

6:35 export. The rules and current code need to stay within reach. A log of unrelated passing test takes up space without explaining the failing comma case. Preserve the exact assertion failure and

6:46 the code it refers to with the complete log available on disk. You're choosing what the model can act on and an omitted requirement can change the solution it proposes. Open code documents automatic

6:57 compaction which uses a smaller summary when a session gets long. It also exposes a separate option for pruning old tool output. Neither is a guarantee that every important detail survives.

7:07 For the export, a useful summary would preserve integer sense, the escaping requirement, the files changed, and the current test result. The repository still holds the full source and tests.

7:18 After compaction, the agent can read them again. If its summary forgets that a label can contain a line break, the test should bring that requirement back into view. That's also why I'd put

7:28 permissions around the tools, even with a local model. Local inference tells you where a model request is processed. It doesn't prevent a shell command from changing files or using the network.

7:38 Open code provides allow, ask, and deny controls. Use a disposable project copy for an initial trial. Keep credentials outside it and inspect the diff. Privacy and correct code are separate things to

7:50 check. Then count what the workflow costs. Cursor's pricing page lists Pro at $20 a month before applicable taxes. GitHub's documentation lists Copilot Pro at $10 a month. If you pay for both

8:02 monthly, that's $30 or 360 across a year at those prices. This illustrative pair has different features. Identify which job each plan does for you before calling the payments redundant. A fully

8:14 local model request has no hosted inference invoice. You still supply the machine, electricity, and time. The local agent run is unmeasured here, so its energy use and completion time

8:24 remain unknown. The test command's execution time can't tell you how long a model would take to produce the repair. Keep those empty cells empty. Here's a simple budget you can adapt. Suppose you

8:34 value your time at $30 an hour, and the local workflow takes one extra hour of intervention each month. That alone equals the illustrative $30 subscription pair before electricity or a hardware

8:45 purchase. It's arithmetic using stated assumptions, not a claim about your wage or a measured slowdown. With hardware you already own and little intervention, the calculation can favor local use.

8:56 With a new machine and repeated repairs, the saving needs more evidence. Measure completed changes, including review, rather than the speed of text appearing. An open source agent with a hosted model

9:07 is another arrangement. It lets you choose the interface while the provider still charges for service. For a useful comparison, keep the task and review standard fixed. Record model settings,

9:18 failed attempts, and manual corrections. A result you can rerun tells you more than a successful final message. So, I'd move routine well- tested changes into an open workflow first. I wouldn't

9:28 cancel every coding subscription based on this example because the local agent result is still unmeasured. Keep the service that resolves work your local setup can't yet finish reliably. Then

9:38 check whether you actually use it. The next useful step is the same feature on your own machine with the same acceptance test and a complete cost record. Our publish PI agent and

9:48 llama.cpp CPP setup video covers the runtime side. Use this export as the small job that tells you whether that setup is doing useful work.

Frontier News · by Hyperjump Technology