Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
WebLLM allows developers to run AI models directly in the browser on the visitor's GPU, eliminating inference API costs for each request. However, the trade-off is a significant initial download (model weights, libraries) and reliance on WebGPU support, making it only practical for repeated tasks on capable devices. The real value is for bounded, reusable features like text rewriting, where the setup cost amortizes over many uses.
Key points
WebLLM runs inference in the browser using the visitor's GPU via WebGPU, avoiding server-side API charges.
First load requires downloading model weights from Hugging Face and compiled libraries, which can take substantial time.
Cached model assets via Browser Cache API allow reuse, but the runtime still needs to prepare for execution.
Quantization reduces model size (e.g., from 2 GB to 0.5 GB for a 1B parameter model) but can affect output quality.
WebLLM supports running inference in a web worker to keep the UI responsive during generation.
Tools mentioned
Techniques
- quantization
- prefill
- web workers
- structured output
- hybrid inference
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Usually, when someone uses AI on your website, their message travels to a server, and another request gets added to your bill. Web LLM can move that generation onto the visitor's own GPU.
So, how does a browser run an AI model, and what are you giving up? The useful promise is no inference API bill for requests that run locally. Your website still needs hosting, the model needs
delivering, and somebody's computer still does the work. Take a writing assistant that rewrites selected text. Someone highlights a sentence, clicks rewrite, and gets a version to review.
Each use adds another generation request. In the usual cloud setup, the browser sends that selection to your back end. Your back end calls an inference provider, which runs the model
and sends the result back along the same route. That provider can charge for the input and output. A popular rewrite button can become a recurring expense. Web LLM changes where that request
finishes. Your website provides the editor and loads a supported model. Once it's ready, the selection goes to an inference engine inside the browser,
which returns an answer without calling your inference provider. If the website isn't paying for the computation, who is, and is that trade worth it? Follow the rewrite button through its first
visit, when the browser version has to prepare the model. Before it can answer, the browser needs the model. The model's learned parameters are stored in files called weights. Web LLM also needs a
compatible model library, the tokenizer that turns text into model inputs, and configuration telling those pieces how to work together. Those assets arrive over the internet. In the project's
standard configuration, model files come from Hugging Face, while compiled model libraries have their own URLs. A developer can choose other locations. The model wasn't already hiding inside
the browser. Web LLM downloads missing assets and prepares the runtime. Its documentation warns that a first load can take substantial time. Show loading progress before someone wonders
whether the rewrite button has broken. By default, Web LLM caches model assets using the browser's cache API. That's storage managed by the browser on the visitor's device. So, a later load can
reuse those files instead of downloading them again. It still needs to prepare the model for execution, so cache doesn't mean instantly ready. Think of the cached weights as a book on a shelf.
To use them, the runtime still has to open the book and put the working material on the desk. Here, the shelf is browser storage and the desk is working memory. They're different resources with
different limits. Browser storage also isn't a permanent installation promise. People can clear site data and browsers can evict stored data when space is tight. Your app needs to handle a
missing cache, and if you want the whole website to reopen offline, its page and application code need an offline loading strategy, too. To test offline generation, load the app and model,
disconnect the network, then submit a fresh rewrite. Keep the original and the new answer visible. An old answer in a chat window wouldn't prove generation happened offline. Once loaded, the model
doesn't need an internet round trip for each new token. So, what turns a sentence into a different sentence inside a webpage? There are three main parts to keep track of. The weights
contain the learned parameters. Web LLM coordinates the inference runtime. WebGPU gives that runtime access to the device's GPU for the numerical work. The page supplies the instruction and shows
the response as it arrives. WebAssembly helps with CPU work in the runtime. It's a way to run compiled code inside the browser. While WebGPU handles GPU operations, the browser is providing the
execution environment. It isn't supplying a language model of its own. The particular model still has to be supported by the engine. For example, select this sentence: We might ship on
Friday if the tests pass. Ask the assistant to make it clearer while preserving its meaning. That last instruction matters. A rewrite that promises a Friday release has changed
the product's message, even if it sounds more confident. First, the runtime turns the instruction and selected sentence into tokens, the chunks of text the model works with. It processes that
input using the loaded weights. That initial processing is often called prefill. It prepares the state the model needs to start generating a response. Then the runtime runs the model to
produce a next token, adds it to the response, and continues from that updated state. Web GPU carries the GPU operations to the visitor's hardware. Text appears in the editor as those
generated pieces return until generation reaches a stopping condition or its output limit. The weights stay available while that happens. The browser isn't downloading a fresh model for every
word. It's repeatedly applying the same learned parameters to an input and a growing response. That's why moving the weights onto the device can remove the server call from this part of the
feature. Heavy processing still has to coexist with typing, scrolling, and clicking. Web LLM supports running its engine in a web worker, a separate thread inside the browser. The interface
sends the worker a request. The worker manages generation, and messages bring the output back to the page. That separation helps keep the interface responsive. It doesn't create a second
GPU or reduce the model's memory needs. You can give the user a working cancel button while generation is running, but the computer still has a real job to finish or stop. Now the useful part is
deciding which jobs deserve that download. Rewriting a short selection is a reasonable task to evaluate because the input is bounded and the user can compare the suggestion with the
original. Keep the original available and let the user accept or reject the change. For our Friday sentence, check that the rewrite preserves uncertainty and the condition about passing tests.
Might and if aren't clutter to remove. Turning them into a firm promise fails this test. The same editor could summarize a selected passage. Evaluate whether the summary keeps the
important facts, exceptions, and names. It only receives the text your application gives it, so selecting one paragraph doesn't automatically give the model the rest of the document or
current information from the web. You could also ask it to classify a message, perhaps as a bug report or a feature request. Check the chosen label against examples you've reviewed. If you
need a fixed response format, Web LLM supports structured output, but a valid format doesn't guarantee the classification is correct. Build a collection of real inputs, including
awkward ones, and decide what counts as a useful answer. Run the same collection when you change the model or its settings. Model size sets a limit on how easy that will be to deliver. A larger
model may offer capabilities you need, but the visitor has to download its files and fit its working state into memory. Shrinking the weights can help. That's where quantization comes in,
storing parameters with fewer bits of precision. Take an illustrative model with 1 billion parameters. At 16 bits per parameter, the raw weights account for 2 GB. At 4 bits, that arithmetic
becomes half a gigabyte. That's a storage calculation before extra metadata and runtime memory. It isn't a measurement of a particular Web LLM model. The actual memory requirement
also includes working buffers and the state used while processing the prompt and generating text. For attention models, keeping more context can increase that state. So, a download that
fits on disk can still fail when the runtime tries to allocate memory for it. Lower precision can affect output quality, too. The file getting smaller doesn't tell you whether the Friday
sentence still means the same thing. Recheck the task after changing the model's precision. Otherwise, you might save memory by making the feature less useful, which is a fairly expensive way
to save money. Compatibility is another check to make before downloading a large file. Web GPU needs a suitable browser and device environment, and production pages
need a secure context, normally HTTPS. Detecting the API is only the start. The request for a usable GPU adapter can still fail. The model can also require capabilities
that the available adapter doesn't provide. Web LLM's model records include requirements such as GPU features alongside memory estimates. Treat those estimates as guidance, then
test loading and generation on the devices your visitors actually use. The user's machine pays in memory, processing time, and energy. That can mean battery use in competition with
other work. Measure these costs on representative devices with the model you intend to ship. But your side of the bill changes, too. Delivering a large model file has bandwidth costs under the
hosting arrangement you choose. Model updates can require another download. If visitors only need one short rewrite, the setup cost may be a poor exchange for avoiding that single cloud request.
Privacy needs the same attention to the whole application. The inference engine can process the selection locally, but analytics, error reporting, save documents, or a cloud
fallback can still send text away. Inspect those routes before saying that a person's writing stays on their device. A network inspection should include loading, generation, and what
happens afterward. An offline answer shows that this generation didn't need the network. It doesn't prove the app will never upload anything when connectivity returns. Keep that claim as
narrow as the evidence you actually collected. To evaluate the feature, separate the first visit from a return visit. On the first, record the download and the time until the assistant is
ready. On the return visit, check what gets reused. Then measure the wait for the first visible output and the time until a complete useful rewrite. Check quality alongside timing, and keep
failed loads in the results. Give unsupported devices a clear outcome, and preserve the selected text if initialization fails. For this writing assistant, local
inference is worth trying when the chosen model passes the task tests, the device can run it comfortably, and repeated use makes the download worthwhile. Cloud inference makes sense
when the needed capability or device support isn't there. A hybrid can offer both routes explicitly. If local loading fails, offer a server option and explain that it sends the selection off the
device. Let the user choose that route before the text moves. You can keep the editor useful while respecting the privacy choice made at the rewrite button.
Local generations don't create an inference provider charge. The visitor supplies the computing resources and accepts the download. Start with one bounded task, test its answers and its
device costs, then choose where that work belongs.