[NFBCS] Local AI
Paul York
paul at yorkfamily.com
Thu May 21 05:35:30 UTC 2026
I've been knee deep in local llm setup for the better part of the last two
weeks. To give you perspective on my hardware, I'm playing with two
moderately beefy "consumer" machines: a Windows 11-based Ultra 7, 64GB RAM,
RTX 4070 w/ 12GB VRAM and a linux-based Ryzen AI 9 HX370 mini pc with 64GB
RAM (both bought before prices went bonkers thankfully).
TLDR: I'm keeping my Claude and Gemini subscriptions.
I think a longer discussion will hinge on what you want to do with it. Are
you programming? Running OpenClaw/agentic stuff? Just chatting? Creating
documents and presentations? Doing NotebookLM kind of things? Because
here's the deal. After a LOT of tweaking, I'm getting:
- around 25 tokens per second output on my iGPU (Ryzen) using some
pretty high quality models (Qwen 3.6 35b and Gemma 4 26b) by pushing the
VRAM up to 48GB.
- anywhere between 65 and 95 tokens per second output on my RTX GPU
using much lower quality models (Qwen 3.5 9b and Gemma 4 4b).
In both cases, if I don't take steps to optimize the model such that it
stays 100% in VRAM, it slows to an entirely unusable rate.
UP FRONT WARNING--I'm a noob with this, so take my explanation with a grain
of salt.
What does that actually mean? Well especially if you use a "reasoning"
model like Qwen, then a simple query response (like "tell me a funny dad
joke") can take up to a minute to respond. This is because approximately
every word of every "thought" is an output token. It "talks to itself"
until if decides it has found a reasonable answer. And it adds up quick.
Here are some basic results for this exact query on all 4 models / hardware:
- Qwen on RTX 4070: required 1200 tokens and 18 seconds to respond
- Gemma on RTX 4070: required 300 tokens and 3.5 seconds to respond
- Qwen on iGPU: required 630 tokens and 23 seconds to respond
- Gemma on iGPU: required 450 tokens and took 18 seconds to respond
Again that's moderately beefy hardware and a lot of tweaking. But I could
also do far better if I accepted much dumber models. Which may be just
find for basic agentic work. But much less good for coding or reasoned
synthesis. And the "smartest" model took 23 seconds to reason through a dad
joke. 7 seconds just to figure out how to respond to "hello". Working on
truly complex reasoning can take a bathroom+coffee break to give you back
results.
Note too that this doesn't take into account context size and context
caching. Context is the LLM's active memory. Models have maximums (I think
they are tuned for these sizes). But in most cases you'll likely have to
accept something lower. However, to be even moderately useful for much of
anything, you can't go terribly low. Coding tools and agentic tools just
blow up if they can't remember things from one thought to the next.
The numbers I'm getting on the RTX are decent. Almost usable. BUT the
context sizes to achieve that make it basically unusable for the kind of
work I want to do. If I bump up the context window to a usable level with
these models, I leak out into RAM (far slower than VRAM) and my performance
tanks to unusable levels (like < 5-10 tps...at least 80-90% or more slower).
The iGPU with huge VRAM is slower than the dedicated GPU, but because I can
crank up the context window, they actually become usable for what I want to
use them for. However...speed. Claude Sonnet or Gemini Flash are easily 10x
faster at everything. And more like 20x-30x faster for most reasoning work.
So at some point it's a question of how much you value your time.
I will be using local models for some basic stuff, I think. I'm starting
down a personal knowledge management path with AnythingLLM or something
similar. I think it'll pair perfectly with this. And I'll likely find more
ways to leverage it. But I won't be abandoning the big boys any time soon.
And sadly, although your Ultra 7 w/ 32GB RAM is an awesome PC, I fear your
experience with local llms for anything other than experimentation and
learning will prove frustratingly slow. And with prices the way they are
right now, getting your PC spec'd to perform moderately well will certainly
cost around the same as a full year of one of the "ultimate" plans.
Hope this was helpful. And that I didn't show my ignorance too badly.
Best,
Paul York
On Wed, May 20, 2026 at 11:35 PM Lewis Wood via NFBCS <nfbcs at nfbnet.org>
wrote:
> I am currently learning as well.
>
>
>
> I am now doing Ollama playlist lessons #2 currently.
>
> https://www.youtube.com/playlist?list=PLvsHpqLkpw0fIT-WbjY-xBRxTftjwiTLB
>
>
>
>
>
> I did my initial research on Lm Studio before I learned about Ollama CLI
>
>
>
> This was my first Lm Studio and it was an excellent one regarding
> resources, models, agents, etc. Even discussed how to load partial in
> differing areas gpu and ddr.
>
> https://www.youtube.com/watch?v=UngVdAsQEiU
>
>
>
>
>
> You can search youtube “lm studio”
>
>
>
> Lewis Wood
>
>
>
>
>
>
>
> *From:* NFBCS <nfbcs-bounces at nfbnet.org> *On Behalf Of *Joe Orozco via
> NFBCS
> *Sent:* Wednesday, May 20, 2026 10:08 PM
> *To:* 'NFB in Computer Science Mailing List' <nfbcs at nfbnet.org>
> *Cc:* Joe Orozco <jsorozco at gmail.com>
> *Subject:* [NFBCS] Local AI
>
>
>
> Hello,
>
>
>
> With Google following in Claude’s footsteps in terms of usage
> restrictions, can anyone speak to their experience using local LLM options?
> I’m looking at Jemma 4 and trying to understand how accessible this route
> might be with JAWS on Windows.
>
>
>
> I’m on a fairly decent machine: 32 GB RAM, Ultra 7 processor, 4 TB SSD. I
> see they’re recommending GPU for some of the more robust models, but I want
> to think most of what I’m doing shouldn’t require gaming machine specs. If
> you beg to differ though, let me know.
>
>
>
> If anyone can speak to Jemma alternatives, I’d also be interested. I don’t
> think I’ll suspend my subscriptions, but with these usage limitations
> feeling like the new standard, I want to spread my usage a little so that I
> don’t feel like I need to be hitting the top subscriptions just to get more
> mileage out of the five-hour increments.
>
>
>
> Thanks in advance for any tips,
>
>
>
> Joe
>
>
>
> --
>
> Joe Orozco: Your Message, My Mission
>
> https://joeorozco.com/services/
>
>
> _______________________________________________
> NFBCS mailing list
> NFBCS at nfbnet.org
> http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
> To unsubscribe, change your list options or get your account info for
> NFBCS:
> http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/paul%40yorkfamily.com
>
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://nfbnet.org/pipermail/nfbcs_nfbnet.org/attachments/20260521/d3025936/attachment.htm>
More information about the NFBCS
mailing list