[NFBCS] Local AI
Ty Littlefield
tyler at tysdomain.com
Thu May 21 07:01:15 UTC 2026
I agree with this. I bought a data center quality GPU for fun and
testing. It's a referb, and I have 128 gb ram and I'm running everything
off of em.2 drives. Even with that much ram, context windows don't last
as long as the fronteer models. I suspect that the fronteer models are
cramming everything into vector db type data storage services and using
something to prefetch data when needed, but I could be very wrong there.
Just in raw performance, I got a lot of this stuff cheap and on sale. It
would cost probably 3x the amount if not more right now. You need a
minimum of 64 gb ram, and that's on the low end.
It's worth doing, but you might do better at these current prices buying
a beefed up Mac server vs trying to build your own system.
On 5/20/2026 11:35 PM, Paul York via NFBCS wrote:
> I've been knee deep in local llm setup for the better part of the last
> two weeks. To give you perspective on my hardware, I'm playing with
> two moderately beefy "consumer" machines: a Windows 11-based Ultra 7,
> 64GB RAM, RTX 4070 w/ 12GB VRAM and a linux-based Ryzen AI 9 HX370
> mini pc with 64GB RAM (both bought before prices went bonkers thankfully).
>
> TLDR: I'm keeping my Claude and Gemini subscriptions.
>
> I think a longer discussion will hinge on what you want to do with it.
> Are you programming? Running OpenClaw/agentic stuff? Just chatting?
> Creating documents and presentations? Doing NotebookLM kind of things?
> Because here's the deal. After a LOT of tweaking, I'm getting:
>
> * around 25 tokens per second output on my iGPU (Ryzen) using some
> pretty high quality models (Qwen 3.6 35b and Gemma 4 26b) by
> pushing the VRAM up to 48GB.
> * anywhere between 65 and 95 tokens per second output on my RTX GPU
> using much lower quality models (Qwen 3.5 9b and Gemma 4 4b).
>
> In both cases, if I don't take steps to optimize the model such that
> it stays 100% in VRAM, it slows to an entirely unusable rate.
>
> UP FRONT WARNING--I'm a noob with this, so take my explanation with a
> grain of salt.
>
> What does that actually mean? Well especially if you use a "reasoning"
> model like Qwen, then a simple query response (like "tell me a funny
> dad joke") can take up to a minute to respond. This is because
> approximately every word of every "thought" is an output token. It
> "talks to itself" until if decides it has found a reasonable answer.
> And it adds up quick. Here are some basic results for this exact query
> on all 4 models / hardware:
>
> * Qwen on RTX 4070: required 1200 tokens and 18 seconds to respond
> * Gemma on RTX 4070: required 300 tokens and 3.5 seconds to respond
> * Qwen on iGPU: required 630 tokens and 23 seconds to respond
> * Gemma on iGPU: required 450 tokens and took 18 seconds to respond
>
> Again that's moderately beefy hardware and a lot of tweaking. But I
> could also do far better if I accepted much dumber models. Which may
> be just find for basic agentic work. But much less good for coding or
> reasoned synthesis. And the "smartest" model took 23 seconds to reason
> through a dad joke. 7 seconds just to figure out how to respond to
> "hello". Working on truly complex reasoning can take a bathroom+coffee
> break to give you back results.
>
> Note too that this doesn't take into account context size and context
> caching. Context is the LLM's active memory. Models have maximums (I
> think they are tuned for these sizes). But in most cases you'll likely
> have to accept something lower. However, to be even moderately useful
> for much of anything, you can't go terribly low. Coding tools and
> agentic tools just blow up if they can't remember things from one
> thought to the next.
>
> The numbers I'm getting on the RTX are decent. Almost usable. BUT the
> context sizes to achieve that make it basically unusable for the kind
> of work I want to do. If I bump up the context window to a usable
> level with these models, I leak out into RAM (far slower than VRAM)
> and my performance tanks to unusable levels (like < 5-10 tps...at
> least 80-90% or more slower).
>
> The iGPU with huge VRAM is slower than the dedicated GPU, but because
> I can crank up the context window, they actually become usable for
> what I want to use them for. However...speed. Claude Sonnet or Gemini
> Flash are easily 10x faster at everything. And more like 20x-30x
> faster for most reasoning work. So at some point it's a question of
> how much you value your time.
>
> I will be using local models for some basic stuff, I think. I'm
> starting down a personal knowledge management path with AnythingLLM or
> something similar. I think it'll pair perfectly with this. And I'll
> likely find more ways to leverage it. But I won't be abandoning the
> big boys any time soon.
>
> And sadly, although your Ultra 7 w/ 32GB RAM is an awesome PC, I fear
> your experience with local llms for anything other than
> experimentation and learning will prove frustratingly slow. And with
> prices the way they are right now, getting your PC spec'd to perform
> moderately well will certainly cost around the same as a full year of
> one of the "ultimate" plans.
>
> Hope this was helpful. And that I didn't show my ignorance too badly.
>
> Best,
> Paul York
>
> On Wed, May 20, 2026 at 11:35 PM Lewis Wood via NFBCS
> <nfbcs at nfbnet.org> wrote:
>
> I am currently learning as well.
>
> I am now doing Ollama playlist lessons #2 currently.
>
> https://www.youtube.com/playlist?list=PLvsHpqLkpw0fIT-WbjY-xBRxTftjwiTLB
>
> I did my initial research on Lm Studio before I learned about
> Ollama CLI
>
> This was my first Lm Studio and it was an excellent one regarding
> resources, models, agents, etc. Even discussed how to load
> partial in differing areas gpu and ddr.
>
> https://www.youtube.com/watch?v=UngVdAsQEiU
>
> You can search youtube “lm studio”
>
> Lewis Wood
>
> *From:*NFBCS <nfbcs-bounces at nfbnet.org> *On Behalf Of *Joe Orozco
> via NFBCS
> *Sent:* Wednesday, May 20, 2026 10:08 PM
> *To:* 'NFB in Computer Science Mailing List' <nfbcs at nfbnet.org>
> *Cc:* Joe Orozco <jsorozco at gmail.com>
> *Subject:* [NFBCS] Local AI
>
> Hello,
>
> With Google following in Claude’s footsteps in terms of usage
> restrictions, can anyone speak to their experience using local LLM
> options? I’m looking at Jemma 4 and trying to understand how
> accessible this route might be with JAWS on Windows.
>
> I’m on a fairly decent machine: 32 GB RAM, Ultra 7 processor, 4 TB
> SSD. I see they’re recommending GPU for some of the more robust
> models, but I want to think most of what I’m doing shouldn’t
> require gaming machine specs. If you beg to differ though, let me
> know.
>
> If anyone can speak to Jemma alternatives, I’d also be interested.
> I don’t think I’ll suspend my subscriptions, but with these usage
> limitations feeling like the new standard, I want to spread my
> usage a little so that I don’t feel like I need to be hitting the
> top subscriptions just to get more mileage out of the five-hour
> increments.
>
> Thanks in advance for any tips,
>
> Joe
>
> --
>
> Joe Orozco: Your Message, My Mission
>
> https://joeorozco.com/services/
>
> _______________________________________________
> NFBCS mailing list
> NFBCS at nfbnet.org
> http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
> To unsubscribe, change your list options or get your account info
> for NFBCS:
> http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/paul%40yorkfamily.com
>
>
> _______________________________________________
> NFBCS mailing list
> NFBCS at nfbnet.org
> http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
> To unsubscribe, change your list options or get your account info for NFBCS:
> http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/tyler%40tysdomain.com
--
*Ty Littlefield (he/him/his)*
* My Website <https://tysdomain.com>|
* Linkedin <https://www.linkedin.com/in/ty-lerlittlefield/>|
* Github <https://github.com/tlfdev>
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://nfbnet.org/pipermail/nfbcs_nfbnet.org/attachments/20260521/881a11f2/attachment.htm>
More information about the NFBCS
mailing list