[NFBCS] Local AI
Chris Nestrud
ccn at chrisnestrud.com
Thu May 21 17:06:37 UTC 2026
A little off-topic, but I've had good experience with runpod.io for
short-term tasks needing substantial GPU VRAM. Also Open Router (openrouter.ai)
provides API-based access to many models for small percentage increase.
Chris
On Thu, May 21, 2026 at 10:33:41AM -0600, Ty Littlefield via NFBCS wrote:
> My video card isn't 128 gb, I wouldn't have done that (also don't know that
> those exist, even in the A100 models). Ram was cheaper when I bought it, so
> it sounds more impressive than it really is. now if I say I bought 128 gb
> for my home server people think I'm the musky Elmo because data centers are
> gobbling up ram and sending prices through the roof.
>
> *Ty Littlefield (he/him/his)*
>
> * From Bytes to Bites <https://tysdomain.com>|
> * Linkedin <https://www.linkedin.com/in/ty-lerlittlefield/>|
> * Github <https://github.com/tlfdev>
>
> On 5/21/2026 7:05 AM, Lewis Wood via NFBCS wrote:
> >
> > WOW, You guys make my proposed little purchase of a 16gb nvidia card
> > very small. 128gb are very expensive in the very least. I refuse to go
> > above $1,000.00.
> >
> > I currently have 32gb ddr Windows Lenova computer with a 4gb nvidia
> > t1000 card.
> >
> > I figured adding a 166gb nvidia card would give me 20gb of vram total.
> >
> > I may not be able to run the very 70b models, but I think the smaller
> > should do for 20gb vram and 32gb ddr ram.
> >
> > I am thinking it will run most, I can always offshore some larger tasks
> > to online providers (clawd or co-pilot).
> >
> > I am just a newbie so any advise is quite welcome.
> >
> > Lewis Wood
> >
> > *From:*NFBCS <nfbcs-bounces at nfbnet.org> *On Behalf Of *Elijah Massey via
> > NFBCS
> > *Sent:* Thursday, May 21, 2026 3:04 AM
> > *To:* nfbcs at nfbnet.org
> > *Cc:* Elijah Massey <emassey0135 at gmail.com>
> > *Subject:* Re: [NFBCS] Local AI
> >
> > I tried running local LLMs for a while on my Ryzen AI 9 HX 370 mini PC
> > with 32 GB of RAM, on Arch Linux with Ollama. It was very slow at first
> > until I got the Ollama Vulkan backend working and it started using the
> > iGMU, and then it became pretty fast (at least 20 or 30 tokens per
> > second I think). I was using 13B models mostly because that's what fit
> > in my 32 GB of RAM. I did not need to increase my VRAM for it to be fast
> > like this and it was set to 4 GB by default, although I heard there is a
> > mechanism for dynamically increasing RAM that the iGPU can access
> > somehow. However, there were definitely limitations compared to the
> > cloud models. I noticed that the context window on the local models I
> > tried was always shorter so it would have trouble remembering earlier
> > messages in the conversation, and the cloud models just often created
> > better responses and better did what I asked. I think it would be better
> > if I have more RAM so I could run 70b models or larger. By the way, if
> > you want to run local LLMs, I think an Apple Silicon Mac with a lot of
> > RAM would be a great option, because of their unified memory. The GPU
> > and NPU can access all of the RAM which is very good for local AI. Also,
> > I would get the Max chips over the Pro can they have double the GPU
> > cores and support higher RAM. I'm waiting for the M6 Max MacBook
> > Pro/Ultra to come out later this year and then I'll probably get that.
> >
> > Sent from my iPhone
> >
> >
> >
> > On May 21, 2026, at 02:02, Ty Littlefield via NFBCS
> > <nfbcs at nfbnet.org> wrote:
> >
> >
> >
> > I agree with this. I bought a data center quality GPU for fun and
> > testing. It's a referb, and I have 128 gb ram and I'm running
> > everything off of em.2 drives. Even with that much ram, context
> > windows don't last as long as the fronteer models. I suspect that
> > the fronteer models are cramming everything into vector db type
> > data storage services and using something to prefetch data when
> > needed, but I could be very wrong there.
> >
> > Just in raw performance, I got a lot of this stuff cheap and on
> > sale. It would cost probably 3x the amount if not more right now.
> > You need a minimum of 64 gb ram, and that's on the low end.
> >
> > It's worth doing, but you might do better at these current prices
> > buying a beefed up Mac server vs trying to build your own system.
> >
> > On 5/20/2026 11:35 PM, Paul York via NFBCS wrote:
> >
> > I've been knee deep in local llm setup for the better part of
> > the last two weeks. To give you perspective on my hardware,
> > I'm playing with two moderately beefy "consumer" machines: a
> > Windows 11-based Ultra 7, 64GB RAM, RTX 4070 w/ 12GB VRAM and
> > a linux-based Ryzen AI 9 HX370 mini pc with 64GB RAM (both
> > bought before prices went bonkers thankfully).
> >
> > TLDR: I'm keeping my Claude and Gemini subscriptions.
> >
> > I think a longer discussion will hinge on what you want to do
> > with it. Are you programming? Running OpenClaw/agentic stuff?
> > Just chatting? Creating documents and presentations? Doing
> > NotebookLM kind of things? Because here's the deal. After a
> > LOT of tweaking, I'm getting:
> >
> > * around 25 tokens per second output on my iGPU (Ryzen)
> > using some pretty high quality models (Qwen 3.6 35b and
> > Gemma 4 26b) by pushing the VRAM up to 48GB.
> > * anywhere between 65 and 95 tokens per second output on my
> > RTX GPU using much lower quality models (Qwen 3.5 9b and
> > Gemma 4 4b).
> >
> > In both cases, if I don't take steps to optimize the model
> > such that it stays 100% in VRAM, it slows to an entirely
> > unusable rate.
> >
> > UP FRONT WARNING--I'm a noob with this, so take my explanation
> > with a grain of salt.
> >
> > What does that actually mean? Well especially if you use a
> > "reasoning" model like Qwen, then a simple query response
> > (like "tell me a funny dad joke") can take up to a minute to
> > respond. This is because approximately every word of every
> > "thought" is an output token. It "talks to itself" until if
> > decides it has found a reasonable answer. And it adds up
> > quick. Here are some basic results for this exact query on all
> > 4 models / hardware:
> >
> > * Qwen on RTX 4070: required 1200 tokens and 18 seconds to
> > respond
> > * Gemma on RTX 4070: required 300 tokens and 3.5 seconds to
> > respond
> > * Qwen on iGPU: required 630 tokens and 23 seconds to respond
> > * Gemma on iGPU: required 450 tokens and took 18 seconds to
> > respond
> >
> > Again that's moderately beefy hardware and a lot of tweaking.
> > But I could also do far better if I accepted much dumber
> > models. Which may be just find for basic agentic work. But
> > much less good for coding or reasoned synthesis. And the
> > "smartest" model took 23 seconds to reason through a dad joke.
> > 7 seconds just to figure out how to respond to "hello".
> > Working on truly complex reasoning can take a bathroom+coffee
> > break to give you back results.
> >
> > Note too that this doesn't take into account context size and
> > context caching. Context is the LLM's active memory. Models
> > have maximums (I think they are tuned for these sizes). But in
> > most cases you'll likely have to accept something lower.
> > However, to be even moderately useful for much of anything,
> > you can't go terribly low. Coding tools and agentic tools just
> > blow up if they can't remember things from one thought to the
> > next.
> >
> > The numbers I'm getting on the RTX are decent. Almost usable.
> > BUT the context sizes to achieve that make it basically
> > unusable for the kind of work I want to do. If I bump up the
> > context window to a usable level with these models, I leak out
> > into RAM (far slower than VRAM) and my performance tanks to
> > unusable levels (like < 5-10 tps...at least 80-90% or more
> > slower).
> >
> > The iGPU with huge VRAM is slower than the dedicated GPU, but
> > because I can crank up the context window, they actually
> > become usable for what I want to use them for.
> > However...speed. Claude Sonnet or Gemini Flash are easily 10x
> > faster at everything. And more like 20x-30x faster for most
> > reasoning work. So at some point it's a question of how much
> > you value your time.
> >
> > I will be using local models for some basic stuff, I think.
> > I'm starting down a personal knowledge management path with
> > AnythingLLM or something similar. I think it'll pair perfectly
> > with this. And I'll likely find more ways to leverage it. But
> > I won't be abandoning the big boys any time soon.
> >
> > And sadly, although your Ultra 7 w/ 32GB RAM is an awesome PC,
> > I fear your experience with local llms for anything other than
> > experimentation and learning will prove frustratingly slow.
> > And with prices the way they are right now, getting your PC
> > spec'd to perform moderately well will certainly cost around
> > the same as a full year of one of the "ultimate" plans.
> >
> > Hope this was helpful. And that I didn't show my ignorance too
> > badly.
> >
> > Best,
> >
> > Paul York
> >
> > On Wed, May 20, 2026 at 11:35 PM Lewis Wood via NFBCS
> > <nfbcs at nfbnet.org> wrote:
> >
> > I am currently learning as well.
> >
> > I am now doing Ollama playlist lessons #2 currently.
> >
> > https://www.youtube.com/playlist?list=PLvsHpqLkpw0fIT-WbjY-xBRxTftjwiTLB
> >
> > I did my initial research on Lm Studio before I learned
> > about Ollama CLI
> >
> > This was my first Lm Studio and it was an excellent one
> > regarding resources, models, agents, etc. Even discussed
> > how to load partial in differing areas gpu and ddr.
> >
> > https://www.youtube.com/watch?v=UngVdAsQEiU
> >
> > You can search youtube “lm studio”
> >
> > Lewis Wood
> >
> > *From:*NFBCS <nfbcs-bounces at nfbnet.org> *On Behalf Of *Joe
> > Orozco via NFBCS
> > *Sent:* Wednesday, May 20, 2026 10:08 PM
> > *To:* 'NFB in Computer Science Mailing List'
> > <nfbcs at nfbnet.org>
> > *Cc:* Joe Orozco <jsorozco at gmail.com>
> > *Subject:* [NFBCS] Local AI
> >
> > Hello,
> >
> > With Google following in Claude’s footsteps in terms of
> > usage restrictions, can anyone speak to their experience
> > using local LLM options? I’m looking at Jemma 4 and trying
> > to understand how accessible this route might be with JAWS
> > on Windows.
> >
> > I’m on a fairly decent machine: 32 GB RAM, Ultra 7
> > processor, 4 TB SSD. I see they’re recommending GPU for
> > some of the more robust models, but I want to think most
> > of what I’m doing shouldn’t require gaming machine specs.
> > If you beg to differ though, let me know.
> >
> > If anyone can speak to Jemma alternatives, I’d also be
> > interested. I don’t think I’ll suspend my subscriptions,
> > but with these usage limitations feeling like the new
> > standard, I want to spread my usage a little so that I
> > don’t feel like I need to be hitting the top subscriptions
> > just to get more mileage out of the five-hour increments.
> >
> > Thanks in advance for any tips,
> >
> > Joe
> >
> > --
> >
> > Joe Orozco: Your Message, My Mission
> >
> > https://joeorozco.com/services/
> >
> > _______________________________________________
> > NFBCS mailing list
> > NFBCS at nfbnet.org
> > http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
> > To unsubscribe, change your list options or get your
> > account info for NFBCS:
> > http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/paul%40yorkfamily.com
> >
> >
> >
> > _______________________________________________
> >
> > NFBCS mailing list
> >
> > NFBCS at nfbnet.org
> >
> > http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
> >
> > To unsubscribe, change your list options or get your account info for NFBCS:
> >
> > http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/tyler%40tysdomain.com
> >
> > --
> >
> > *Ty Littlefield (he/him/his)*
> >
> > * My Website <https://tysdomain.com>|
> > * Linkedin <https://www.linkedin.com/in/ty-lerlittlefield/>|
> > * Github <https://github.com/tlfdev>
> >
> > _______________________________________________
> > NFBCS mailing list
> > NFBCS at nfbnet.org
> > http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
> > To unsubscribe, change your list options or get your account info
> > for NFBCS:
> > http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/emassey0135%40gmail.com
> >
> >
> > _______________________________________________
> > NFBCS mailing list
> > NFBCS at nfbnet.org
> > http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
> > To unsubscribe, change your list options or get your account info for NFBCS:
> > http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/tyler%40tysdomain.com
> _______________________________________________
> NFBCS mailing list
> NFBCS at nfbnet.org
> http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
> To unsubscribe, change your list options or get your account info for NFBCS:
> http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/ccn%40chrisnestrud.com
More information about the NFBCS
mailing list