[NFBCS] Local AI

Ty Littlefield tyler at tysdomain.com
Thu May 21 16:33:41 UTC 2026


My video card isn't 128 gb, I wouldn't have done that (also don't know 
that those exist, even in the A100 models). Ram was cheaper when I 
bought it, so it sounds more impressive than it really is. now if I say 
I bought 128 gb for my home server people think I'm the musky Elmo 
because data centers are gobbling up ram and sending prices through the 
roof.

*Ty Littlefield (he/him/his)*

  * From Bytes to Bites <https://tysdomain.com>|
  * Linkedin <https://www.linkedin.com/in/ty-lerlittlefield/>|
  * Github <https://github.com/tlfdev>

On 5/21/2026 7:05 AM, Lewis Wood via NFBCS wrote:
>
> WOW, You guys make my proposed little purchase of a 16gb nvidia card 
> very small. 128gb are very expensive in the very least. I refuse to go 
> above $1,000.00.
>
> I currently have 32gb ddr Windows Lenova computer with a 4gb nvidia 
> t1000 card.
>
> I figured adding a 166gb nvidia card would give me 20gb of vram total.
>
> I may not be able to run the very 70b models, but I think the smaller 
> should do for 20gb vram and 32gb ddr ram.
>
> I am thinking it will run most, I can always offshore some larger 
> tasks to online providers (clawd or co-pilot).
>
> I am just a newbie so any advise is quite welcome.
>
> Lewis Wood
>
> *From:*NFBCS <nfbcs-bounces at nfbnet.org> *On Behalf Of *Elijah Massey 
> via NFBCS
> *Sent:* Thursday, May 21, 2026 3:04 AM
> *To:* nfbcs at nfbnet.org
> *Cc:* Elijah Massey <emassey0135 at gmail.com>
> *Subject:* Re: [NFBCS] Local AI
>
> I tried running local LLMs for a while on my Ryzen AI 9 HX 370 mini PC 
> with 32 GB of RAM, on Arch Linux with Ollama. It was very slow at 
> first until I got the Ollama Vulkan backend working and it started 
> using the iGMU, and then it became pretty fast (at least 20 or 30 
> tokens per second I think). I was using 13B models mostly because 
> that's what fit in my 32 GB of RAM. I did not need to increase my VRAM 
> for it to be fast like this and it was set to 4 GB by default, 
> although I heard there is a mechanism for dynamically increasing RAM 
> that the iGPU can access somehow. However, there were definitely 
> limitations compared to the cloud models. I noticed that the context 
> window on the local models I tried was always shorter so it would have 
> trouble remembering earlier messages in the conversation, and the 
> cloud models just often created better responses and better did what I 
> asked. I think it would be better if I have more RAM so I could run 
> 70b models or larger. By the way, if you want to run local LLMs, I 
> think an Apple Silicon Mac with a lot of RAM would be a great option, 
> because of their unified memory. The GPU and NPU can access all of the 
> RAM which is very good for local AI. Also, I would get the Max chips 
> over the Pro can they have double the GPU cores and support higher 
> RAM. I'm waiting for the M6 Max MacBook Pro/Ultra to come out later 
> this year and then I'll probably get that.
>
> Sent from my iPhone
>
>
>
>     On May 21, 2026, at 02:02, Ty Littlefield via NFBCS
>     <nfbcs at nfbnet.org> wrote:
>
>     
>
>     I agree with this. I bought a data center quality GPU for fun and
>     testing. It's a referb, and I have 128 gb ram and I'm running
>     everything off of em.2 drives. Even with that much ram, context
>     windows don't last as long as the fronteer models. I suspect that
>     the fronteer models are cramming everything into vector db type
>     data storage services and using something to prefetch data when
>     needed, but I could be very wrong there.
>
>     Just in raw performance, I got a lot of this stuff cheap and on
>     sale. It would cost probably 3x the amount if not more right now.
>     You need a minimum of 64 gb ram, and that's on the low end.
>
>     It's worth doing, but you might do better at these current prices
>     buying a beefed up Mac server vs trying to build your own system.
>
>     On 5/20/2026 11:35 PM, Paul York via NFBCS wrote:
>
>         I've been knee deep in local llm setup for the better part of
>         the last two weeks. To give you perspective on my hardware,
>         I'm playing with two moderately beefy "consumer" machines: a
>         Windows 11-based Ultra 7, 64GB RAM, RTX 4070 w/ 12GB VRAM and
>         a linux-based Ryzen AI 9 HX370 mini pc with 64GB RAM (both
>         bought before prices went bonkers thankfully).
>
>         TLDR: I'm keeping my Claude and Gemini subscriptions.
>
>         I think a longer discussion will hinge on what you want to do
>         with it. Are you programming? Running OpenClaw/agentic stuff?
>         Just chatting? Creating documents and presentations? Doing
>         NotebookLM kind of things? Because here's the deal. After a
>         LOT of tweaking, I'm getting:
>
>           * around 25 tokens per second output on my iGPU (Ryzen)
>             using some pretty high quality models (Qwen 3.6 35b and
>             Gemma 4 26b) by pushing the VRAM up to 48GB.
>           * anywhere between 65 and 95 tokens per second output on my
>             RTX GPU using much lower quality models (Qwen 3.5 9b and
>             Gemma 4 4b).
>
>         In both cases, if I don't take steps to optimize the model
>         such that it stays 100% in VRAM, it slows to an entirely
>         unusable rate.
>
>         UP FRONT WARNING--I'm a noob with this, so take my explanation
>         with a grain of salt.
>
>         What does that actually mean? Well especially if you use a
>         "reasoning" model like Qwen, then a simple query response
>         (like "tell me a funny dad joke") can take up to a minute to
>         respond. This is because approximately every word of every
>         "thought" is an output token. It "talks to itself" until if
>         decides it has found a reasonable answer. And it adds up
>         quick. Here are some basic results for this exact query on all
>         4 models / hardware:
>
>           * Qwen on RTX 4070: required 1200 tokens and 18 seconds to
>             respond
>           * Gemma on RTX 4070: required 300 tokens and 3.5 seconds to
>             respond
>           * Qwen on iGPU: required 630 tokens and 23 seconds to respond
>           * Gemma on iGPU: required 450 tokens and took 18 seconds to
>             respond
>
>         Again that's moderately beefy hardware and a lot of tweaking.
>         But I could also do far better if I accepted much dumber
>         models. Which may be just find for basic agentic work. But
>         much less good for coding or reasoned synthesis. And the
>         "smartest" model took 23 seconds to reason through a dad joke.
>         7 seconds just to figure out how to respond to "hello".
>         Working on truly complex reasoning can take a bathroom+coffee
>         break to give you back results.
>
>         Note too that this doesn't take into account context size and
>         context caching. Context is the LLM's active memory. Models
>         have maximums (I think they are tuned for these sizes). But in
>         most cases you'll likely have to accept something lower.
>         However, to be even moderately useful for much of anything,
>         you can't go terribly low. Coding tools and agentic tools just
>         blow up if they can't remember things from one thought to the
>         next.
>
>         The numbers I'm getting on the RTX are decent. Almost usable.
>         BUT the context sizes to achieve that make it basically
>         unusable for the kind of work I want to do. If I bump up the
>         context window to a usable level with these models, I leak out
>         into RAM (far slower than VRAM) and my performance tanks to
>         unusable levels (like < 5-10 tps...at least 80-90% or more
>         slower).
>
>         The iGPU with huge VRAM is slower than the dedicated GPU, but
>         because I can crank up the context window, they actually
>         become usable for what I want to use them for.
>         However...speed. Claude Sonnet or Gemini Flash are easily 10x
>         faster at everything. And more like 20x-30x faster for most
>         reasoning work. So at some point it's a question of how much
>         you value your time.
>
>         I will be using local models for some basic stuff, I think.
>         I'm starting down a personal knowledge management path with
>         AnythingLLM or something similar. I think it'll pair perfectly
>         with this. And I'll likely find more ways to leverage it. But
>         I won't be abandoning the big boys any time soon.
>
>         And sadly, although your Ultra 7 w/ 32GB RAM is an awesome PC,
>         I fear your experience with local llms for anything other than
>         experimentation and learning will prove frustratingly slow.
>         And with prices the way they are right now, getting your PC
>         spec'd to perform moderately well will certainly cost around
>         the same as a full year of one of the "ultimate" plans.
>
>         Hope this was helpful. And that I didn't show my ignorance too
>         badly.
>
>         Best,
>
>         Paul York
>
>         On Wed, May 20, 2026 at 11:35 PM Lewis Wood via NFBCS
>         <nfbcs at nfbnet.org> wrote:
>
>             I am currently learning as well.
>
>             I am now doing Ollama playlist lessons #2 currently.
>
>             https://www.youtube.com/playlist?list=PLvsHpqLkpw0fIT-WbjY-xBRxTftjwiTLB
>
>             I did my initial research on Lm Studio before I learned
>             about Ollama CLI
>
>             This was my first Lm Studio and it was an excellent one
>             regarding resources, models, agents, etc. Even discussed
>             how to load partial in differing areas gpu and ddr.
>
>             https://www.youtube.com/watch?v=UngVdAsQEiU
>
>             You can search youtube  “lm studio”
>
>             Lewis Wood
>
>             *From:*NFBCS <nfbcs-bounces at nfbnet.org> *On Behalf Of *Joe
>             Orozco via NFBCS
>             *Sent:* Wednesday, May 20, 2026 10:08 PM
>             *To:* 'NFB in Computer Science Mailing List'
>             <nfbcs at nfbnet.org>
>             *Cc:* Joe Orozco <jsorozco at gmail.com>
>             *Subject:* [NFBCS] Local AI
>
>             Hello,
>
>             With Google following in Claude’s footsteps in terms of
>             usage restrictions, can anyone speak to their experience
>             using local LLM options? I’m looking at Jemma 4 and trying
>             to understand how accessible this route might be with JAWS
>             on Windows.
>
>             I’m on a fairly decent machine: 32 GB RAM, Ultra 7
>             processor, 4 TB SSD. I see they’re recommending GPU for
>             some of the more robust models, but I want to think most
>             of what I’m doing shouldn’t require gaming machine specs.
>             If you beg to differ though, let me know.
>
>             If anyone can speak to Jemma alternatives, I’d also be
>             interested. I don’t think I’ll suspend my subscriptions,
>             but with these usage limitations feeling like the new
>             standard, I want to spread my usage a little so that I
>             don’t feel like I need to be hitting the top subscriptions
>             just to get more mileage out of the five-hour increments.
>
>             Thanks in advance for any tips,
>
>             Joe
>
>             --
>
>             Joe Orozco: Your Message, My Mission
>
>             https://joeorozco.com/services/
>
>             _______________________________________________
>             NFBCS mailing list
>             NFBCS at nfbnet.org
>             http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
>             To unsubscribe, change your list options or get your
>             account info for NFBCS:
>             http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/paul%40yorkfamily.com
>
>
>
>         _______________________________________________
>
>         NFBCS mailing list
>
>         NFBCS at nfbnet.org
>
>         http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
>
>         To unsubscribe, change your list options or get your account info for NFBCS:
>
>         http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/tyler%40tysdomain.com
>
>     -- 
>
>     *Ty Littlefield (he/him/his)*
>
>       * My Website <https://tysdomain.com>|
>       * Linkedin <https://www.linkedin.com/in/ty-lerlittlefield/>|
>       * Github <https://github.com/tlfdev>
>
>     _______________________________________________
>     NFBCS mailing list
>     NFBCS at nfbnet.org
>     http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
>     To unsubscribe, change your list options or get your account info
>     for NFBCS:
>     http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/emassey0135%40gmail.com
>
>
> _______________________________________________
> NFBCS mailing list
> NFBCS at nfbnet.org
> http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org
> To unsubscribe, change your list options or get your account info for NFBCS:
> http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/tyler%40tysdomain.com
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://nfbnet.org/pipermail/nfbcs_nfbnet.org/attachments/20260521/68db74f2/attachment.htm>


More information about the NFBCS mailing list