<div dir="ltr"><p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">Hi Paul, and all,</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">Thank you for the comprehensive response and subsequent discussion.</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">My goal is twofold: First, try to save a little money now
that the usage is moving more toward a metered standard. Right now I\u2019m using
Perplexity Pro for academic research and Gemini Pro for everyday queries. I
think the time saved has made both subscriptions a worthwhile investment. I
just recently began with Claude and like that I\u2019ve been able to train it to
write in my style and tone. Perhaps Gemini and Perplexity allow for similar
capability, but Claude has made this a bit more intuitive and has resulted in
excellent working drafts I can use to finalize what I\u2019m going for. I think I\u2019ll
likely end up biting the bullet and paying for a subscription there as well,
but the type of subscription is what leads to my second goal.</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">My second goal is to move more into the agentic space. There
are a set of repetitive tasks I would just assume dump onto a virtual
assistant. I was, still am, on the fence about learning Python to tackle these
tasks, and if I hesitate at all, it is because I fear diving into an area that does
not really contribute to my optimal productivity. Coding would be great in
terms of understanding the underlying method that created what I needed, but in
the end, a coder is not what I need to be a part of my core set of skills.</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">But, the idea of maintaining spreadsheets based on PDF files
received and using both to create reports and invoices is highly appealing.
This is working with data I would rather not share on an open web.</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">After all this, there are marketing, creative projects, and
other proprietary data that I would like to keep offline where possible. It\u2019s
not data so sensitive that I will absolutely refuse to put it through the
online systems, but if I could keep it offline and rely on a reasonably fast
response, that would be excellent.</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">Based on what I\u2019m reading here, it sounds like my offline
dream is not going to come true. My workstation may not be up for running an
LLM strong enough to engage with me in the way I need. If I\u2019m reading that incorrectly,
I\u2019m happy to be told otherwise.</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">To summarize, at this stage, I am not doing any coding. Although
I am tempted by the vibe coding trend, I understand it is always better to understand
the programming going on so I can fix the errors myself. I am mostly relying on
document analysis, research, and coaching to help turn my outlines into usable drafts.
I do some image generation and occasionally ask Gemini to review audio/video
files, but I\u2019m okay with keeping those firmly in the paid subscription buckets.</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">I hope this has made some kind of sense and hope that my
questions are helping other non-technical people figure out how to best maximize
the available options. The open weight options sound really good on paper, but
I don\u2019t want to go through the trouble of installing local models if my machine
won\u2019t make a substantial difference and if, in the end, I would be better off
making the best of the online models.</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">Very appreciative,</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif">Joe</p>
<p class="MsoNormal" style="margin:0in 0in 8pt;font-size:12pt;font-family:"Times New Roman",serif"> </p></div><br><div class="gmail_quote gmail_quote_container"><div dir="ltr" class="gmail_attr">On Thu, May 21, 2026 at 1:36\u202fAM Paul York via NFBCS <<a href="mailto:nfbcs@nfbnet.org">nfbcs@nfbnet.org</a>> wrote:<br></div><blockquote class="gmail_quote" style="margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"><div dir="ltr"><div class="gmail_default" style="font-size:large">I've been knee deep in local llm setup for the better part of the last two weeks. To give you perspective on my hardware, I'm playing with two moderately beefy "consumer" machines: a Windows 11-based Ultra 7, 64GB RAM, RTX 4070 w/ 12GB VRAM and a linux-based Ryzen AI 9 HX370 mini pc with 64GB RAM (both bought before prices went bonkers thankfully).</div><div class="gmail_default" style="font-size:large"><br></div><div class="gmail_default" style="font-size:large">TLDR: I'm keeping my Claude and Gemini subscriptions.</div><div class="gmail_default" style="font-size:large"><br></div><div class="gmail_default" style="font-size:large">I think a longer discussion will hinge on what you want to do with it. Are you programming? Running OpenClaw/agentic stuff? Just chatting? Creating documents and presentations? Doing NotebookLM kind of things? Because here's the deal. After a LOT of tweaking, I'm getting:</div><div class="gmail_default" style="font-size:large"><ul><li>around 25 tokens per second output on my iGPU (Ryzen) using some pretty high quality models (Qwen 3.6 35b and Gemma 4 26b) by pushing the VRAM up to 48GB.</li><li>anywhere between 65 and 95 tokens per second output on my RTX GPU using much lower quality models (Qwen 3.5 9b and Gemma 4 4b).</li></ul><div>In both cases, if I don't take steps to optimize the model such that it stays 100% in VRAM, it slows to an entirely unusable rate.</div><div><br></div><div>UP FRONT WARNING--I'm a noob with this, so take my explanation with a grain of salt.</div><div><br></div><div>What does that actually mean? Well especially if you use a "reasoning" model like Qwen, then a simple query response (like "tell me a funny dad joke") can take up to a minute to respond. This is because approximately every word of every "thought" is an output token. It "talks to itself" until if decides it has found a reasonable answer.
And it adds up quick.
Here are some basic results for this exact query on all 4 models / hardware:</div><div><ul><li>Qwen on RTX 4070: required 1200 tokens and 18 seconds to respond</li><li>Gemma on RTX 4070: required 300 tokens and 3.5 seconds to respond</li><li>Qwen on iGPU: required 630 tokens and 23 seconds to respond</li><li>Gemma on iGPU: required 450 tokens and took 18 seconds to respond</li></ul><div>Again that's moderately beefy hardware and a lot of tweaking. But I could also do far better if I accepted much dumber models. Which may be just find for basic agentic work. But much less good for coding or reasoned synthesis. And the "smartest" model took 23 seconds to reason through a dad joke. 7 seconds just to figure out how to respond to "hello". Working on truly complex reasoning can take a bathroom+coffee break to give you back results.</div></div><div><br></div><div>Note too that this doesn't take into account context size and context caching. Context is the LLM's active memory. Models have maximums (I think they are tuned for these sizes). But in most cases you'll likely have to accept something lower. However, to be even moderately useful for much of anything, you can't go terribly low. Coding tools and agentic tools just blow up if they can't remember things from one thought to the next.</div><div><br></div><div>The numbers I'm getting on the RTX are decent. Almost usable. BUT the context sizes to achieve that make it basically unusable for the kind of work I want to do. If I bump up the context window to a usable level with these models, I leak out into RAM (far slower than VRAM) and my performance tanks to unusable levels (like < 5-10 tps...at least 80-90% or more slower).</div><div><br></div><div>The iGPU with huge VRAM is slower than the dedicated GPU, but because I can crank up the context window, they actually become usable for what I want to use them for. However...speed. Claude Sonnet or Gemini Flash are easily 10x faster at everything. And more like 20x-30x faster for most reasoning work. So at some point it's a question of how much you value your time.</div><div><br></div><div>I will be using local models for some basic stuff, I think. I'm starting down a personal knowledge management path with AnythingLLM or something similar. I think it'll pair perfectly with this. And I'll likely find more ways to leverage it. But I won't be abandoning the big boys any time soon.</div><div><br></div><div>And sadly, although your Ultra 7 w/ 32GB RAM is an awesome PC, I fear your experience with local llms for anything other than experimentation and learning will prove frustratingly slow. And with prices the way they are right now, getting your PC spec'd to perform moderately well will certainly cost around the same as a full year of one of the "ultimate" plans.</div><div><br></div><div>Hope this was helpful. And that I didn't show my ignorance too badly.</div><div><br></div><div>Best,</div><div>Paul York</div></div></div><br><div class="gmail_quote"><div dir="ltr" class="gmail_attr">On Wed, May 20, 2026 at 11:35\u202fPM Lewis Wood via NFBCS <<a href="mailto:nfbcs@nfbnet.org" target="_blank">nfbcs@nfbnet.org</a>> wrote:<br></div><blockquote class="gmail_quote" style="margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"><div><div lang="EN-US"><div><p class="MsoNormal"><span style="font-size:11pt">I am currently learning as well.<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">I am now doing Ollama playlist lessons #2 currently.<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><a href="https://www.youtube.com/playlist?list=PLvsHpqLkpw0fIT-WbjY-xBRxTftjwiTLB" target="_blank">https://www.youtube.com/playlist?list=PLvsHpqLkpw0fIT-WbjY-xBRxTftjwiTLB</a><u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">I did my initial research on Lm Studio before I learned about Ollama CLI<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">This was my first Lm Studio and it was an excellent one regarding resources, models, agents, etc. Even discussed how to load partial in differing areas gpu and ddr.<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><a href="https://www.youtube.com/watch?v=UngVdAsQEiU" target="_blank">https://www.youtube.com/watch?v=UngVdAsQEiU</a><u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">You can search youtube \u201clm studio\u201d<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">Lewis Wood<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><div><div style="border-width:1pt medium medium;border-style:solid none none;border-color:rgb(225,225,225) currentcolor currentcolor;padding:3pt 0in 0in"><p class="MsoNormal"><b><span style="font-size:11pt;font-family:Calibri,sans-serif">From:</span></b><span style="font-size:11pt;font-family:Calibri,sans-serif"> NFBCS <<a href="mailto:nfbcs-bounces@nfbnet.org" target="_blank">nfbcs-bounces@nfbnet.org</a>> <b>On Behalf Of </b>Joe Orozco via NFBCS<br><b>Sent:</b> Wednesday, May 20, 2026 10:08 PM<br><b>To:</b> 'NFB in Computer Science Mailing List' <<a href="mailto:nfbcs@nfbnet.org" target="_blank">nfbcs@nfbnet.org</a>><br><b>Cc:</b> Joe Orozco <<a href="mailto:jsorozco@gmail.com" target="_blank">jsorozco@gmail.com</a>><br><b>Subject:</b> [NFBCS] Local AI<u></u><u></u></span></p></div></div><p class="MsoNormal"><u></u> <u></u></p><p class="MsoNormal"><span style="font-size:11pt">Hello,<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">With Google following in Claude\u2019s footsteps in terms of usage restrictions, can anyone speak to their experience using local LLM options? I\u2019m looking at Jemma 4 and trying to understand how accessible this route might be with JAWS on Windows.<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">I\u2019m on a fairly decent machine: 32 GB RAM, Ultra 7 processor, 4 TB SSD. I see they\u2019re recommending GPU for some of the more robust models, but I want to think most of what I\u2019m doing shouldn\u2019t require gaming machine specs. If you beg to differ though, let me know.<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">If anyone can speak to Jemma alternatives, I\u2019d also be interested. I don\u2019t think I\u2019ll suspend my subscriptions, but with these usage limitations feeling like the new standard, I want to spread my usage a little so that I don\u2019t feel like I need to be hitting the top subscriptions just to get more mileage out of the five-hour increments.<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">Thanks in advance for any tips,<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11pt">Joe<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11pt"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-family:"Times New Roman",serif">--<u></u><u></u></span></p><p class="MsoNormal"><span style="font-family:"Times New Roman",serif">Joe Orozco: Your Message, My Mission<u></u><u></u></span></p><p class="MsoNormal"><span style="font-family:"Times New Roman",serif"><a href="https://joeorozco.com/services/" target="_blank">https://joeorozco.com/services/</a><u></u><u></u></span></p><p class="MsoNormal"><u></u> <u></u></p></div></div>_______________________________________________<br>
NFBCS mailing list<br>
<a href="mailto:NFBCS@nfbnet.org" target="_blank">NFBCS@nfbnet.org</a><br>
<a href="http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org" rel="noreferrer" target="_blank">http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org</a><br>
To unsubscribe, change your list options or get your account info for NFBCS:<br>
<a href="http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/paul%40yorkfamily.com" rel="noreferrer" target="_blank">http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/paul%40yorkfamily.com</a><br>
</div></blockquote></div>
_______________________________________________<br>
NFBCS mailing list<br>
<a href="mailto:NFBCS@nfbnet.org" target="_blank">NFBCS@nfbnet.org</a><br>
<a href="http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org" rel="noreferrer" target="_blank">http://nfbnet.org/mailman/listinfo/nfbcs_nfbnet.org</a><br>
To unsubscribe, change your list options or get your account info for NFBCS:<br>
<a href="http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/jsorozco%40gmail.com" rel="noreferrer" target="_blank">http://nfbnet.org/mailman/options/nfbcs_nfbnet.org/jsorozco%40gmail.com</a><br>
</blockquote></div><div><br clear="all"></div><div><br></div><span class="gmail_signature_prefix">-- </span><br><div dir="ltr" class="gmail_signature"><div dir="ltr">--<div><br><div>Joe Orozco: Your Message, My Mission<br><a href="https://joeorozco.com/services/" target="_blank">https://joeorozco.com/services/</a> <br></div></div></div></div>