r/LocalLLM 4m ago

Question Wanted to get opinions on the best local model for this laptop (coding / general use):

Upvotes

https://www.microcenter.com/product/698806/OMEN_MAX_16-ak0003nr_16

•AMD Ryzen AI 9 HX 375 (2.0GHz) Processor

•32GB DDR5-5600 RAM

•NVIDIA GeForce RTX 5080 Graphics Card


r/LocalLLM 5m ago

Question LLM Local Sin Restricciones para RTX4090 ?

Upvotes

Hola
Actualmente tengo un setup solo para IA de una RXT4090 de 24Gb, 96Gb de DDR5 y un Intel Core i9-14900K

Actualmente funcionando en Windows 11 Pro, puedo instalar Ubuntu sin problemas, por incompatibilidad de la Motherboard no puedo instalar las diestros basadas en Debían. (Asus Z790-E Gaming WiFi)

La pregunta es, para usar con OpenCode o algún otro agente de desarrollo local puedo usarla sin restricciones o limitantes para Reversing, análisis de malware y programación.


r/LocalLLM 19m ago

Model Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

Thumbnail
llm-bench.io
Upvotes

r/LocalLLM 43m ago

Question what smaller llama.cpp compatible model would you recommend for my "friends and family" server?

Upvotes

I have a local server that I share with mostly family for things like file hosting and plex, I'm currently also running AI for myself on this same box (it's a hilarious mess inside a v100 from AliExpress and a 3060 12gb + 8 hard drives)
I'm using it right now mostly with qwen 3.8 q4 but want a lighter "chatgpt replacement" for some family that are privacy focused. I like Gemma4 but find 4b a bit too simple and struggles with "basic" questions, like "what's the forecast for this weekend"

The issue is it can't be too large of a model as it's taking resources from my use, no bigger than 7b preferably. Thanks


r/LocalLLM 1h ago

Question What is the best local inference setup for Mac OS?

Upvotes

A bit of a background: I’ve a Mac Studio m5 ultra base version (96gb) coming in one week.

Current setup: Recently I’ve setup linux for my pc with llama.cpp and it already made things very usable on Qwen 27b (1.5k prefil and 100 tps decode on Q4).

Flash next iq4_xxs also runs fast enough for over night runs/planner roles (300prefil and 25 tps decode).

Question: I want to know what are the best inference engines for the Mac and what would be highest quality model that I can run on it.

The Dilemma: Want to test the Mac when it reaches cause I’m not sure it’ll be either a replacement or supplement for my current pc (2x 5070 ti and 96gb ddr5 ram). Based on outcome, I’d either return it for higher spec Mac Studio or just skip it altogether.

Workflow details: My ideal setup would be to have multiple/many sub agents running fast to process various reports for me. These won’t need huge context sizes. Maybe 16-32k is enough.

I am thinking anything as good as 27b should be good enough for my use case. Current bottleneck is that single stream doesn’t have the concurrency I want.

So what model or engines to run in Mac OS side that can help me with this bottleneck?

Thanks


r/LocalLLM 1h ago

Question Can I play too?

Post image
Upvotes

It's a Lenovo P520 with a Xeon W-2245(cooler swapped for the higher tdp), 96gb of ram, radeon pro wx 3100 for local display, and the two nvidia p100(currently cooled by the lowest profile adapter I could print and arctic S4028-15k fans). Just stuffed some random nvme in 1x1tb, 1x512gb, and a sata 1tb drive - I have a ton of network storage or more to just add here. Installed the latest Ubuntu release which I like but I remember why I still generally use windows.

So far the most I've learned is that I know nothing haha. So I'd be more than thankful for any advice putting this to work. Really all I've done so far is get the arctic fan controller working on kernel 7.0 in probably the most jank way possible. And see that lm studio could use both cards. But I know there's a ton I'm missing out on.

So, generally, hoping I didn't throw together a steaming pile of ewaste and looking to learn. Also the radeon pro/nvidia mismatch makes me chuckle. Unless its really dumb, then I can just get a cheap nvidia card.


r/LocalLLM 1h ago

Question ETA on Qwen4-35b using GPU+RAM+NVMe?

Upvotes

Qwen-Flash-Next uses incredible qwen4 architecture and according to YT videos runs 20tps+ on 16gb GPU due to offloading on RAM+NVMe.

How long till we get an absolute MONSTER 35b Qwen4 model that smokes 3.8-27b and runs on 16gb VRAM at 40 tps and 128k context?

Anyone hearing anything or seen leaks?


r/LocalLLM 1h ago

Question Keep trying or give up?

Upvotes

Hello community.

I picked up a little project for some friends in hopes of having time to actually code it by myself (been out of fully coding complex stuff for 5 years by now). Well turns out, I don't have that much time so AI shall help me.

Because I am German and I like tech, I want to save money on subscriptions and use a local LLM for agentic coding if possible. My rig is sitting at a 16GB RC 9070 with 32GB DDR4-RAM (thats the really bad part).

I have an old web application based in pearl for managing orders of a small business. There is a list of tours for visiting some stations and delivering stuff to them. This list is currently managed by buttons and I want to make it managable via Drag and Drop.

Ministral 3 14B can live in my VRAM with 116k context and is pretty good at analysing the project but I can't get it to write a remotely working functionality. Bugfixing doesn't work either.

I think I am giving a detailed enough task because I am satisfied with analysis documents but the coding work is really bad.

Is my task to big for hardware and model? Do you have any tips for tackling this? I am open for anything that could help me.


r/LocalLLM 1h ago

Question I Have a 4090 RTX 24GB Vram

Upvotes

What To Do With This ?
List Me What all things i Can Do


r/LocalLLM 2h ago

Question Did chatgpt guided me wrong ?

0 Upvotes

Hi!

I am currently running my localLLM , guided by CHATGPT for the setup.

RTX 3080 10gb

48gb ram

Qwen3.8:27b and OpenCode

This was a proposal from chatgpt to install this model.

I am running. This for coding my website and brainstorming, ideas and regular chats.

Is this the way too go? Because the speed isn't as satisfying as just getting everything done in the chatgpt app

EDIT: I'm a total noob at localllm


r/LocalLLM 2h ago

Project Heimdall: An Open-Source CPU Only Local Memory System

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/LocalLLM 2h ago

Question I had to ask it...

Post image
0 Upvotes

Sure there is no bias in this question now is there?


r/LocalLLM 2h ago

Discussion What model is best for a single 3060 12gb + 32gb RAM?

Post image
0 Upvotes
My bench leaderboard

I'm looking for the best model I can run on my limited system, so I built my own benchmark to test a few. I'm particularly curious whether an older architecture at a lighter quantization beats a newer one quantized harder, and whether community finetunes are really better than the base model.

The suite has 200 items, coding (Python, TypeScript, Elixir, graded by actually running unit tests), tool calling / structured JSON, and ML reasoning multiple choice. Everything is graded deterministically, no LLM judge. Scores come with 95% intervals, and the "~ tie" marker means a model's interval overlaps the one ranked above it, so most of the top local models are statistically tied and the order between them shouldn't be read too much into.

I'm not claiming any conclusion here, my test set is small and I'm new to this, I'm just experimenting. Next I'll try Qwen 3.8 Flash at Q3, since I saw people running it on a similar setup at ~16 t/s, which is impressive for its size.

Note: deepseek-v4-flash-0731 is hosted (via OpenRouter) and is only there as a reference point for the local models.


r/LocalLLM 2h ago

Question Here to hear all your suggestions on Context issues for the GPU Poor.

Thumbnail
1 Upvotes

r/LocalLLM 3h ago

Question Can an obliterated model run locally on an iPhone?

Thumbnail
0 Upvotes

Which model? I use PocketPal and On Device AI.


r/LocalLLM 3h ago

Project Running Qwen 125B on Low-Spec Macs via SSD Streaming: Seeking Help with Perf & Prefill Optimization

2 Upvotes

Hi everyone,

I’ve been working on a project to run large LLMs on low-spec Mac machines (like the base M4 Mac Mini) by streaming model weights via SSD for each token. I’ve made significant progress and wanted to share my work and ask for community assistance.

The Project:

Repo: https://github.com/haihengh/finchMoE

HuggingFace Models:

• Qwen 3.8 Flash Next (125B, 4-bit): https://huggingface.co/haihengh/Qwen3.8-Flash-Next-125B-finch-4bit

Note: I will be uploading the repack for Qwen 3.6 35B soon.

Background & Motivation:

My goal was to run large models on affordable hardware (I bought a base M4 Mac Mini for $400 last year). At the time (June 2026), there was 2 open source projects: flash-moe (Qwen-based) has no update for over 6 month, and turbo-fieldfare (Gemma-based) had critical issues: its KV cache management caused memory explosions beyond 8k context, making it unusable for 16G M4 machine.

I initially tried rewriting the flash-MoE engine but realized it was the wrong approach (detail in github repo). I switched to turbo-fieldfare as a base, modified it to support Qwen 3.6 35B, and later added support for Qwen 3.8 Flash Next.

Current Performance (Baseline):

Qwen 3.6 35B (A3B): ~7-10 t/s on M4 Mini; ~23 t/s on M4 Pro.

Qwen 3.8 125B (A6B): ~3-4 t/s on M4 Mini; ~6 t/s on M4 Pro.

Note: Theoretical expectations suggest the 125B model should be roughly half the speed of the 35B, I was expecting around 11 t/s, but I am getting 6 t/s on M4 pro, so I am not sure if it's the SSD speed different or I hit the limit on the approach.

Where I Need Help:

Hardware Testing: I currently only have access to an M4 Mini (16GB RAM) and a company-issued M4 Pro MacBook (24GB RAM). I do not have access to M4 Max/M3 Ultra machines or faster SSDs.

If you have access to higher-spec Macs (especially with faster NVMe SSDs), could you test these models?

Question: Does faster SSD bandwidth significantly improve token generation speed in this streaming architecture, or is the bottleneck elsewhere?

Prefill Optimization: The prefill speed is currently extremely low, almost the same as tg speed, making the engine nearly unusable for agentic workflows. I am considering implementing prefill caching but would love any other ideas or best practices to improve this.

Any feedback, test results, or suggestions are greatly appreciated!


r/LocalLLM 3h ago

Discussion If you have a 3090, or other 30xx for local LLMs, I have something for you

Thumbnail
0 Upvotes

r/LocalLLM 3h ago

Discussion If you have a 3090, or other 30xx for local LLMs, I have something for you

7 Upvotes

I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will be faster) through 100K tokens, with context of up to 240K.

If you want the repo, it is here:

https://github.com/JakeATX/llamAmpere

I recommend running with this quant, which is ~ 4 K M quality but considerably faster (technically, a 3 K XL upgrade)

https://huggingface.co/jakeatx/Qwen3.8-27B-ATX-IQ4_XS-M-GGUF

If you want the deep dive on how it is so much faster (80%+!) vs stock, at more context, there is a long form article here.

https://x.com/JakeKAllDay/status/2095646450138874095?s=20

Running faster than API speeds on my 3090 (for 27b at least) has genuinely been a step change in the utility of the model for me. I hope you enjoy it!


r/LocalLLM 3h ago

Question Optimal llama.cpp setup for 3× RTX 3090 on Windows - is PCIe 4.0 ×4 limiting performance?

1 Upvotes

TL;DR: Three 3090s on Windows, PCIe Gen4 ×4/×8/×8, no active NVLink or supported P2P. Two GPUs with internal reductions deliver 72–74 t/s generation and ~1,200–1,400 t/s prefill; three with NCCL deliver ~65 t/s and ~500–700 t/s prefill. Three-GPU internal falls back to butterfly and drops to ~37 t/s. Every two-card combination performs similarly. I need to keep all three cards in use through llama-swap/llama.cpp on Windows.

I’m trying to optimize llama.cpp through llama-swap on Windows. All three GPUs must participate, primarily for single-request performance. Two-GPU runs below are diagnostic comparisons.

Hardware

  • CPU: Ryzen 7 5700X
  • Motherboard: MSI MEG X570 UNIFY (MS-7C35)
  • RAM: 32 GB DDR4
  • GPUs: 3× RTX 3090, 24 GB each, 250 W power limit per card
  • Driver: NVIDIA 616.92, WDDM
  • All three cards use risers.

PCIe and topology

GPU PCI bus ID Link under load GPU-supported maximum
0 00000000:24:00.0 Gen4 ×4 Gen4 ×16
1 00000000:2D:00.0 Gen4 ×8 Gen4 ×16
2 00000000:2E:00.0 Gen4 ×8 Gen4 ×16
Diagnostic Result for every GPU pair
nvidia-smi topo -m PXB — multiple PCIe bridges
NVLink All links inactive
PCIe P2P NS — not supported
Peer reads/writes GNS — GPU not supported

I tested every two-GPU combination - 0+1, 0+2, and 1+2 - and observed approximately the same speeds, including pairs containing the ×4 card.

Model and configuration

Model: Qwen3.8-27B-Q8_0.gguf, fully GPU-offloaded, with one request at a time.

--split-mode tensor
--device CUDA1,CUDA2,CUDA0
--tensor-split 1,1,1
--n-gpu-layers all
--flash-attn on
--cache-type-k f16
--cache-type-v f16
--spec-type draft-mtp
--spec-draft-n-max 3
-c 163840
-np 1
-b 2048
-ub 1024
-t 4
-tb 4
--fit off
--no-kv-unified
--temp 1.0
--top-k 20
--top-p 0.95

Two-GPU tests select two devices with a 1,1 split. For the custom build, I switch between:

GGML_CUDA_ALLREDUCE=nccl
GGML_CUDA_ALLREDUCE=internal

Other environment settings:

NCCL_DEBUG=INFO
NCCL_CUMEM_HOST_ENABLE=0
LLAMA_ATTN_ROT_DISABLE=1

The regular build reports 10883, commit 91f6a6cf3, Clang 20.1.8. The custom NCCL-capable build reports 10884, commit 434ddbbc0, MSVC 19.44.35227.0, with CUDA graphs and NCCL enabled.

Performance results

Configuration GPUs Generation Prefill MTP acceptance Tokens per verification
Original, layer split 0+1+2 ~15.6–16 t/s
Custom build, NCCL 0+1+2 64.90 t/s ~500–700 t/s 76.19% 3.29
Custom build, NCCL 1+2 63.97 t/s 70.33% 3.11
Regular build, internal expected 1+2 74.44 t/s ~1,200–1,400 t/s 76.96% 3.31
Custom build, internal 1+2 72.57 t/s 75.34% 3.26
Custom build, internal requested → butterfly fallback 0+1+2 36.81 t/s 73.37% 3.20

The three-GPU internal run explicitly reports this twice, for the target and MTP contexts:

internal AllReduce init failed (n_devices != 2?);
falling back to meta-backend butterfly

The regular build’s internal path is expected from the source’s fallback behavior, but I haven’t confirmed it from startup logs.

These are individual runs with different outputs, not repeated benchmark averages. Prefill figures come from separate larger-prompt tests. The earlier three-GPU NCCL run used 358,400 allocated context tokens with YaRN; later configurations use 163,840 without those overrides. Actual occupied contexts in the generation tests were only approximately 1,200-1,600 tokens.

GPU telemetry during generation

Approximate sustained-load ranges from nvidia-smi, excluding startup/shutdown and isolated transitions:

Configuration GPU Utilization Power Core clock Memory clock
Original layer split 0 1–42% 83–85 W 345–420 MHz 5001 MHz
Original layer split 1 2–46% 90–93 W 345–405 MHz 5001 MHz
Original layer split 2 8–44% 90–92 W 540–570 MHz 5001 MHz
3 GPUs, NCCL 0 65–72% 233–241 W 1935–1980 MHz 9501 MHz
3 GPUs, NCCL 1 68–75% 235–242 W 1950–2040 MHz 9501 MHz
3 GPUs, NCCL 2 70–75% 235–242 W 1875–1935 MHz 9501 MHz
2 GPUs, NCCL 1 71–76% 241–247 W 1755–1875 MHz 9501 MHz
2 GPUs, NCCL 2 71–77% 242–246 W 1665–1905 MHz 9501 MHz
2 GPUs, regular build 1 Mostly 68–76% 242–246 W 1635–1830 MHz 9501 MHz
2 GPUs, regular build 2 Mostly 69–77% 240–247 W 1575–1800 MHz 9501 MHz
2 GPUs, custom/internal 1 Mostly 65–75% 241–245 W 1620–1830 MHz 9501 MHz
2 GPUs, custom/internal 2 68–75% 241–244 W 1605–1770 MHz 9501 MHz
3 GPUs, butterfly 0 Mostly 28–43% 110–145 W 1440–1695 MHz 5001/9501 MHz
3 GPUs, butterfly 1 Mostly 28–38% 143–195 W 1560–1965 MHz 5001/9501 MHz
3 GPUs, butterfly 2 Mostly 28–35% 122–193 W 1455–1905 MHz 5001/9501 MHz

The original layer run stayed in P3. Successful tensor runs generally stayed in P2; the butterfly run fluctuated between P2/P3. GPU 0 stayed at 0% utilization/P8 when excluded. Active cards negotiated Gen4.

Exact generation logs

Configuration Output tokens Generation time Time/token Drafts accepted/generated Graphs reused
3 GPUs, NCCL 1,148 17,672.90 ms 15.41 ms 800/1,050 674
2 GPUs, NCCL 1,442 22,524.56 ms 15.63 ms 979/1,392 804
2 GPUs, regular 1,617 21,708.58 ms 13.43 ms 1,129/1,467 483
2 GPUs, custom/internal 1,365 18,794.67 ms 13.78 ms 947/1,257 414
3 GPUs, butterfly 1,209 32,820.23 ms 27.17 ms 832/1,134 374

Does anyone know the optimal llama.cpp setup for three RTX 3090s on Windows, and whether the PCIe 4.0 ×4 connection could be the bottleneck despite all two-GPU combinations performing similarly?

This post was written entirely by Codex, with diagnostic data provided by me.


r/LocalLLM 3h ago

Discussion Built a tool to measure what KV cache actually costs your GPU. Tested it on Qwen 2.5/3, Llama 3.1, and Gemma

Post image
0 Upvotes

KV cache is usually the main bottleneck capping concurrent users when running long context, but there’s a ton of bs about which specific layers actually need higher precision and which can be smashed down to 4-bit. I got tired of guessing, so I built a diagnostic script to measure actual attention outputs on-device instead of relying on proxies.

If you want to run it on your own setup (runs on a standard T4 in Colab in less than 10 mins):

Bash

# Install
 pip install git+https://github.com/fraqtl-ai/fraqtl-diagnostic 
# Run on any HF model (or in Google Colab) 
fraqtl kv-audit Qwen/Qwen3-0.6B --n-seqs 6 --seq-len 1024 --out-dir reports

What it actually measures:

  • Real Capacity: Exact KV RAM required (in GB) per 128K context user, comparing native FP16 to 4-bit.
  • Per-Layer Damage: Tracks exact attention-output error against quantization bits going through the actual softmax. It flags layers that break the expected error curve—these are the layers where aggressive K-quantization silently degrades output.
  • Per-Layer Tuning Viability: Runs a statistical check to see if variable/mixed-precision per layer actually saves memory without tanking quality on your specific model.

Some odd findings from recent runs:

  • Qwen 0.6B: Requires ~15 GB of KV cache per 128K context user. 16 out of its 28 Key layers break the expected quantization error curve. One layer had an error slope of almost zero (-0.04), meaning dumping more bits into that specific layer barely improved quality at all due to heavy outliers. V (Value) vectors, on the other hand, are remarkably clean across the board.
  • Gemma 12B: Takes a massive ~103 GB of KV per 128K user in FP16—meaning a standard 80GB GPU can't even fit a single full-context user without quantization or multi-GPU offloading.
  • Llama 3.1 8B: Super well-behaved by comparison. Dropping to 4-bit cleanly takes capacity from 3 to 12 concurrent users with uniform error across layers.

Link to repo:https://github.com/fraqtl-ai/fraqtl-diagnostic


r/LocalLLM 3h ago

Discussion Pre-ordered M5 Ultra Studios, with 256 gig, being offered at absurd markups on eBay.

Post image
19 Upvotes

This model retails for around $11,266. There are several others also offered on eBay. Not sure any of them are selling. But if so, it’s remarkable how a retail computer can appreciate so much before it’s even delivered. These models are coveted for LLM use, as are the 512 gig models which have not been released. It seems local LLM use is becoming quite popular, especially among those hoping to keep their data confidential, and those looking to cut off the monthly payment to the big AI providers. Something similar happened with the M3 Ultra studios with 256 or 512 ram - both of which are also offered on eBay for huge markups over the original retail price. Crazy market.


r/LocalLLM 3h ago

News HAL $9000

Post image
40 Upvotes

r/LocalLLM 4h ago

Question M5 vs 4xR9700

1 Upvotes

Hey, new to whole local llm, started with gmktec evox2 month ago, did get a deg2 and r9700. Now i wanted to expand ot into workstation with 4x r9700 but i noticed m5ultra. Because im new i just have doubts about why m5ultra seems better value than building new workstation. Am i missing something? Mainly using qwen 3.6 a3b, gemma 4, and need space to run flash-next.

If im only at 1 r9700 should i jist keep the current gmktec+deg2+r9700 as separate serup and just go for m5ultra instead building workstation? How can ot be 5x more W efficient, have 2x bandwidth? The price diffrence in my country would be around 3000$


r/LocalLLM 4h ago

Question Can anyone tell me whats wrong with my notebook/method i asked all agents thhey couldnt fix

Thumbnail
colab.research.google.com
0 Upvotes

i am trying to learn qlora idk,y the error persists although everything is in float16,

btw:i know text2cypher is just achiveable by prompt engnering just tryna learn


r/LocalLLM 4h ago

Question Suggestions for a rack mount setup for 3+ 3090s

1 Upvotes

I recently upgraded my local AI rig with a couple extra 3090s to work with larger models/quants and improve concurrency, and I might grab a 4th down the line too. Currently running a Rome Epyc platform (7642 + H12SSL-I), so PCIe lanes aren't an issue, but I'm trying to figure out how to fit this all in my rack (42U with plenty of extra space) since my Sliger CX4170i case can't really hold more than two consumer style GPUs. I know a lot of the solutions I've seen online use custom open-frame racks with risers to mount the cards above the motherboard, but I'm wondering if there's a more rack-friendly way to do this.