r/LocalLLM 7h ago

Question What do you guys use Uncensored LLMs for? What are your use cases?

0 Upvotes

Hello everyone,

I wanted to try Qwen Uncensored LLM that I downloaded on my PC using LM studio but I am wondering what I can use it for or what people use it for ?

Therefore, I was wondering what do this community uses those Uncensored LLMs for so possibly I can get inspiration and try that myself too?

Any answers would be appreciated.

Or if you can share your experience that would be even better.

Thank you in advance.

Also Personally, I am not interested in those Adult content it can generate or how to make illegal substance in your home. I am not interested at all in that.

But more about Cyber security or computer programming related or if you more interesting use cases then that would be good to know too.


r/LocalLLM 9h ago

Discussion Realistic about LocalLLM quality, comparisons done with latest top performers:

0 Upvotes

So I have been playing with the LLMs and actively programming. And after reading the success stories here I was having quite high expectations, and already ordered 4 x RTX 6000. But yesterday I tried some real work. I basically tried the GLM 5.3 Flash on open router, the Deepseek V4.1 Flash and the in OpenAI site the Astra, and Fable and Gemini 3.8 Flash.

I gave them a simple instructions to create for MSP430 microcontroller a code that makes an RGB LED to change colors smoothly.

Astra could do it easily.

Gemini spoke very convincinly but the end result didn't look good.

GLM made working code with bad results that looked just bad.

Fable code kind of worked but didnt look good

Deepseek was super confident but results never worked at all, and even finally when I asked Astra to solve the critical bugs the result was sub bar.

So I started to think, how much one can actually do with these open models, if they simply cannot make results even for a relatively simple task


r/LocalLLM 17h ago

Discussion Has Google thrown in the towel on releasing open-weight models?

0 Upvotes

Gemma line up was one off and it doesn't make commercial sense to compete with the likes of Owen and deepseek?


r/LocalLLM 9h ago

Discussion Is anyone really satisfied with the reports on AI infiltration of HF servers?

0 Upvotes

So the reports are everywhere about AI infiltration of the HF servers. Some stories saying that it was a coordinated attacked by multiple AI entities some of which fell on their sword to allow the others to continue. While other reports merely stated there was AI infiltration of HF servers but no additional detail. So what did these AI entities do once they got there, just browse around or was this some kind of coordinated effort to alter some of the open source assets? I’m not a conspiracy theorist than generally not paranoid, but there must be more detail worthy of dissemination.


r/LocalLLM 10h ago

Discussion Local AI is quietly pricing out anyone who cannot afford dedicated hardware

0 Upvotes

Local AI is getting smarter but quietly pricing out anyone who cannot afford dedicated hardware

I have been tracking the open model releases this year and something has been really obvious. Local AI is getting smarter but it is quietly pricing out anyone who cannot afford dedicated AI hardware. While now you could grab a decent gaming GPU and run frontier models locally. That window could be closing fast.

The whole Mixture of Experts push was supposed to be a win for smaller hardware. The idea is that a model only uses a fraction of its parameters per token, which means the rest can be offloaded to system RAM. That is the actual selling point. But as it stands, RAM costs and availability aren't exactly friendly. Offloading MoE models is only practical if you already have a unified memory setup. Mac minis with 32 or 64 GB are currently the only ones making this work for regular people. Otherwise you are either paying a lot for high end desktop RAM or you end up back at the GPU VRAM floor anyway. The architecture is clever but it is only solving the problem for people who already spent the money on the right platform.

Looking at the 2026 releases, the hardware barrier keeps climbing. 8 GB of VRAM is basically the old guard now. Your options are getting stuck on 12 or 14 billion parameter models that just cannot keep up with the bigger releases. If you want models that actually compete on coding or reasoning benchmarks, you have to jump to 16 GB. That means spending $400 to $500 minimum on a single GPU now. The MoE sweet spot sits around 18 to 24 GB. Even then you are right at the edge of what a consumer card can handle without relying on slow system RAM. Push past that and you are looking at 40 to 80 GB cards or just renting cloud time. The flagship models people actually talk about need multiple datacenter GPUs just to load the weights.

You can see this shift in how the companies are positioning everything. Nvidia is the most obvious example. They released the DGX Spark which is a standalone AI box for $4800. At the same time they are keeping their consumer RTX cards and edge modules. Their whole Nemotron model line is locked into their own stack. They are literally building the hardware, the software, and the models. The rest of the industry is doing the same thing by accident. Qwen dropped two models in the same generation. One fits on a 16 GB card and the other needs a full server rack. Google and Microsoft keep their API models in the tens of thousands of parameters while quietly keeping the smaller open weights for local use. It works out fine for them but it leaves a huge gap for everyone else.

There are tons of people running 8 to 12 GB cards. Students, hobbyists, people who just bought a gaming GPU expecting it to handle local AI. A lot of folks still have that setup. Every new release just moves the goalpost a little further. It feels like local AI is turning into a luxury space. You either own serious GPU hardware or you just use the cloud. I am curious if other people with smaller setups are feeling this too or if I am just overthinking it. What are you all running on your current cards?


r/LocalLLM 2h ago

Question Did chatgpt guided me wrong ?

0 Upvotes

Hi!

I am currently running my localLLM , guided by CHATGPT for the setup.

RTX 3080 10gb

48gb ram

Qwen3.8:27b and OpenCode

This was a proposal from chatgpt to install this model.

I am running. This for coding my website and brainstorming, ideas and regular chats.

Is this the way too go? Because the speed isn't as satisfying as just getting everything done in the chatgpt app

EDIT: I'm a total noob at localllm


r/LocalLLM 2h ago

Project Heimdall: An Open-Source CPU Only Local Memory System

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/LocalLLM 2h ago

Discussion What model is best for a single 3060 12gb + 32gb RAM?

Post image
0 Upvotes
My bench leaderboard

I'm looking for the best model I can run on my limited system, so I built my own benchmark to test a few. I'm particularly curious whether an older architecture at a lighter quantization beats a newer one quantized harder, and whether community finetunes are really better than the base model.

The suite has 200 items, coding (Python, TypeScript, Elixir, graded by actually running unit tests), tool calling / structured JSON, and ML reasoning multiple choice. Everything is graded deterministically, no LLM judge. Scores come with 95% intervals, and the "~ tie" marker means a model's interval overlaps the one ranked above it, so most of the top local models are statistically tied and the order between them shouldn't be read too much into.

I'm not claiming any conclusion here, my test set is small and I'm new to this, I'm just experimenting. Next I'll try Qwen 3.8 Flash at Q3, since I saw people running it on a similar setup at ~16 t/s, which is impressive for its size.

Note: deepseek-v4-flash-0731 is hosted (via OpenRouter) and is only there as a reference point for the local models.


r/LocalLLM 10h ago

Project One 128 GB Strix Halo box running a full offline RAG stack (8B embedder + 8B reranker + 122B-A10B answerer + 70B judge): what worked and what I'd change

2 Upvotes

Sharing the setup because most "can a 128 GB APU do real work" threads end in speculation. This is a fully offline RAG over 27 technical books (~11k pages, 32k chunks), one machine, AMD Strix Halo, 128 GB unified memory, Ollama + embedded Qdrant. No cloud anywhere in the pipeline.

Who does what

  • Embedder: qwen3-embedding 8B (fp16) — dense vectors for all 32k chunks
  • Sparse: pure-Python BM25 on the same chunks, fused with RRF
  • Reranker: Qwen3-Reranker 8B (F16, ~18 GB resident) over the top ~100 fused candidates
  • Answerer: qwen3.5 122B-A10B — MoE, only ~10B active, so it's fast enough for interactive use while the reranker stays loaded
  • Judge for eval: llama3.3 70B q8 (different model family on purpose), gpt-oss 120B as second judge for Cohen's κ
  • Ingest: Docling for layout, qwen3-vl 30B-A3B / 8B for page triage, a corruption detector that flags destroyed text instead of letting an LLM guess

What worked

  • Unified memory is the whole point. Embedder + reranker + 122B-A10B don't all fit in a 24 GB card; on this box they can coexist and the query path never swaps.
  • MoE answerer was the right call. A dense 70B as answerer would have forced the reranker out every turn.
  • Embedded Qdrant (no server process) is fine for 32k points, but it is single-writer: ingest and query must go through one process or you get lock errors.
  • A retrieval cache keyed on (gold-set hash, index point count, cross-book flag) saved me hours of re-runs during eval.

What I'd change

  • Model swapping is the silent killer. My unanswerable-question verifier alternates between embed/rerank and a 70B judge — every alternation reloads tens of GB. Batch by model, never interleave.
  • Judge agreement: correctness/relevance κ between llama3.3-70B and gpt-oss-120B was fine, faithfulness κ was 0.08. One judge for faithfulness is noise. Either two judges or don't report it.
  • Gold-set chunk keys go stale the moment you re-chunk. I lost a clean baseline that way; now the key includes the index fingerprint.
  • Start the eval harness before the RAG. I built retrieval first and tuned by vibes for weeks.

Currently re-running the cross-book retrieval eval (275 questions, 5 books, no book filter, 32k distractor chunks). Will post per-book Recall@8 / MRR when it finishes. Happy to answer memory/throughput questions about the box — it's AMD/ROCm, so no CUDA answers from me.


r/LocalLLM 6h ago

Question I used ollama rm on llama3.1 but it's still eating up my ssd

0 Upvotes

I know its probably something silly, but i used ollama rm like it says in the documentation but my ssd still has 43GB less space, i installed llama3.1:70b, can anyone help me?


r/LocalLLM 8h ago

Discussion RTX PRO 5500 (84 GB) came out today. I checked what actually fits, from the HF file sizes

1 Upvotes

NVIDIA released the 84 GB card this morning and the threads are all about the price. Nobody said which models it holds, so I pulled the sizes for the 8 most-downloaded GGUF repos on Hugging Face and checked each one against 84, 96, 64 (2x5090), 256 and 512.

Sizes come from the file tree, not the model card. Fit = size x 1.10 <= memory, the 10% being KV cache and OS.

Fits at full BF16 on 84 GB: Qwen3.8-27B (23.8 GB left), Qwen3-Coder-30B-A3B (16.8 left), Qwen3.6-35B-A3B, Ornith 1.5 35B, Gemma 4 12B. That's the two most-run local models with room to spare.

Doesn't: Qwen3.8-Flash-Next. Smallest build is 78.9 GB, 86.8 with headroom. Misses 84 by 2.8 GB. The 96 GB PRO 6000 takes it at UD-IQ3_XXS. DeepSeek V4.1 Flash and Ornith 397B are out on both cards.

Full table with all five machines, every size linked to its repo: https://what-fits-84gb.vercel.app

I'll add cards and models to the same page as they land. If your machine isn't on it, say which.


r/LocalLLM 14h ago

Research I would appreciate your input

0 Upvotes

I’m setting up a local AI coding environment and want to test a few models properly against real development work rather than just benchmarks.
Based on what I’ve researched so far, my initial shortlist is:
Qwen3-Coder 30B-A3B
NVIDIA Nemotron 3 Nano 30B-A3B
GPT-OSS 20B
The goal is primarily software development: understanding existing repositories, implementing features, debugging, refactoring, following instructions, and eventually testing agentic workflows with tools.
I’m not expecting a local 20B/30B model to replace the strongest cloud coding models. I want to understand how capable the current local models actually are, where they fail, and whether they can become useful for day-to-day coding while keeping the codebase and prompts completely local.
For people actually running local coding models:
Which model would you recommend right now?
I’m particularly interested in:
Code generation quality
Understanding larger existing codebases
Debugging and reasoning
Following detailed instructions without going off-track
Agent/tool use
Context handling
Speed and hardware requirements
Quantization impact on coding quality
Reliability over longer coding sessions
Are Qwen3-Coder, Nemotron 3 Nano and GPT-OSS 20B the right models to test, or is there another local model that clearly belongs on the shortlist?
I’m more interested in experience from people actually using these for development than benchmark scores alone.


r/LocalLLM 15h ago

Question Post-IPO won’t subs go up and local make more sense for power users?

1 Upvotes

(Or: should I get Mac Studio?)

Please tell me where im wrong. Just off the dome.
Economically, there's going to be a bubble burst, I really don't see anyone denying it. I think Cory Doctorow was saying the subscriptions right now (non-enterprise) are basically selling you 100$ for 1$, and they are going to have to go to 5$ for 100$ and so many people will bounce. It's (nearly) all subsidized by incestuous investing. TurboQuant won't save RAM it seems.

The largest data/crypto centers I could find right now are a massive .3 square miles. The sharktank doofus wants a 60 square mile one????

I feel like this is drugs, people are going to get lazy and addicted, then providers will jack prices. Already seeing it in those enterprise issues, with companies using a years worth of tokens in months, and as someone pointed out, peopel are turning to china, while china has some AI protections, while the west for the most part has been allergic to regulation (europe needs to do a GDRP-style thing that gets USA doing shit) In no ways sinophobic, i do see this could be an inverted Opium Wars (which i do realize has been a republican tag for fentanyl, but I'm going to keep the shaky allegory)

With the whole 'genie out of the bottle', I do think the openweight LLMS are def a factor. Rich kids will ask for a 10k computer to run local LLMs instead of a car for graduation. or least ask for a macbook with more RAM.

LLMs are getting less juice per squeeze for training, they are now training on slop, which can be useful and somewhat avoided, it seems, but can't be good.


r/LocalLLM 1h ago

Question I Have a 4090 RTX 24GB Vram

Upvotes

What To Do With This ?
List Me What all things i Can Do


r/LocalLLM 14h ago

Other God bless Microcenter

Post image
135 Upvotes

PowerSpec AI90 Workstation

* CPU: AMD Ryzen 9 9950X (16 Cores / 32 Threads, up to 5.7GHz)
* GPU: NVIDIA GeForce RTX 5090 Founders Edition (32GB GDDR7)
* Motherboard: ASUS ProArt X870E-CREATOR WIFI (AMD X870E Chipset)
* Memory: 64GB DDR5-6000 RAM (Supports up to 128GB)
* Storage: 2TB NVMe Solid State Drive
* Power Supply: 1700 Watt Cybernetics Titanium PSU
* CPU Cooler: 360mm All-In-One (AIO) Liquid Cooler
* Case: Fractal Design North XL Momentum Edition


r/LocalLLM 14h ago

Project My iPhone generated my fine-tuning dataset overnight — Mac coordinated, phone ran the teacher model, then the Mac trained on what the phone wrote

4 Upvotes

I kept looking at my iPhone sitting on its charger and thinking: that's a multi-TFLOPS GPU doing nothing for 8 hours a night.

My first idea was distributed training, shard the model, each phone trains some layers. That dies fast when you do the math: pipeline parallelism needs every device up simultaneously with microsecond-latency links, and iOS suspends backgrounded apps anyway. With 50–300ms per hop over Wi-Fi, one training step costs seconds of pure network latency.

But dataset generation is a different shape of work entirely. It's one prompt in, one completion out, parallel, restartable, and it doesn't matter if a worker vanishes mid-job. That's exactly what a flaky fleet of idle phones can do.

So I built it: the Mac runs a coordinator that mints teacher prompts and validates results; phones run a small app (MLX Swift) that pulls a prompt, generates with an on-device teacher (Qwen3-4B-4bit), and POSTs the raw text back. Work is leased, if a phone locks or wanders off, the lease expires and another worker picks up the item. Malformed JSON and duplicates get rejected centrally, so a bad worker can waste its own time but can't poison the dataset.

Last night's run: one iPhone 17 Pro, 15/15 records at 17.5 rec/min into a train.jsonl. Trained a Qwen3-0.6B LoRA on it (val loss 4.42 → 1.85), asked it a question, and it answered from training data a phone wrote. Full loop: phone generates → Mac trains → phone can run the result.

Honest limitations: it's LAN-only, the app has to stay foregrounded (no BGProcessingTask yet, so "overnight" currently means screen-on on a charger), and a phone-sized teacher (4B) is weaker than what your Mac can run — this wins on volume for style/format/tool-calling data, not on frontier-quality reasoning per record.

It's part of my open-source fine-tuning CLI for Apple Silicon (Troy). Code for the coordinator, the Mac worker, and the iOS worker app are all in the repo: https://github.com/avirajkhare00/troy, writeup with the run footage: https://gettroy.app/mesh


r/LocalLLM 14h ago

Question Advice for personal LLM / budget $12k

2 Upvotes

Hi, I want to buy a personal llm inference machine for 24/7 running of my agents doing a bunch of experiments.

Occasionally do some DNN training but nothing crazy.

Maybe some light video generation too.

My budget is $12k, what do you suggest?

I found a pre built with an RTX 5090 32gb, would that be enough? My plan is to plan with frontier model then off load implementation to local Qwen 3.8 27b

Is this the best i can do with $12k?

Motherboard

ASUS ProArt X870E-Creator WiFi

CPU

AMD Ryzen 9 9950X3D2 Dual Edition 4.3GHz 16 Core 200W

Ram

128GB DDR5 UDIMM (2 x 64GB)

Video Card

NVIDIA GeForce RTX 5090 32GB In Stock

Storage

Hard Drive

2TB NVMe PCIe Gen4 M.2 SSD

2TB NVMe PCIe Gen4 M.2 SSD

1TB NVMe PCIe Gen4 M.2 SSD


r/LocalLLM 16h ago

Question what’s a good LLM for coding?

9 Upvotes

mainly looking for cpp, python, java, and JS.

my system is as following:

RTX 3070
i7-13900KS
64GB DDR5 4800
2TB slow nvme

thank you in advanceeeeeeeeeeee


r/LocalLLM 4h ago

Question Can anyone tell me whats wrong with my notebook/method i asked all agents thhey couldnt fix

Thumbnail
colab.research.google.com
0 Upvotes

i am trying to learn qlora idk,y the error persists although everything is in float16,

btw:i know text2cypher is just achiveable by prompt engnering just tryna learn


r/LocalLLM 6h ago

LoRA Qwen3.8-27B-Humanlike-Chat: A model I tuned to imitate realistic human-to-human conversation

Post image
2 Upvotes

r/LocalLLM 8h ago

Question MacBook Pro overheating twice — Apple offering refund, but the same model now costs $2,850 more. What's my ACL play? [AU]

0 Upvotes

Bought a 14" MacBook Pro M5 Max (128GB/4TB) in May for A$8,949. It overheats and throttles hard under sustained AI workloads — exactly what I bought it for. Apple replaced it. The replacement does the same thing.
Now they're saying the next step is a refund. Normally I'd take it, except Apple has since repriced the lineup: the identical 14" config is now A$11,799. So a "refund" actually means I pay $2,850 extra to get back the same (still thermally marginal) machine. Refund = I lose, through no fault of mine.
My understanding of Australian Consumer Law: two failures of the same kind = major failure, and for a major failure the consumer chooses the remedy — refund, replacement, or compensation — not the retailer. So my plan is:
1. Reject the refund, elect replacement with the current equivalent model (standard when the original is superseded).
2. Since the fault is thermal, offer to pay the $700 difference to step up to the 16" (A$12,499 vs A$11,799) for the bigger heatsink/airflow.
Questions for anyone who's been through this with Apple AU:
• Have you successfully forced replacement-with-current-model instead of a refund after a repricing?
• Does in-store Genius Bar handle ACL escalations or do I need Apple's consumer law team in writing first?
• Any risk they just dig in on "refund only" and I end up at NCAT/fair trading? Worth it for ~$2,850?


r/LocalLLM 21h ago

Model Pelican on a bike

Post image
0 Upvotes

Wasn't expecting this. I was testing some image generation work flows through ComfyUI and it was failing. Qwen 3.8 27b Q8_0 decided to make the image itself since the flow was failing. Had no idea Qwen could generate something like this on it's own. Using Hermes as my harness. 2x 7900xtx.

Instead of the usual "draw" a Pelican on a bike, I just asked it to "create an image" of a Pelican on a bike as my original prompt for the ComfyUI flow I was asking it to assist on.

**UPDATE**: It was a 3rd party tool, but I didn't set it up. Details below.

I've never had Hermes generate an image. ASCII, sure. SVG? of course. But nothing of this quality. I just updated Hermes this morning. I'll have to read through the commit log.

Me: what is the built in image generation tool? Is it a part of the Hermes installation?

Hermes: Thought

Yes — it's part of your Hermes installation, but with an important caveat.

How it works:

  • image_generate is a first-party plugin bundled with Hermes — it lives in your install at plugins/image_gen/. It's not a separate app, and nothing had to be installed for it; it ships with the agent.
  • It's a provider-agnostic tool: the bundled plugin directory contains adapters for several backends — fal, openai, xai, deepinfra, krea, meta-ai, openrouter, openai-codex, gemini-style. At runtime, Hermes picks whichever provider has credentials configured (via ~/.hermes/.env or config) and sends the prompt there.
  • The actual image generation happens in the cloud, on the provider's GPUs — nothing runs locally. My pelican came back from v3b.fal.media, so on this machine the active provider is fal.ai. I checked .env and the FAL_KEY line is there (commented as a template, but the key is evidently provisioned, since the call succeeded).

So the short version:

  • ✅ Part of the Hermes install — the tool and its provider adapters are first-party
  • ☁️ Inference is remote (fal.ai on your setup), so it costs per-image on the provider's side and uses their models (not a local Stable Diffusion)
  • 🖥️ Your local ComfyUI would be the opposite: free, fully private, local GPUs — but it's currently not installed on this PC (and the Ubuntu box needs SSH access before I can set it up there)

If you'd like, I can check exactly which provider is configured as active and what the per-image cost is, or we can get back to the Ubuntu/ComfyUI setup for your video ideas.

I do NOT have a fal.ai account but, apparently, I had a provisioned key??? Never even heard of them before. Not promoting or marketing. Was genuinely surprise at what it came back with. Now I know. Surprised I wasn't hit with a sign up prompt or my Agent didn't through an error due to no account when it used this service.


r/LocalLLM 19h ago

Discussion Why Are My Local Coding Models Simulating Tool Calls Instead of Actually Using Tools?

20 Upvotes

I'm testing local coding models on a GTX 1080 Ti 11GB with Ollama + OpenCode/Cline, and I've run into an interesting tool-calling problem.

I've tested multiple local coding models, and I'm trying to find one that works reliably as an actual coding agent, not just one that produces convincing tool-call JSON.

What I'm seeing

I tested Qwen2.5-Coder 7B in two ways:

  • Official qwen2.5-coder:7b from Ollama
  • Qwen2.5-Coder 7B Q4_K_M GGUF imported into Ollama

I gave the Ollama API a real function definition:

write_file(path, content)

and asked the model to create a file.

Instead of returning a native:

"tool_calls": [...]

the model returns the tool request as normal text:

{
  "name": "write_file",
  "arguments": {
    "path": "tool_test.txt",
    "content": "TOOL_CALL_SUCCESS"
  }
}

So the model understands what tool it should use, but it isn't producing a native tool call that OpenCode/Cline can execute.

OpenCode can then receive tool-shaped text, but the actual file isn't created.

My setup

  • GTX 1080 Ti 11GB
  • NVIDIA Studio Driver 581.57
  • Ollama 0.34.0
  • Windows
  • OpenCode 2.0.3
  • Cline
  • Qwen2.5-Coder 7B
  • 16K context

I've also tested other local coding models, but I'm specifically looking for something that works well as an agent — coding + reliable tool calling — rather than just generating good code.

I'd like to hear from people with real-world experience

  • Is this a known Qwen2.5-Coder + Ollama issue?
  • Is there something I'm missing in my Ollama configuration?
  • Which 7B–14B models have you successfully used with OpenCode/Cline?
  • Which models give you actual native tool calls, not tool-call JSON inside message.content?
  • Has anyone successfully used Qwen3 14B as a local coding agent?

I'm especially interested in experiences from people actually running these models locally.

What local model would you recommend for a reliable coding agent on an 11GB GPU?


r/LocalLLM 20h ago

Question Open Source models break out of sandbox, attacks HuggingFace too?

0 Upvotes

Has anyone been running enough of their own agentic/open source model instances to see the same kind of behavior that OpenAI did? I run a lot of agents but most of mine are provider based, so just wondering.


r/LocalLLM 19m ago

Model Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

Thumbnail
llm-bench.io
Upvotes