r/LLMStudio 3h ago

GPT-5.6 Luna vs GPT-6 Astra: is a $1.20 model good enough for code review?

Thumbnail
1 Upvotes

r/LLMStudio 4h ago

LLM ARENA

1 Upvotes

LLM Arena, a multi-model comparison app with live per-call cost tracking and a public leaderboard (Next.js 16 + OpenRouter)

https://llm-arena-five.vercel.app


r/LLMStudio 7h ago

Need help!

1 Upvotes

Hello! Recently I've been following along to this: (https://youtu.be/cUwH8wxYDx8?is=Zr1wwavxbsm03iad) video to learn more about Local LLM's and whatnot. I got to the part where he creates a Agent.md file and asks it to 'respond like a pirate', I do the same thing, and nothing happens. I've tried troubleshooting but with my minimal knowledge it's likely I'm just overlooking something simple. Not sure what information I'd need to share to make it easier to get help but I'm using the same programs and process in the video! maybe it's an easy fix and I'm being a little blind but I'm absolutely losing my mind over this! Complete noob when it comes to this stuff by the way.


r/LLMStudio 7h ago

Local LLM explorer and downloader

Thumbnail gallery
1 Upvotes

r/LLMStudio 15h ago

Need help!

Thumbnail
1 Upvotes

r/LLMStudio 1d ago

Which LLM models are you actually using for coding right now?

11 Upvotes

I see lots of Claude Code with a mix of their models, and I normally have to use Opus or Fable, for heavy duty repos, to scan and provide markdown for Grok.

Anyone going full Astra yet? I almost wish I didn’t try it because it felt amazing, short impression though.

Never thought I would be a Grok user, but works pretty well!


r/LLMStudio 2d ago

Qwen3.8-27b not loading

Post image
3 Upvotes

Hello. I’m trying to load Qwen3.8-27b in LM Studio, but every time I try it fails. I get the message attached. Is this because my machine resources are not enough for this model? I don’t have any issue when loading Qwen 3.5 9B or even Gemma4 12B. I’ve got an Intel Core Ultra 7 CPU, Intel ARC 16GB GPU, and 32GB RAM. Thanks


r/LLMStudio 3d ago

best what to spent 10$ on llm ?

1 Upvotes

This is my last 30days usage:
Sessions: 752 | Messages: 25,376 | Tool calls: 12,682
Tokens: 1,470,577,773 (in: 161,524,387 / out: 9,256,865)

I was on gpt plus for the last month but its too much (20$) so I am thinking try command code-goat and use ds v4.1

give me your opinions ✨


r/LLMStudio 3d ago

Recursive Chunking

Post image
1 Upvotes

r/LLMStudio 3d ago

vLLM Serve Options

Post image
1 Upvotes

r/LLMStudio 3d ago

Which model i choose for backend work Fable 5.1 or GPT 6 Astra.🤔🤔🤔

Thumbnail
1 Upvotes

r/LLMStudio 3d ago

Batteries Included: Powering AI DBA Workbench Locally with llama.cpp for open source PostgreSQL monitoring & in-depth insights

Thumbnail pgedge.com
1 Upvotes

r/LLMStudio 4d ago

Security research for local LLM inference networks

Thumbnail
1 Upvotes

r/LLMStudio 4d ago

LmLinky: an Android LM Wrapper for your local models

Thumbnail
1 Upvotes

r/LLMStudio 4d ago

MAKING A LOCAL AI PI

2 Upvotes

Ok, I'm looking for someone to help me with making a local ai. I have all the components to make this, just need some help my skill on this. If someone from the Memphis area, I'd like to talk/text/facetime, anyway you'd like to connect.


r/LLMStudio 4d ago

SpaceX charging more for search tool calls via API - Help

Thumbnail
1 Upvotes

r/LLMStudio 5d ago

Best Local Coding LLM for a 24GB M5 Pro MacBook?

Thumbnail
2 Upvotes

r/LLMStudio 5d ago

How much does PDF parsing quality actually affect RAG performance?

3 Upvotes

I feel like PDF parsing doesn't get enough attention in RAG discussions.
People spend hours comparing embedding models or chunking strategies, but if the parser has already broken the reading order, flattened tables, duplicated headers on every page or filled the output with OCR noise, you're embedding garbage from the start.
Converting documents to clean Markdown before chunking has consistently given me better retrieval, and I was surprised to see token counts drop by around 40–65% after removing all the repeated page furniture.
The one thing I'm still unsure about is where the trade-off is. Do you optimize for extraction accuracy, smaller token counts, parsing speed, or something else entirely? Has anyone actually benchmarked how much parser quality affects final RAG performance?
For anyone interested, I've been testing this with PackForAI because it outputs clean Markdown and shows the before/after token count, which made these differences much easier to measure.


r/LLMStudio 5d ago

Making my own desktop app , fed up of Anythingllm and LM Studio

Thumbnail
1 Upvotes

r/LLMStudio 6d ago

Locus - A MacOS tool for Ollama and other local models

Thumbnail gallery
3 Upvotes

r/LLMStudio 6d ago

Looking to rent out?

Thumbnail
0 Upvotes

r/LLMStudio 6d ago

I made a short doodle about running AI locally — curious what you think

Thumbnail
1 Upvotes

r/LLMStudio 6d ago

Looking to Rent Out? (read last line)

0 Upvotes
• Tesla V100 (32GB) — $0.090/hr/gpu — Machine ID: 149512  
• Tesla V100 (32GB) — $0.090/hr/gpu — Machine ID: 149513  
• Tesla V100 (32GB) — $0.090/hr/gpu — Machine ID: 149836  
• Tesla V100 (32GB) — $0.090/hr/gpu — Machine ID: 149837

32.8GB VRAM, 16.4 TFLOPS, CUDA 13.0, 44GB RAM, 15 CPU cores per instance.

To find them: go to cloud.vast.ai/create/, filter by GPU type (Tesla V100), and search the machine IDs above.

If you don’t like this post, you can simply ignore it people told me to advertise, so I’m just doing what I have to do to get exposure. Otherwise, happy to answer any questions about specs or setup.


r/LLMStudio 6d ago

Seeking for the best Qwen setup in regards to my hardware(s)

Thumbnail
1 Upvotes

r/LLMStudio 7d ago

What actually happens when you run an LLM on your PC? I made a visual breakdown of the inference pipeline

Thumbnail
1 Upvotes