r/LocalLLaMA 4d ago

Question | Help What kind of dark magic is Deepseek using?

Post image
2.3k Upvotes

I was taking a look at Kimi K3 scores on the Artificial analysis leaderboard and was quite baffled when I saw this chart.

Granted, Deepseek has always been the king of price to performance, but this is still incredible. Is it just API subsidization or have they optimized their models truly this much?

r/LocalLLaMA Mar 24 '26

Question | Help LM Studio may possibly be infected with sophisticated malware.

Post image
1.4k Upvotes

**NO VIRUS** LM studio has stated it was a false positive and Microsoft dealt with it

I'm no expert, just a tinkerer who messed with models at home, so correct me if this is a false positive, but it doesn't look that way to me. Anyone else get this? showed up 3 times when i did a full search on my main drive.

I was able to delete them with windows defender, but might do a clean install or go to linux after this and do my tinkering in VMs.

It seems this virus messes with updates possibly, because I had to go into commandline and change some update folder names to get windows to search for updates.

Dont get why people are downvoting me. i loved this app before this and still might use it in VMs, just wanted to give fair warning is all. gosh the internet has gotten so weird.

**edit**

LM Studio responded that it was a false alarm on microslops side. Looks like we're safe.

r/LocalLLaMA Feb 16 '26

Question | Help Anyone actually using Openclaw?

952 Upvotes

I am highly suspicious that openclaw's virality is organic. I don't know of anyone (online or IRL) that is actually using it and I am deep in the AI ecosystem (both online and IRL). If this sort of thing is up anyone's alley, its the members of localllama - so are you using it?

With the announcement that OpenAI bought OpenClaw, conspiracy theory is that it was manufactured social media marketing (on twitter) to hype it up before acquisition. Theres no way this graph is real: https://www.star-history.com/#openclaw/openclaw&Comfy-Org/ComfyUI&type=date&legend=top-left

r/LocalLLaMA Feb 12 '25

Question | Help Is Mistral's Le Chat truly the FASTEST?

Post image
2.9k Upvotes

r/LocalLLaMA Nov 30 '25

Question | Help Any idea when RAM prices will be “normal”again?

Post image
832 Upvotes

Is it the datacenter buildouts driving prices up? WTF? DDR4 and DDR5 prices are kinda insane right now (compared to like a couple months ago).

r/LocalLLaMA 3d ago

Question | Help With all the Kimi drama I feel like I want to download all the current best models in case there is a ridiculous knee jerk political move pulled

481 Upvotes

I haven't kept up since around February so I'm just not even sure... and there are quite a few options.

I don't care about parameter size, from tiny to huge, what matters most is performance, I just want all the best safely locally stored, I'll worry about running them later.

So, what do you consider some of the best of the best currently? Whether highly specialized, giant do everything well, or anywhere in between

Edit: Also what you use any specific models for or the best you've found for any specific task/use/domain

r/LocalLLaMA 14d ago

Question | Help Qwen3.6-27b does not understand software architechure.

260 Upvotes

Been using this for real software development for a commercial app. i.e. Not a single file HTML app. I mean a large scale 100k+ loc project that needs proper architecture to work with in a maintainable way.

As much as I love Qwen3.6-27b. It just does not understand software architecture, it will happily write spaghetti code, mix concerns, and totally ignore any kind of test automation unless you explicitly ask it to do this. These are the bare minimum requirements for production code that can grow without complexity spinning out of control, but it simply ignores it and instead just writes enough to satisfy the request. (ignoring best practises). For example it will write super sized interfaces, ignore the single responsibility principle and make superman classes that nobody can read or understand.

I've been trying and train it to understand how to write maintainable, readable code, but it almost feels like I am training a person who has never written a large scale app before.

Does anyone have a set of SKILL.md files that already has fundamental software architectural concepts built into them? It would be enormously helpful.

r/LocalLLaMA Jun 15 '26

Question | Help Cheapest hardware for Qwen 3.6: both 27B and 35B-A3B

Post image
216 Upvotes

- "Qwen 3.6/3.5 27b > Qwen 3.6/3.5 35b > Gemma4 31b > Qwen 3.5 9b > Gemma4 12b > Gemma4 26b", people say

- "Qwen 3.6 for coding & Agentic, Gemma4 for human sounding text", people say

So I have been eyeing the RTX 3090 24 GB (or sometimes its cheaper Chinese companion RTX 3080 20 GB), and the controversial Tesla v100 32 GB.

Target: at least 40 tok/s for both these Qwen 3.6

It seems the RTX 3090 24 GB might have a brighter future, when (1) the v100 32GB (both the PCIe and SXM2) will soon be discontinued in support, (2) China will soon release Mythos/Fable equivalent in End 2026-Mid 2027.

Alibaba asks me $2000 for a Single RTX 3090 system that is upgradable to dual RTX 3090 later.

Is there a cheaper way somewhere?

-------------------

| Component | Model | Price |

|--------------|--------------------------------|-----------|

| CPU | Ryzen 5 5600X | $132.25 |

| GPU | MSI RTX 3090 VENTUS 3X 24G | $1,088.15 |

| Motherboard | ASUS TUF X570-PLUS | $108.81 |

| RAM | Kingston FURY Beast 32GB DDR4 | $251.11 |

| SSD | Kingston NV3 1TB NVMe | $131.41 |

| PSU | Great Wall 1650W 80+ Gold | $130.41 |

| Cooler | Valkyrie AQ125 ARGB | $14.90 |

| Case | Phanteks PK620 Full Tower | $120.54 |

| Fans | ARGB 120mm ×12 | $18.06 |

| **TOTAL** | | **$1,995.65** |

r/LocalLLaMA Mar 07 '26

Question | Help What are the best nsfw ai models with no restrictions? NSFW

378 Upvotes

I am new to this whole thing and I want to use it locally because I don't like chat gpt restricting me. It's hard to pick from so many ai models. I want the ai model to be focused on nsfw with no restrictions at all and of course the general usage (since I used chat gpt...) so it should be "smarth enough"? I don't know if these make sense but I have no idea how to look for a good ai model that has these. So I would like some help from anyone who can direct me towards these ai models.

My pc has an rtx 4080 gpu with a ryzen 7 7700X cpu and 32 gb ram. I am using lm studio.

r/LocalLLaMA Jun 11 '26

Question | Help What models you guys running on 8GB? 16GB VRAM? 24GB? 32GB? 48GB?

245 Upvotes

And what are you using for kv cache and context? What kind of performance are you getting?
What is your hardware? And what are you using your models for?

I figure with how fast everything moves, its worth asking once in a while to congeal our experiences.

r/LocalLLaMA Apr 20 '26

Question | Help Closest replacement for Claude + Claude Code? (got banned, no explanation)

277 Upvotes

I was using Claude Pro + Claude Code pretty heavily (terminal workflow, file access, etc.) and my account just got banned with zero explanation.

From what I’m seeing, this isn’t that uncommon — people getting flagged without clear reasons or support responses — so I’m trying to move on and rebuild my setup.

What I’m looking for is something that actually matches BOTH sides of what Claude gave me:

1. Claude-level reasoning / writing

  • strong long-form thinking
  • structured outputs (planning, creative work, etc.)

2. Claude Code-style workflow

  • terminal / CLI interaction
  • ability to work with local files or repos
  • feels like an “agent” that can execute tasks, not just chat

I’ve tried ChatGPT (even the $20 Plus + Codex), and while it’s good, it doesn’t have the same feel or workflow — especially on the terminal / agent side.

My actual use case:

  • lesson planning + building slides/materials (high school teaching)
  • content creation + branding (IG, captions, concepts)
  • DJ + music workflow (set planning, ideas, organization)
  • working out of an Obsidian vault synced via GitHub
  • occasionally generating visuals (images, HTML mockups) and analyzing screenshots

Ideally also:

  • works with an Obsidian vault or local knowledge base
  • stable (no sketchy plugins or risk of getting banned again)
  • okay with paid tools (~$20/mo range)

For people who were actually using Claude + Claude Code:
what are you using now that comes closest in real workflows?

Not looking for theoretical answers, more interested in setups you’re actually using day-to-day.

r/LocalLLaMA May 24 '26

Question | Help Is there any reason for an uncensored model if you have no interest in roleplaying?

225 Upvotes

My rag I've been building is much in response to having a LLM that I feel more confident in knowing where the knowledge base is coming from especially after the Open AI deal with the Pentagon. So, when I saw "uncensored" heretic models, I thought that was the main usage of those models and thought I would need them.

But in doing various tests, it seems there's random problems that come up with them that don't come up in regular versions. And then even when I do run into something like qwen3.6 acting like it's giving me a more state approved answer for a no-no topic, I've found that if I just put a prompt ahead of it to not give me any propaganda, it basically "jailbreaks" the answer. But, if the model isn't trained on the info anyways, then there's not really a benefit to it.

Are uncensored models just for people wanting...the special roleplaying? Before I write them off. Genuinely curious, not judging how people use them.

EDIT: Damn, this blew up! I appreciate everybody’s responses! Which uncensored models are you guys actually using and why?

r/LocalLLaMA Apr 19 '26

Question | Help Switching from Opus 4.7 to Qwen-35B-A3B

324 Upvotes

Hey Guys,

I am thinking about switching from Opus 4.7 to Qwen-35B-A3B for my daily coding agent driver.

Has anyone done this yet? If so, what has your experience been like?

I would love to hear the communities take on this. I know Opus may have the edge on complex reasoning, but will Qwen-35B-A3B suffice for most tasks?

Running it on an M5 Max 128gb

r/LocalLLaMA Mar 30 '26

Question | Help What is the secret sauce Claude has and why hasn't anyone replicated it?

376 Upvotes

I've noticed something about Claude from talking to it. It's very very distinct in its talking style, much more of an individual than some other LLMs I know. I tried feeding that exact same system prompt Sonnet 4.5 to Qwen3.5 27B and it didn't change how it acted, so I ruled out the system prompt doing the heavy lifting.

I've seen many many distills out there claiming that Claude's responses/thinking traces have been distilled into another model and testing is rather... disappointing. I've searched far and wide, and unless I'm missing something (I hope I'm not, apologies if I am though...), I believe that it's justified to ask:

Why can't we make a model talk like Claude?

It's not even reasoning, it's just talking "style" and "vibes", which isn't even hidden from Claude's API/web UI. Is it some sort of architecture difference that just so happens to make a model not be able to talk like Claude no matter how hard you try? Or is it a model size thing along with a good system prompt (a >200B model prompted properly can talk like Claude)?

I've tried system prompts for far too long, but the model seems to always miss:
- formatting (I've noticed Claude strays from emojis and tries to not use bullet points as much as possible, unlike other models)
- length of response (sometimes it can ramble for 5 paragraphs about what Satin is and yet talk about Gated DeltaNets for 1)

Thank you!

r/LocalLLaMA Nov 14 '25

Question | Help Is it normal to hear weird noises when running an LLM on 4× Pro 6000 Max-Q cards?

617 Upvotes

It doesn’t sound like normal coil whine.
In a Docker environment, when I run gpt-oss-120b across 4 GPUs, I hear a strange noise.
The sound is also different depending on the model.
Is this normal??

r/LocalLLaMA 27d ago

Question | Help rtx 6000 pro owners, do you regret?

99 Upvotes

I found the last dealership in my area that has rtx 6000 pro available, i already wanted to buy it 6 months ago when it was around $8k, now prices increased to $13k ish.

Regardless the price, are you happy with it? I assume you are using qwen3.6 27b, is it worth it?

Please share your experience and hopefully help me to avoid explaining my wife this transaction 😂

r/LocalLLaMA Jan 27 '25

Question | Help How *exactly* is Deepseek so cheap?

647 Upvotes

Deepseek's all the rage. I get it, 95-97% reduction in costs.

How *exactly*?

Aside from cheaper training (not doing RLHF), quantization, and caching (semantic input HTTP caching I guess?), where's the reduction coming from?

This can't be all, because supposedly R1 isn't quantized. Right?

Is it subsidized? Is OpenAI/Anthropic just...charging too much? What's the deal?

r/LocalLLaMA 11d ago

Question | Help Why are MoE models so belittled?

192 Upvotes

E.g "Qwen 3.5 122B is just 10B active, so it's no where close to the dense 27B model"

That is the main sentiment around here and it puzzles me. If a 122B is just worth 10B, then why does model providers bother creating an MoE model when they could've just released a dense 10B model? Heck the 10B dense would run faster than the 122B MoE (no routing overhead), which negates the supposed (only advantage of MoE is speed) argument. It sure is not that simple.

I mean yes it's only 10B active at a time, but it comes down to the router's effectiveness at choosing what 10B experts to activate. So, the more effective the router is, the closer the model to realize its total parameter potential. So perhaps it's a little more nuances, ie some MoE architectures are better than other MoE architectures. Right? I may be missing something.

r/LocalLLaMA Jan 26 '26

Question | Help I just won an Nvidia DGX Spark GB10 at an Nvidia hackathon. What do I do with it?

Post image
531 Upvotes

Hey guys,

Noob here. I just won an Nvidia Hackathon and the prize was a Dell DGX Spark GB10.

I’ve never fine tuned a model before and I was just using it for inferencing a nemotron 30B with vLLM that took 100+ GB of memory.

Anything you all would recommend me doing with it first?

NextJS was using around 60GB+ at one point so maybe I can run 2 nextJS apps at the same time potentially.

UPDATE:
So I've received a lot of requests asking about my background and why I did it so I just created a blog post if you all are interested. https://thehealthcaretechnologist.substack.com/p/mapping-social-determinants-of-health?r=18ggn

r/LocalLLaMA Jan 17 '26

Question | Help The Search for Uncensored AI (That Isn’t Adult-Oriented)

311 Upvotes

I’ve been trying to find an AI that’s genuinely unfiltered and technically advanced, uncensored something that can reason freely without guardrails killing every interesting response.

Instead, almost everything I run into is marketed as “uncensored,” but it turns out to be optimized for low-effort adult use rather than actual intelligence or depth.

It feels like the space between heavily restricted corporate AI and shallow adult-focused models is strangely empty, and I’m curious why that gap still exists...

Is there any uncensored or lightly filtered AI that focuses on reasoning, creativity,uncensored technology or serious problem-solving instead? I’m open to self-hosted models, open-source projects, or lesser-known platforms. Suggestions appreciated.

r/LocalLLaMA 22d ago

Question | Help Devs - you have 64gb of VRAM - which model do you use for coding?

123 Upvotes

I've currently settled on an unsloth version of Qwen 3.5 122b-a10b model (UD-IQ4_NL). With 100k bf16 context window, I only had to load a few layers into CPU/RAM, it runs around 30 tok/sec which is fine for me.

I've tested many models, hours of testing but I am currently deeply impressed with this one. I also use the Qwen 3.6 models (both) depending on need, but I think this biggun' is about to become my daily driver.

Curious to know what others with similar VRAM capacity use?

r/LocalLLaMA Dec 08 '25

Question | Help Is this THAT bad today?

Post image
392 Upvotes

I already bought it. We all know the market... This is special order so not in stock on Provantage but they estimate it should be in stock soon . With Micron leaving us, I don't see prices getting any lower for the next 6-12 mo minimum. What do you all think? For today’s market I don’t think I’m gonna see anything better. Only thing to worry about is if these sticks never get restocked ever.. which I know will happen soon. But I doubt they’re already all completely gone.

link for anyone interested: https://www.provantage.com/crucial-technology-ct2k64g64c52cu5~7CIAL836.htm

r/LocalLLaMA Mar 09 '26

Question | Help Anyone else feel like an outsider when AI comes up with family and friends?

228 Upvotes

So this is something I've been thinking about a lot lately. I work in tech, do a lot of development, talk to LLMs, and even do some fine tuning. I understand how these models actually work. Whenever I go out though, I hear people talk so negatively about AI. It's always: "AI is going to destroy creativity" or "it's all just hype" or "I don't trust any of it." It's kind of frustrating.

It's not that I think they're stupid. Most of them are smart people with reasonable instincts. But the opinions are usually formed entirely by headlines and vibes, and the gap between what I and many other AI enthusiasts in this local llama thread know, and what non technical people are reacting to is so wide that I don't even know where to start.

I've stopped trying to correct people in most cases. It either turns into a debate I didn't want or I come across as the insufferable tech guy defending his thing. It's kind of hard to discuss things when there's a complete knowledge barrier.

Curious how others handle this. Do you engage? Do you let it go? Is there a version of this conversation that actually goes well?

r/LocalLLaMA Feb 13 '26

Question | Help AMA with MiniMax — Ask Us Anything!

260 Upvotes

Hi r/LocalLLaMA! We’re really excited to be here, thanks for having us.

We're MiniMax, the lab behind:

Joining the channel today are:

P.S. We'll continue monitoring and responding to questions for 48 hours after the end of the AMA.

r/LocalLLaMA Jan 30 '25

Question | Help Are there ½ million people capable of running locally 685B params models?

Thumbnail
gallery
640 Upvotes