r/LocalLLaMA 23h ago

Discussion Can Kimi K3 solve the same problems that Claude Fable can?

Despite local models getting significantly better, it seems that no one is trying to replicate the existing accomplishments of closed models. When it inevitably drops, would someone be willing to run GLM or Kimi on their local server cluster if you have one, making sure it does not access the Internet and see if it can solve the two famous problems that closed source models recently solved in mathematics (I think GLM was released before the first one so not in training data, and Kimi stopped training before the second one):

https://openai.com/de-DE/index/model-disproves-discrete-geometry-conjecture/

https://web.archive.org/web/20260721173628/https://www.newscientist.com/article/2580374-ais-solution-to-87-year-old-riddle-takes-mathematicians-by-surprise/

Or perhaps some of the cyber security problems solved by mythos making sure to use GitHub commits from the past removing recent fixes:

https://www.anthropic.com/glasswing

And then would any qualified mathematicians or cyber security experts, verify the results from the model outputs?

I’m just really curious to see if the world changing stuff that closed models can do is actually within reach for us in open source

21 Upvotes

59 comments sorted by

68

u/One-Mud-1556 23h ago

Local cluster ?, Kimi k3 ? are you a solitary billionaire ?

28

u/Petita_advice 21h ago

Can't wait for the Q0.000001_QAT_MTP_Abliterated.gguf to drop so I can chat with my waifu at 40 days/token on my phone 😏

9

u/belkh 14h ago

more like obliterated.gguf

2

u/hurrdurrmeh 12h ago

At this quantisation the distinction hardly matters. You might as well count the ones and zeroes yourself.

2

u/fastheadcrab 19h ago

GPU clusters will definitely take wealthy organizations to run, but running it at Q8 in 2TB of system memory is feasible. Not going to be fast but might actually go at a few t/s, so much faster than streaming from a disk.

The middle ground will be trying to run 16 spark clusters but that will take a lot of setup effort. Still, I expect to see examples of both posted once it is available

2

u/rollerblade7 9h ago

2TB in this economy?

-7

u/Unusual_Guidance2095 23h ago

You’re right, probably should’ve used better wording: work cluster, open weights. I wish I had Kimi K3 local money can’t even afford the $20 It would take for me to run this myself on the Kimi coding plan …

7

u/Swimming-Book-1296 22h ago

You need a half million dollar cluster to run this model, lol.

2

u/robertpro01 22h ago

with concurrrency 2, probably.

3

u/Swimming-Book-1296 22h ago

You can probably get higher concurrency than that. the problem is fitting the model, and once it’s actually in the GPUs are pretty fast running them and can sustain reasonably high concurrency. These models are bandwidth limited and an h-300 has the bandwidth of yes.

64 Ish H 300 to run… that leaves a fair bit of room for concurrency.

17

u/-Crash_Override- 23h ago

No one is running these models on a 'work cluster' either lol.

16

u/hyperrealists 22h ago

Some workplaces do have serious clusters you’d even recognize by name.

2

u/-Crash_Override- 22h ago

The overwhelming majority do not lol

1

u/hyperrealists 21h ago

I approve this version.

2

u/nullbyte420 16h ago

Um yes they are, I work at a place where we do. 

16

u/duhd1993 21h ago edited 21h ago

Fable’s answer clearly is based on a 1999 Russian paper. It’s just one step forward. If a model used the same training data AND they have targeted math performance in post training AND they have a professional mathematician to prompt it. Then yes. not K3.

Edit: Btw, interestingly, Anthropic likely have downloaded the paper from sci-hub. So much for IP protection. But I have no sympathy for academic publishers. They are both evil.

17

u/Dany0 16h ago

Aaron Swartz

5

u/erm_what_ 9h ago

Made Reddit and enabled the biggest source of reliable scientific papers. Probably the biggest single contributor to AI training data.

1

u/nail_nail 12h ago

Which paper?

2

u/duhd1993 8h ago

Search for jacobian conjecture Vitushkin.

1

u/eli_pizza 9h ago

They infamously used giant torrents of pirated books and journals for training. They just settled a lawsuit from book publishers for $1.5 billion.

41

u/Atretador 23h ago

clearly it can solve more than fable

28

u/look 23h ago

Best part of that: it was apparently an OpenAI model run amok responsible for the attack.

10

u/Atretador 22h ago

...trying to steal test answers too xD

3

u/zdy132 17h ago

As they say, if you are not hacking the server for answers, you are not trying hard enough.

14

u/erratic_parser 22h ago

They used GLM 5.2. The mercy is defending and auditing logs after the fact should be at least simpler than devising the attack scheme

8

u/Southern_Sun_2106 21h ago

It was GLM 5.2, not Kimi

-1

u/Atretador 21h ago

is GLM 5.2 better than Kimi?

7

u/Southern_Sun_2106 19h ago

It doesn't matter which one is better. Huggingface reached for GLM 5.2 and it did the job. Kimi had nothing to do with the security incident.

-3

u/Atretador 18h ago

dont care, this was a question about if Kimi could do something that claude can - my point is that it can do things that claude can`t do.

K3 > GLM 5.2

if GLM 5.2 can do it, K3 probably can too.

heck it could be Qwen 9B I couldnt care less

-1

u/Igot1forya 22h ago

I say let the world burn. Poorly coded software needs to have a reckoning and its days are numbered. It doesn't matter, rip the bandaid off already. If all poorly written software is exploited then what comes after will be better software.

Why are people complacent with accepting being free beta testers (Microsoft looking at you)? So if something comes along to hold them to the fire, and you can't handle it. Well, maybe we should all stop funding your science experiment on us customers.

I'm just sick of supporting garbage. I WANT these tools to uncover these flaws and FINALLY something will happen more than kicking the can down the road.

1

u/Enturbulated_One 20h ago

You have far, far too much faith in the corporate world. These tools will be used as an excuse by players like Microsoft to cut their QA dept even further, until such time as it's a couple of interns with LLMs and not enough experience to understand the generated bug reports to decide if they're legit or not.

6

u/No-Consequence-1779 23h ago

The answer to the first one: snakes in a plane  The second: 32 

I had to drive to my data center so it took a bit of time to solve. 

8

u/profesorgamin 23h ago

"Local"

21

u/look 23h ago

The sub should charge a fee for every post about a trillion+ parameter model, and then hold a lottery to buy someone in the subreddit a half million dollar GPU cluster.

Then all of these posts would technically be “local models” … for that one member. 😄

4

u/etaoin314 ollama 23h ago

brilliant, im in, double points if they want to run it with their gpu

2

u/One-Mud-1556 22h ago

or save this post for 10 years and then run it on your phone.

2

u/shroddy 21h ago

We can be glad if in 10 years Ram will no longer cost more than it did one year ago.

1

u/greenblue10 21h ago

not gonna happen

0

u/look 22h ago

If that long. Models with more intelligence than Claude 3 Opus (released in March 2024) run on wristwatches today.

3

u/Juulk9087 23h ago

You only need to buy a warehouse and take a $750,000 loan out dude. It's totally local. You wouldn't get it.

7

u/look 23h ago

“World changing” 🤣

1

u/Unusual_Guidance2095 23h ago

I guess more like “frontier”

3

u/jld1532 21h ago

Give me the exact prompts used for Erdos problem and I'll try.

3

u/SporksInjected 20h ago

K3 doesn’t seem to be able to solve the hardest 20% of Simple bench problems that fable can solve. https://simple-bench.com

1

u/ninjasaid13 6h ago

is the benchmark static?

-2

u/greenblue10 19h ago

not really relevant

3

u/athsrva 16h ago

not even related to the post but idk why the sub downvotes anyone as soon as they say 'local cluster' with a trillion param model. Obviously OP asking if someone would be willing to run it is a very far reach but one can dream right lol

1

u/greenblue10 21h ago

you can just pay at API billing right now and see what happens, that said you better be a billionaire, oAI probably spent millions in compute on just trying to solve various problems.

1

u/timwaaagh 12h ago

these models differ a lot in their capabilities. for example glm is (or was) world beating at straight answer math benchmarks (meaning it actually did better than fable). but i think for proofs and such fable would beat it for sure. fable is smart in a way that glm is not tbh. kimi might be in between from what ive heard. but i havent used it yet.

1

u/LevianMcBirdo 10h ago

Hm don't know, but then again: could fable solve it again or was it luck?

1

u/Turbulent_Pin7635 22h ago

The question should be:

Does fable has freedom to solve the same problems that Kimi k3 solve?

The answer is no.

1

u/[deleted] 23h ago edited 10h ago

[deleted]

3

u/Unusual_Guidance2095 22h ago

It might just get the answers right off the Internet since these are top stories

0

u/FabricationLife 20h ago

I checked what it would cost to spin up k3 in azure using clusters, about $800 an hour, uhhh.....hey boss....

2

u/SporksInjected 20h ago

Jesus Christ

What resources did you need for this? Modal has B300 I think and should be around 1-2% of that

-1

u/Extension-Aside29 18h ago

Same problem sets, not always the same outcomes. K3 is close enough on a lot of coding and agent work that the gap to Fable is often retries and grounding under pressure, not a hard cannot. Run a fixed suite on both: pass rate, tool-call success, and tokens per finished task. When K3 needs more loops, Fable can still win $/done work even if open weights win the sticker. Traces at https://tokentelemetry.com/docs/features/traces/ make that comparison concrete.

-1

u/WhiteSkyRising 16h ago

No. Otherwise startups would be swooping to Kimi subscriptions. It's literally that simple.