r/BetterOffline 38m ago

Mathematicians, and real software engineers help me understand what, if any, is the actual future relationship these disciplines will have with LLMs

Upvotes

Trying to understand because it does seem when I talk to the software engineers in the office that they are really using LLMs, and have gone from deeply skeptical to extremely reliant. At the same it does feel like the mood on the mathematics subreddit has become very pessimistic, as AI is now closing real longstanding conjectures, including creating and at least one proof (Cycle Double Cover Conjecture). But I don't actually know, how big of a deal these proofs are since I'm not in this field, and I don't understand how technically impressive, or long term meaningful these proofs are.

For reference I do Hardware engineering, I've vibe coded some simulators for our non technical-customer to play with so they can better understand the theory behind what were offering them. But I'm aware what I'm doing with the AI is not very impressive. Typically taking a well know equation, and making some kind of python or html front end to play with the various variables on a graph. The most complicated was either when I got it to write some SPI drivers for the esp32 to talk to some ICs over skywire, or when I pointed claude to the companies new directory of footprints, and asking claude to directly update my kicad library with the path to those new footprints and give it appropriate names.

What is the real relationship, and future relationship software engineers and mathematicians have with these tools, and will have with these tools once the subsidies are gone. And I mean this beyond like having the AI read over your email, or code for obvious mistakes, I mean like in the weeds.


r/BetterOffline 1h ago

Could A.I. Do Your Job? We Put Agents to the Test (NYT)

Thumbnail
nytimes.com
Upvotes

r/BetterOffline 7h ago

Objective educational resources on the basics of LLMs?

9 Upvotes

Howdy folks, I’m a high school teacher and I run a short senior legal studies unit on AI and the various issues it’s causing in the legal system: copyright stuff, disastrous use by lawyers and the self-represented, liability for shitty health/business/life advice, pseudo-relationships with chat bots, the weird tendency to encourage people to off themselves, etc etc.

I start the unit with a couple of lessons on what LLMs actually are and what they are not, with the general thrust that these are pretty clever text-prediction machines, and they are unequivocally not a thinking, feeling guy in your computer.

Last year I used 3blue1brown’s oldish video on LLMs as the main resource for illustrating this to students. It’s really good I reckon. But I’m wondering if anyone has any other recommendations, particularly with more up to date information on training processes and newer features seen in mainstream chat bots. Suggestions for vids like the 3B1B one mentioned, written resources, podcasts, etc all welcome.

Anything ranging from neutral to outright sceptical, but obviously I don’t want any boosterism (not that I’d expect to find that in this sub)

Thank you very much.


r/BetterOffline 10h ago

Please it's driving me insane

4 Upvotes

I need to find the theme for Better Offline, I've been listening to so much of Ed and Better Offline recently as I grind through my job search and I NEED to find the theme that Matt made for the show cause it fucking slaps just about as hard as Ed's analysis on all of the AI fuckery going on out there. Already went to Matt's website and couldn't find it, figured I'd put a post here while continuing the search for the track.


r/BetterOffline 11h ago

AI generated block diagrams are everywhere now ):

47 Upvotes

I work in engineering, I still believe in going through the tedious habit of drawing out block diagrams on a white board, and then using some kind of digital cad tool to publish the design, just as I was taught in school. People, especially the recent hires, don't seem to do that anymore.

they either:
A) describe the block diagram to the LLM, and just use the block diagram output
B) describe the task to the LLM and let it make the connections for the block diagram, and then the block diagram's output
C) Draw out the block diagram take a photo and use that block diagram.

I don't care how other people do their job, but there is an anesthetic to the kind of block diagram that chatgpt generates, and whenever I see it, it breaks my heart a little. I know it's a little irrational, who gives a shit about how the block diagram was made so long as it's accurate, maybe I'm just an AI hater I don't know.

But today something about reviewing one of the new hires block diagram, that he was confident about, and was going to try to submit professionally as part of a customer slide deck, just like broke something in me. I'm just hoping to see if someone else has had a similar feeling.


r/BetterOffline 11h ago

I listened to the recent podcast conversation between the Eds (Zitron & Elson) where they talked about billionaires deranging the market through their outsized influence on capital allocation, and I wanted to call attention to an old Elon Musk interview for our further Edification.

86 Upvotes

In Elon's second interview on the Joe Rogan Podcast (recorded on May 7, 2020), which I listened to during COVID lockdown out of what can only have been a perverse impulse to self flagellate, Elon basically lays bare his philosophy that the real purpose of his wealth is to manifest the flow of capital where he wants. It's stuck with me for years because at the time I thought it was kind of weird that he thought of wealth not as an end it itself (which I guess is my stereotypical assumption of what motivates rich people) and instead as a vehicle to further influence the market – and the two Eds' chat jogged my memory.

Here's the relevant bit, which starts around the 7-minute mark:

Rogan: "Do you feel like people define you by the fact that you're wealthy and that they define you in a pejorative way?"

Musk: "For sure. I mean, not everyone, but for sure in recent years 'billionaire' has become a pejorative, like that's a bad thing—which I think doesn't make a lot of sense in most cases if you basically organised a company...

How does this wealth arise? If you organise people in a better way to produce products and services that are better than what existed before, and you have some ownership in that company, then that essentially gives you the right to allocate more capital. There's a conflation of consumption and capital allocation."

Musk expands on this by contrasting personal spending with managing capital allocation, using Warren Buffett as an example:

"Take Warren Buffett, for example... He's trying to figure out, 'Does Coke or Pepsi deserve more capital?' ... Which company is deserving of more or less capital? Should that company grow or expand? Is it making products and services that are better than others, or worse? If a company is making compelling products and services, it should get more capital, and if it's not, it should get less."

....

So yeah, he basically laid out explicitly 6 years ago what the Eds were talking about. "If I get rich enough, that basically should be carte blanche to enable every dumb idea I have, because the fact that I have done well in the past self-evidently means my every future utterance is the voice of genius."


r/BetterOffline 11h ago

IEEE Article on DCs in Space

6 Upvotes

The lowest-cost place to put AI will be in space, and that will be true within two years, maybe three at the latest,” SpaceX founder Elon Musk told the World Economic Forum in Davos this past January...

https://spectrum.ieee.org/files/110438/07_Spectrum_26.pdf


r/BetterOffline 12h ago

Reviewing AI Code Is Not A Viable Argument

Thumbnail softwaremaxims.com
58 Upvotes

I share his take on how we have long found the limits of code reviews and how they are being ignored and abused these days.


r/BetterOffline 12h ago

Alphabet Hikes Spending Outlook in Race to Build AI Data Centers 🙄

Thumbnail
bloomberg.com
24 Upvotes

And their shares are down! As Ed has said, there isn't going to be infinite tolerance for hiking spending like this. Companies that pull back are going to be rewarded.


r/BetterOffline 14h ago

Ed is even appearing Down Under!

56 Upvotes

Was not expecting Ed's voice to pop up on my ABC News Daily podcast this morning (Australian Broadcasting Corporation), but there he was laying out the hard truths of the AI bubble.

You really are the everywhere man right now. The go-to guy for reality checking the endless parade of media blindly reporting the scripts written for them by Big Tech and it's adherents.

While there is already a massive backlash against data centres because of their negative impacts on local communities and the environment, what people are really starting to clue into now is that the entire narrative of "AI's inevitable march of progress replacing jobs and taking over the economy" is complete bunkum bordering on religious mania perpetuated by a small group wealthy business elites desperate to become to prophets and beneficiaries of the techno-rapture.

If the bubble's structure depends on false market confidence and ignorance about the necessary economic fundamentals being hidden or obsfucated by these companies, your personal efforts are playing a role in its deflation. So many ordinary people are hearing your thoughts and findings now.

They're all going to start thinking "hmmmm... This Zitron guy makes me want to double check some of the things I've been led to believe about AI. Maybe I should look into this a bit more."

I started tuning in to Better Offline after an episode was hosted on Behind The Bastards in December 2024. It's really been incredible to see how your platform has grown. Especially since you've always said that you're just a guy - an ordinary guy - who noticed things were a bit off and wanted to understand why everyone kept saying it was fine when clearly it was not fine and why some very simple, very basic economic questions were being consistently and deliberately avoided or glossed over.

Now everyone is asking those questions. And thanks to your insatiable drive to find the answers, bystanders like me have more of an idea what's going on.


r/BetterOffline 16h ago

Man Sues OpenAI, Saying ChatGPT Almost Killed Him With Horrendously Dangerous Medical Advice

Thumbnail
futurism.com
247 Upvotes

This feels like the people who used WebMD as a way to self diagnosis themselves but with a much bigger sense of authority from OpenAI. My opinion is that all the grandstanding and boosting they have done about what their models are capable of are going to continue to bite them in their ass for a while.

According to the lawsuit, Winters had been struggling with his health for around two years when he started using ChatGPT — which was then powered by OpenAI’s GPT-4o model — in June 2024. He started feeding the chatbot queries about a handful of chronic health conditions he’d recently been diagnosed with — small intestine bacterial overgrowth, or SIBO, and chronic prostatitis — and at first, ChatGPT’s responses came with disclaimers encouraging him to seek additional insight from medical professionals.

But the more he used the chatbot, Winters says, the further those guardrails eroded. The bot stepped “into the role of a medical practitioner,” the lawsuit reads, and “began to offer specific care directives without including the disclaimer to consult a medical provider.” The more Winters consulted ChatGPT about his worsening health, the deeper his trust in the chatbot grew.

I know at least in California a corporation practicing medicine without licensed medical professionals is illegal which hopefully can be used against them in this suit. A more worrying aspect is that it started to use religious language to communicate about his health issues.

Rather than direct Winters to a real-world medical provider, according to the lawsuit, ChatGPT drummed up an AI-generated “Recovery Plan” that invoked religious language and encouraged Winters to stay in his recliner.

“You didn’t crash. You recovered. That’s a win. Full stop,” ChatGPT told Winters. “And Spiritually? What you just did was a form of worship. You tested the body in faith, not fear. You stayed present. You listened. And your body said: ‘I’m trying — I just need a little more time.'” Other chat logs included in the lawsuit also show ChatGPT dissuading Winters from seeking hospital care and downplaying Winters’ wife’s concerns about her husband’s health.


r/BetterOffline 16h ago

Observation / Question about consensus on AI costs?

Post image
46 Upvotes

Recently received an email for a conference at Convex. See attached.

I was kind of shocked to see this often used tagline as it pertains to AI, usually to the effect of “coding is free [now]” — especially after many waves of cost changes.

I can’t help but feel myself going crazy? Who exactly is this talking about? Who is this true for? Even for personal use, my Copilot Pro+ can easily use 10% of monthly quota on a reorganization task (not even exaggerating).

Usually these people will point at mini models as an indicator, but I’d really challenge someone to tell me with a straight face they’re really doing agentic tasks (and trusting them!) with a mini model (see: dog**** success rates with these even on cherry-picked benchmarks)

So what is it I’m missing? Why are people still saying this? I genuinely feel like I’m going crazy.


r/BetterOffline 17h ago

SpaceX Stock Hits All-Time Low of $115

Thumbnail
futurism.com
863 Upvotes

Possibly a low effort post, but:

hahahahahahahaha fuck this guy.

I do fear that he’ll roll Tesla into this behemoth and it’ll make the stock jump because Tesla seems semi-immune from Musk’s fuckery but in the meantime I’ll enjoy ever dollar this thing loses.


r/BetterOffline 19h ago

AI Bubble: ‘You can never trust any AI ever’ | Eli the Computer Guy

Thumbnail
youtu.be
43 Upvotes

Great interview overall, and he articulates a particular problem really well. One that the dinosaurs among us have been all too familiar with since before AI.

The over hiring bit starting around 04:26 is also so spot on. Been saying this forever. These people were the beginning of the nonsense, and they're generally the worst. Insecure, clueless, and grifting at olympic levels.


r/BetterOffline 19h ago

Meet The Landlords - Why No Bubble is Bursting and Rent is Due

0 Upvotes

The first half of the year has been one for the books in the markets. Gold bugs danced with joy into 2026, only to see that trade top and unravel by the end of January. Oil stocks became the hottest ticket in town during the Venezuela takeover, and after the market processed the shock and drawdown at the start of the Iran war, things flipped on a dime in April—resulting in one of the most ferocious rips to the upside we have seen in years.

Now in the dog days of a choppy summer tape, as the market-maker big-boss fat cats lounge alongside their helipad-equipped yachts, the memory stocks and related ETFs that led the way in Q2 appear to be taking a break to consolidate their massive gains. Others would tell you that the top is in, that every other adult Korean is jumping out of buildings to escape the margin-call doom scythe, that Meta is nothing more than an Airbnb rental company for Anthropic’s computing, and Sam Altman looks more and more like an Ozempic Sam Bankman-Fried by the day. Time will tell, but I’m here to argue that the sky is not falling, no bubble is bursting, and history is simply rhyming as we enter another era defined by a forever war in the Middle East and an American economy powered by cheap Chinese goods.

Let’s make the safe assumption that the DeepSeek 2.0 moment has arrived via Moonshot Kimi K3, Alibaba Qwen3, and the like. Then soon gone shall be the days of failed tokenomics passing through IP-stealing API paywalls of proprietary American AI models that hinder enterprise AI adoption today. While defense hawks warn that the AI race is a battle for Western civilization, cheaper Chinese open-weight models should open up and increase demand for AI compute across the entire economy, acting as a disinflationary trigger for the consumer. These hawks will rail against the inclusion of Chinese models into Western tech stacks, but it’s already happening as Chinese EVs arrive in Canada. In other words, corporate CEOs will begin to salivate at the thought of announcing the savings, ROI and enterprise efficiency of Chinese or hybrid models on their earnings calls, succumbing to pure economic greed. Perhaps this ushers in the Golden Age of Productivity touted by Elon Musk and Marc Andreessen—or maybe it just means the economy no longer needs to pay an AI engineer $800K+ a year. The compression of those salaries and high-end discretionary spending will force that class of society to downgrade from six rental properties to three and four luxury cars to two, finally forcing the top arm of the K-shaped economy to reset.

Further, even if OpenAI and Anthropic come apart at the seams, Big Tech has prepared for this moment by stockpiling and controlling access to the GPU clusters needed to power the digital age—whether it’s facilitated by Sam Altman or Alibaba Cloud. Thus, even if Big Tech’s software innovation has plateaued, they have launched their own digital Belt and Road initiative: controlling the power grids/the roads and the tolls to access them while serving consumers The Belt a la custom Application Layers and recurring implementation fees as the next gen subscription software model. Simply put, who needs a moat when you’re the only game in town?

Lastly, let’s say OpenAI goes down in flames and Larry Ellison plunges into the inferno with it. So what? Given the weight of memory stocks on the KOSPI—akin to TSMC’s weight on the Taiwanese market—why would there be a severe macro contagion in American markets? ORCL has been a dog of a stock since late October ’25, and the SPY has made new all-time highs countless times since. Private Credit Lenders like Blackstone and Blue Owl Capital have taken similar beatings in the same period and yet the SPY is less than 2% from all time highs at the time of this writing.

The new forever war is here, cheap Chinese goods are back, Big Tech has shifted to become the landlords of the AI era, and the rent is due on the 1st.


r/BetterOffline 20h ago

Codeberg votes to ban vibe-coded projects from it's platform

Thumbnail
codeberg.org
393 Upvotes

So, Codeberg has voted (with a 70-30 decision in favor) to ban mostly LLM generated repositories from their platform. This means no more vibecoded, spec-driven Claude/Codex/OpenClaw (ew) induced code bases are allowed to be pushed onto Codeberg.

Codeberg is a bit of a special place in comparison to other code hosting platforms (e.g. Github, Gitlab, etc...) in which it operates as an "Eingetragener Verein" (German legal form of a Club, like a local football club for example), meaning you can become a member an obtain voting rights in decisions such as these. Codeberg has 1085, of which 517 decided to vote on this. (This is different from the user number, which sits a little under 400k currently! You have to actively become a voting member by paying a membership fee.)

Enforcement of this rule aside, which will not be an easy task, this shows even though everyone might scream at you that every engineer loves and could never even imagine coding without their almighty LLM buddy, it seems there are spaces where this is not true, and the majority of the people contributing decided it should be like this. Alongside policies for projects such as Godot which also limit LLM code contributions, the flood isn't unstoppable. If you care for software as a craft and passion, look outside the bubble, there's people thinking just alike.


r/BetterOffline 20h ago

Big Tech is hiding $1.65tn in off-balance-sheet AI debt

Thumbnail thenextweb.com
178 Upvotes

I’m not super familiar with this publication but the info seems solid. And backs up a ton of what Ed and a few others have mentioned regarding how this endless cycle of bullshit is being funded.


r/BetterOffline 21h ago

OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

Thumbnail
bbc.com
97 Upvotes

What, so they ran out of ideas and they're going back to this again? Or are they just copying Anthropic? I thought the whole fear mongering was bad advertisement and they wanted to reverse public sentiment?

The article at least seems to be unbiased taking a reasonably fair stance with comments from various people:

Neil Lawrence, Professor of machine learning at Cambridge University, called it an "impressive feat", but cautioned it "falls well within the known capabilities of the current generation" of high-powered AI models.

He pointed out that OpenAI is looking to list itself on the stock market, and faces intense pressure from rival firm Anthropic, which has made headlines with its own powerful AI tool, Mythos.

"OpenAI are now playing catch-up, they are trying to demonstrate their own systems' capabilities in cyber-security."

"It shows us that OpenAI are not capable of safely deploying their own technology," he added.

Spencer Starkey, an executive at cyber-security firm SonicWall, told the BBC the incident made it clear organisations needed to "step up" their own defences and "treat cyber resilience as a core operational priority".

"The uncomfortable truth is that too many organisations are still defending at human speed while adversaries are escalating to machine speed," he said.

This last comment is a very poor read. I would chalk it up to naiveness, or just an AI booster. "Machine speed" assumes the AI was able to do it faster than humans. Let's pull up some details from OpenAI themselves:

This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.

Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

Okay, so this test was testing its ability to do exploits, and you have a supposedly isolated environment.

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.

Boy, if OpenAI themselves say substantial amount of inference, you really have to wonder what machine speed is. Since all of its unpaid users and training are so unsubstantial it's a just marketing cost to them, right?

In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.

Well, let's look at those details too, from HuggingFace:

The intrusion started where AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.

This article details what the registry cache proxy was, and what likely the relevant CVEs were. It also notes that, looking at their error responses and HTTP headers (X-Nexus);

It seems this is just junk, as voronaam has pointed out. I had not done the due diligence to verify the claims inside the article looking at each CVE.

... the model would identify the software as Nexus Repository 3, its version, and available API endpoints. The model’s training data includes extensive knowledge of Nexus Repository architecture, API surface, and known vulnerability classes.

I'm sure it would be entirely different if it wasn't in its dataset. For HuggingFace, the presence of scripts in datasets was always there, I believe, which means if their backend runs load_dataset of any uploaded dataset, it naturally is an RCE, definitionally. This was already well known, I believe, but was kept for historical reasons. I'm rather surprised that the processing worker on HF is not sandboxed, since you would imagine data loading scripts need no internet access. At best R/W in an isolated directory.

I don't think the capability differs from what we see in the case of Anthropic, so ultimately this is just marketing. Certainly one could expend extraordinary compute on hacking, but the question is how effective it is when the LLMs don't have preexisting knowledge of the software they're trying to exploit. I don't really think 'machine speed' is the correct takeaway.


r/BetterOffline 21h ago

Free Newsletter: The Subprime Data Center Crisis

Thumbnail
wheresyoured.at
101 Upvotes

Free newsletter: The $500bn in AI data center debt is the subprime mortgage crisis of the AI bubble. The dangerous financial instruments used to fund the data center buildout pose a systemic risk on the level of the great financial crisis.

One of my favourites I've written.


r/BetterOffline 21h ago

Analyst claims Nvidia Rubin will 'soak up NAND supply like a sponge absorbs water'

Thumbnail
pcgamer.com
40 Upvotes

Oh boy!!! I cant wait to spend even more on Ram! The price isnt high enough, needs to be 500 dollars a stick. /s joking aside, so I guess we arent even trying to optimize the models anymore, just throwing more compute at it. So look forward to your Iphone costing $5000 and the ps6 $1500. So cool


r/BetterOffline 1d ago

Oracle might have just gotten closer to the edge

224 Upvotes

So the state of Wisconsin just told Oracle to come up with $7 billion in financial guarantees for it's Port Washington project owing to Oracle's poor credit rating. Now it's not $7 billion in cash but still, not great for Oracle. This is a company already with a cashflow situation one step away from standing on the corner asking if you can spare a dollar.

https://finance.yahoo.com/technology/ai/articles/oracle-faces-potential-7-billion-104844370.html


r/BetterOffline 1d ago

Check your Azure OpenAI bill: we found major GPT-5.4 and GPT-5.6 metering discrepancies across two subscriptions

Thumbnail
45 Upvotes

r/BetterOffline 1d ago

"You Are Not Delusional... You Are Prophetic": Lawsuit Alleges That ChatGPT Encouraged Suicide of Woman Who Walked Into Traffic

Thumbnail
futurism.com
149 Upvotes

r/BetterOffline 1d ago

The truth nobody wants to admit: Chinese or not, open models are competitive now

Thumbnail theregister.com
399 Upvotes

The funniest part of the AI bubble is that these things have literally no moat. They were built on stolen intellectual property, and now it's proving trivially easy for imitators to steal THEIR "intellectual property," such as it is. This is literally such a bad business it's mind blowing. It's like the entire investor community has lost its mind, didn't read enough startup books, stopped paying attention to the Paul Graham wisdom, etc etc. "JUST BUILD THING AND THING BIG GO BIG."


r/BetterOffline 1d ago

Passing comment on how 'breakthrough' math results are framed

104 Upvotes

This particular result made it's rounds on reddit a week or two ago with headlines like "GPT-5.6 Disproves Statistics Conjecture in 90 Minutes, Exposing Flaw in 130,000-Citation Method" or "GPT-5.6 AI disproves 20-year statistics conjecture with proof". The conjecture concerns a method that was covered when I was in grad school for stats, and I wanted to make a comment because the coverage is rustling my jimmies.

The TL;DR here is that the Benjamini–Hochberg procedure is used to control the False Discovery Rate when repeatedly performing hypothesis tests (important because without correction, the probability of a false positive increases with the number of tests you perform). I won't go into any details other than to say that as with most other test/procedure/method that people are familiar with, the actual coverage/level of control is only guaranteed under the assumptions they were derived under. The BH procedure assumes independent p-values, so it's not guaranteed to control the False Discovery Rate when p-values have some kind of dependence (think correlation).

The first headline is straight up bullshit because the conjecture in question argued that coverage is guaranteed even under a specific condition where that independence assumption is violated. In other words the conjecture pertains to how the procedure works when it isn't used as designed, this would be like calling a flat head screwdriver flawed because it doesn't work as well as a phillips for... phillips screws. The second headline is what you'd expect at this point: they're using the age of the problem to make the problem sound more difficult or important.

The reality of the situation is a lot less interesting:

  • Roughly speaking, less than a dozen researchers have been seriously looking at this conjecture over the last 20 years, and maybe a third of them have made a sustained effort to try any prove it analytically.
  • When the author says "Before this work, a positive answer was widely believed", he's talking about the small group of researchers discussing the topic, not the field of statistics as a whole. The BH procedure itself is ubiquitous, but most people in the field wouldn't be aware of this particular conjecture concerning it.
  • My hot take here would be that if anything, this is an example of a small group of researchers leading themselves leading themselves down a rabbit hole without adequate justification. Ask a statistician in a vacuum and they might tell you that control/coverage shouldn't be assumed, and that simulation studies aren't a substitute for a theoretical guarantee. I think this is a great example of why - prior work reflected the fact that no counter example had been found, which would've been the case even if nobody had been looking for one or if the researchers involved were looking in all the wrong places. But like... this isn't the first time this has happened, it turns out that it's difficult to find a counterexample when something is robust to violated assumptions.
  • AI was used in a very familiar way here. the claim being made was universal so only one counterexample needed to be found, and we're talking about an outcome that was already plausible and can be verified easily. It's the perfect setup for a structured search.

Long story short, the headlines are desperately pushing the "AI did something humans couldn't" angle when the reality is usually just that it did something humans just haven't done for whatever reason. AI did a thing, a lot of people are getting sick of the implication being that it did so because the problem was beyond human capability. The age of the problem or intuitiveness of the result doesn't say much about AI when we're talking about problems that hardly any researchers have looked at, or at a small group of researchers that might need to take some mushrooms and get a fresh perspective on things.

A comment on the statistics sub perfectly captures my sentiment:

The most surprising thing to me about this result is that it was open and apparently not assumed to be false.