18 Comments
User's avatar
Edmund  Nelson's avatar

The Best way to ask an LLM a question is to ask it "Do you have good sources for ____" and read those. This is something LLMs are extremely good at, as they act like epic tier librarians. Ask 2-3 different LLMs for sources and also ask for "do you have any highly regarded but not widely known sources" (sometimes it returns drivel though)

You're right that if you ask it to read something like RCA's post it will make nothing relevant (I replicated this myself with Grok 4.6 and it still refenced a few mistakes that RCA prebunked)

In the using LLMs as a spot check thing, I think that someone should mostly use them as a spot check. @Grok please confirm has been a great part of twitter as it calls out actual charlatans pretty well. But again mostly by *citing sources for you to read* rather than trusting the output blindly. But LLMs are so much better than humans at avoiding common errors and they at least give you a mildly responsible fact check on your own post before making a mistake. I've caught a bunch of errors I've made by asking Claude and ChatGPT to fact check my own posts, and it's *extremely* helpful compared to asking an editor. Since if I make a trivial mistake they'll catch it and reduce the workload on the editor/fact checker.

Humans are at this point more unreliable than LLMs but Humans make *different* mistakes from LLMs meaning that you need a Human and and LLM to fact check you, but an LLM fact check is a heck of a lot better than nothing (which was the old state of affairs a lot of the time)

Performative Bafflement's avatar

Just based on observation, LLM's pretty clearly adjust their answers based on your own vocabulary, precision, and complexity of thought, too. When I've seen some of my wife's and friends' queries compared to how I would ask the same question, I inwardly wince, and then have seen it borne out in the quality of the answer they get. I've even tested asking the same model and tier, asking it how I would ask it, and the quality of the answer does indeed change noticeably!

This could be another factor in why Cremieux can get smarter answers, and regular people can't.

And this is, at least partly, actionable - I know the people I've seen are *capable* of more thoughtful and nuanced queries, they just didn't want to put the effort in at the moment. But if it becomes generally known that "level of effort in crafting the question partially determines the quality and rightness of answer" when it comes to LLM's, people would start putting more effort in.

Silverbyte's avatar

I mean, isn't this exactly what people have always been talking about when they call LLM's stochastic parrots? They're always trying to predict the words that next fit your query and so they will usually try to reply in a tone and style that matches. (I've found much joy when chatgpt was first released seeing it do exactly that when I asked the same question in a lot of different styles)

To me this feels like the first basic thing you learn when you try to understand how LLM's function - although I understand that the majority of users don't really understand what they are working with.

Sol Hando's avatar

Damn. So you’re telling me Yarvin’s 3 hour dialogue with Claude affirming his theory of history could have been sycophancy?

Kathleen Weber's avatar

"LLMs will state mainstream views and interpretations,"

That's what they they're built to do. If they did anything else, it would be a miracle of acausality.

I call them mega blabbers. We're trapped in a stagecoach with A superficially informed fellow traveler who knows everything about everything and won't shut up until we get to San Francisco, two days from now.

Joshua Born's avatar

This phenomenon of outsourcing _actual understanding_ to LLMs is unfortunately something I am seeing not just in online discourse, but in my workplace as well. As you describe in this article, this will often lead people to be confidently wrong.

LLMs are great at pattern-matching language. They don't understand anything. They are great for code searching, web searching, and code generation (with harnesses), but they do not substitute for actually understanding a topic yourself.

One thing that struck me about your descriptions is that the LLMs were behaving like douchie Internet know-it-alls, one of whom appears to have used an LLM to criticize your work because he didn't understand it. I have seen plenty of people engaged in pre-LLM online discourse behave in the exactly the same incognizant way you described, throwing out terms and concepts they half-remembered from an article they once read or a college course they might have taken once, only to arrive at a non-sequitur that anyone who actually remembers Econ 101, etc, can easily diagnose as flawed -- but this convinces others they are smart (of course, others of the same intellectual tribe).

Crissman Loomis's avatar

Indeed. I find similar issues when addressing clear (to me) facts like moderate drinking increases longevity or hacking is defense dominant. Each time, I can point out the facts and studies, and eventually Claude will capitulate. I'm not positive that's because I'm correct, or rather that Claude has been trained to be corrigible.

Also, Claude shows poor argumentation skills. It defends its initial position instead of truth seeking, floods the response (always gotta have rule of 3 counter arguments!), and plays motte and bailey in shifting the argument when its original point is disproved.

Anna Krupitsky's avatar

give the LLM an expert reasoning playbook like :

Large observational effect - Check confounding and selection bias

Relative risk sounds dramatic - Look at absolute risk

Claimed population epidemic - Inspect population incidence directly

Effect supposedly grew over time - Check whether aggregate statistics moved

Small early studies show large effects - Look for effect-size shrinkage in larger studies

Mechanism proposed - Ask what observable implications follow

Study conflicts with population data - Investigate representativeness/generalizability

Famous paper cited - Check subsequent criticism and replication

Counterargument raised - Check whether original author already addressed it

Two variables correlate - Construct plausible DAGs before causal interpretation

Adam's avatar

I made this skill to help Claude stay inside the frame rather than try so hard to be mainstream: https://github.com/adamisom/scriptorium/tree/main/skills/inside-frame

Swami's avatar

As a retired executive (previously in charge of directing people to create and launch innovative new products and features in the financial world), the introduction of LLMs is like getting an expanded team of smart employees.

They can bring insights, data, counterarguments, new ideas, do tons of work, and so on. But to be effective, a leader must direct, override and orchestrate the process. The final outcome belongs to the person running the AI, and they better not forget this. And yes, putting an idiot in charge of an AI risks getting confident nonsense.

I will say that I see the models getting incredibly better, incredibly fast. I suspect that in a few years the models will be able to come up with CRA’s analysis on their own, and will be able to improve what he wrote with minimal prompting.

Swami's avatar

I just ran RCA’s write up through Sol and it found the counterarguments and RCA’s follow up posts. It made none of the mistakes mentioned above and ended with what I think is a very useful summary that adds nuanced value on top of what RCA wrote.

katherine's avatar

I tried to replicate your findings regarding David Sinclair's discussion of sudden blindness with Opus 5 (high thinking) and couldn't reproduce your results. You can read the discussion here, and it seems to be substantially in line with what you're saying: https://claude.ai/share/bb9ac9fd-af7b-4381-9394-720e143dd082

In particular, Opus identifies that the NAION hazard signal has only been identified in diabetic GLP-1ra users and not non-diabetic GLP-1ra users and that you can't extropalate the signal found in diabetic users to the general population. There more here too if you care to read it.

This makes me curious what models you're using since you didn't specify. In my own use, I've found that the quality of LLM responses varies greatly with model and the amount of thinking effort. I do still have occasional errors with high thinking mode on frontier models, but this is becoming less and less frequent.

Cremieux's avatar

That's good. If you click into the comments, you'll find people using ChatGPT, Grok, and Claude and getting wrong answers with each of them. You'll also find Grok eventually getting it consistently right after a few comments are made.

JaziTricks's avatar

My prompt:

Us high healthcare costs.

A blogger Random Critical Analysis has presented the argument over multiple blog pieces that the US is simply richer. And that given the utility curve of health expenditure (going up much faster than income) and free discretionary spending (way higher in the US than naive GDP comparison), the US is just on trend with the other Rich countries.

Can you read his blogs carefully and see if he is right Vs his critics?

Don't defer to authority or "generally trustworthy" etc. Look at the detailed data and the arguments and replies from both sides

Claude (opus 5 max) gave a pretty reasonable reply IMHO.

Also

"Cremieux's claim this was "prebunked" is too strong" LOL

https://claude.ai/share/8aab96f8-7a63-42d6-855f-89ca40ea9198

Cremieux's avatar

It says that in the context of not even understanding the issue I'm talking about and referring to something else instead.

It is truly incredible how bad these bots are at reasoning about this stuff.

Rob's avatar

I have also been disappointed in Claude's weak quantitative reasoning skills and have been toying with solutions. A few days ago I added this to my instructions: Before settling on a narrative for a statistical or quantitative problem, consider what the data would look like under each candidate explanation and use that to determine which candidates the data can rule out and which it can't. Use the sign and rough magnitude the mechanism could produce as checks. Distinguish "consistent with" and "conceptually related to" from "sufficient to produce."

With this addition, I have seen Claude catching itself making bad arguments (of some varieties). It proposes something on pattern-matchy qualitative grounds but then actually checks if it works as an explanation. In the few instances this happened, it had a big effect on the use of my time because it kept working and considering other explanations without me having to read bullshit and then draft an explanation why it was wrong and what should come next.

This success was on narrower, easier problems than the RCA post, so I obviously do not expect a few sentences in the instructions to solve your problem. But I do wonder whether some kind of quantitative thinking Skills document with heuristics to use, worked out examples with possible failure modes, etc. could make the LLMs much more useful.

Coel Hellier's avatar

Rather off topic, but isn’t it distinctly weird to put a comma after an m-dash, doesn’t the m-dash suffice?

延续存在's avatar

I think the central problem here may not be whether the LLM possesses enough knowledge, but whether the person asking the question has already built a sufficiently good model.

An expert can keep interrogating an answer not simply because they know more facts. They know what constraints an answer must satisfy, which variables should be compared, where an inference lacks evidence, and where the next question should arise.

That also explains the paradox you identify: the people least able to detect an LLM’s errors are often the people who need its help most. An LLM can provide an answer, but it cannot give the questioner, in advance, the model needed to evaluate that answer.

This may first be a problem of Modeling.