Hi Adam:
Thanks – I hadn’t realised that URL links to Perplexity’s answers aren’t transferable.
It is certainly the case that Perplexity (which is based on ChatGPT4 I believe) is much more “knowledgeable” about Percy Ludgate than ChatGPT was – not least because it has ingested
subsequent additions to the Internet.
Indeed, given the huge amount of effort and money being poured into LLM research we surely should expect the latest models to be significantly improved.
And Perplexity’s answer to my question “Are the references that Perplexity provides always real ones” was very reasonable, one might
even say “thoughtful”, though its cautionary comments have of course to be applied to this answer itself – a nice example of recursion.
However, the NY Times had a very interesting, and surely worrying (for the Artificial Intelligentsia, especially) piece yesterday entitled: “A.I. Is Getting More Powerful, but Its Hallucinations
Are Getting Worse: A new wave of “reasoning” systems from companies like OpenAI is producing incorrect information more often. Even the companies don’t know why.”
I hope this link to it will work:
And on the subject of “hallucinations” (or rather “fabrications”), I also strongly recommend Gary Marcus’s very recent piece:
Why
DO large language models hallucinate?
(https://garymarcus.substack.com/p/why-do-large-language-models-hallucinate)
Let me end by quoting from it:
“Because LLMS statistically mimic the language people have used, they often fool people into thinking that
they operate like people. But they don’t operate like people. They don’t, for example, ever fact check (as humans sometimes, when well motivated, do). They mimic the kinds of things of people say in various contexts. And that’s essentially all
they do. . .
Of course the systems are probabilistic; not every LLM will produce a hallucination every time. But the problem is not going away; OpenAI’s recent o3 actually
hallucinates more than
some its predecesssors.
The chronic problem with creating fake citations in research papers and faked
cases in legal briefs is a manifestation of the same problem; LLMs correctly “model” the structure of academic references, but often make up titles, page numbers, journals and so on — once again failing to sanity check their outputs against information
(in this case lists of references) that are readily found on the internet. So to is the rampant problem with numerical errors in financial
reports, documented in a recent benchmark.
Just how bad is it? One recent study showed rates of hallucinations of between 15% and 60% across
various models on a benchmark of 60 questions that were easily verifiable relative to easily found CNN source articles that were directly supplied in the exam. Even the best performance (15% hallucination rate) is, relative to an open-book exam with
sources supplied, pathetic. That same study reports that, “According to Deloitte, 77% of businesses who joined the study are concerned about AI hallucinations”.
If I can be blunt, it is an absolute embarrassment that a technology that has collectively cost about half a trillion dollars can’t do something as basic as (reliably) check its output against wikipedia or a CNN
article that is handed on a silver plattter. But LLMs still cannot - and on their own may never be able to — reliably do even things that basic.”
Cheers
Brian
--
School of Computing, Newcastle University, 1 Science Square, Newcastle upon Tyne, NE4 5TG
EMAIL =
Brian.Randell@ncl.ac.uk PHONE = +44 (0)786 7805578
URL =
https://www.ncl.ac.uk/computing/staff/profile/brianrandell.html