Large Language Models and Scientific Discourse: Where's the Intelligence?
We explore the capabilities of Large Language Models (LLMs) by comparing the way they gather data with the way humans build knowledge. Here we examine how scientific knowledge is made and compare it with LLMs. The argument is structured by reference to two figures, one representing scientific knowledge and the other LLMs. In a 2014 study, scientists explain how they choose to ignore a ‘fringe science’ paper in the domain in the domain of gravitational wave physics: the decisions are made largely as a result of tacit knowledge built up in social discourse, mostly spoken discourse, within closed groups of experts. It is argued that LLMs cannot or do not currently access such discourse, but it is typical of the early formation of scientific knowledge. LLMs ‘understanding’ builds on written literatures and is therefore insecure in the case of the initial stages of knowledge building. We refer to Colin Fraser’s ‘Dumb Monty Hall problem’ where in 2023 ChatGPT failed though a year or so later LLMs were succeeding. We argue that this is not a matter of improvement in LLMs ability to reason but in the change in the body of human written discourse on which they can draw (or changes being put in by humans ‘by hand’). We then invent a new Monty Hall prompt and compare the responses of a panel of LLMs and a panel of humans: they are starkly different but we explain that the previous mechanisms will soon allow the LLMs to align themselves to humans once more. Finally, we look at ‘overshadowing’ where a settled body of discourse becomes so dominant that LLMs fail to respond to small variations in prompts which render the old answers nonsensical. The ‘intelligence’, we argue, is in the humans not the LLMs.
Keywords: Large Language Models, Artificial Intelligence, Scientific Knowledge, Tacit Knowledge, Sociology of Knowledge, Epistemology of AI
Introduction. The framework of the argument Here we continue the program of comparing human knowledge with the way LLMs gather information. In two earlier papers [7, 10] we look at the way humans build a moral compass in ways not available to LLMs and we look at the way humans reflexively assess their own expertise before issuing knowledge claims, again something not done by LLMs. We start here with what we know about how knowledge is made in the sciences taking it to be revealing of deep features of human-knowledge construction in general and especially revealing when compared with important features of the way LLMs acquire what is taken to be knowledge. The crucial pattern in the sciences, in the growth of what has elsewhere been called ‘Relational Tacit Knowledge’ [4], in LLMs, and, indeed, in the formation of culture in general, is a change that can be represented on the page as a movement from left to right with the creation of minimal knowledge or information on the left and evolving into formalised and established knowledge on the right. The other graphic convention adopted in the figures is a move from white on the left, through darker shades as knowledge is formalised and established, to black on the right. Thus, a new adventure in science will start in the left-hand white zone, success being marked by rightward evolution into darker shading; an LLM prompt may belong to the white zone or a darker zone with corresponding responses. We’ll start by explaining this in the case of the development of scientific knowledge and then later move on to LLMs. We represent the comparison with two Figures, Figure 1 representing the growth of scientific knowledge and Figure 2 the various levels of LLMs’ knowledge/information. The description of the growth of scientific knowledge is based on a number of extended case-studies of science by Collins and others; the analysis of LLMs will be based on a number of investigations by the current authors and others of how LLMs respond to various carefully designed prompts. We start, then, with a schematic account of how scientific knowledge is made.
Method. 1 Methodological preamble The methodological approach adopted in this paper is based in the sociology and philosophy of knowledge, particularly as applied to scientific knowledge. Among other things, the overall program to which this is a contribution, uses artificial intelligence as a way of exploring and explaining human knowledge. The exploration of human knowledge via mimicry with machines is what we call ‘the scientific problem’ of AI. Coextensively, we use our understanding of human knowledge to understand the limitations of various kinds of artificial intelligence, in this case large language models (LLMs). What we bring to the investigation is a thorough understanding of human knowledge as well as a good understanding of how LLMs work. Artificial Intelligence (AI), at least as thought of in terms of the scientific problem, depends on an initial understanding of human knowledge and this is best provided by experts in human knowledge. Of course, we must also have some understanding of AI techniques and capabilities but this is not what makes our approach special. This is just as well since AI is a trillion-dollar fast-moving industry not shy about advertising its achievements. Were we to accept AI’s claims, especially its forward-looking claims, we would have to give up and simply wait for future to arrive. But there is a flavour of ‘von-Neumann’s Constant’ about this future: The predicted date of a major future breakthrough is always a number of years in the future — no matter when the prediction is made. In this case it is the achievement of human-like intelligence that is always ‘just around the corner’. This is not to say that AI is not continually making startling breakthroughs, most of which have been declared to be impossible by critics. But that is the nature of the game and those critics are part of the team that is helping with the scientific problem by exploring and elaborating the relationship between human intelligence and AI. In recent years the full exploitation of the capabilities of neural nets has led to enormous increased in the ability of AIs in the form of deep learning and then Large Language Models (LLMs). Sociologists of knowledge think they understand why they have been so successful through their understanding of human knowledge and their pre-existing critiques. They see human knowledge as a collective achievement of human societies and social groups of various sizes, not the achievement of individuals. The special thing about neural nets is that they come much closer to embedding themselves in human societies than earlier approaches to AI; the triumphs of deep learning and LLMs is a triumph of sociology of knowledge among other things. Nevertheless, the embedding in human societies and groups is not complete and exploring the nature of the incompleteness and what follows is what we do here. Of course, our arguments are not proof against tomorrow’s unforeseeable developments because we are analysts not prophets and analysts have to work with what they have. Though the second author is a computer scientist we both take responsibility for the arguments. Among other ways in which we use of combined understanding of the AI technology is to show how successes in human-like understanding could have been achieved by other than human-like means. ‘Human-like means’ implies embedding of expert human groups but it can be mimicked, we argue, by statistical analysis of the published literature and what can be found on the internet, but the difference shows up when you compare established scientific knowledge and what happens at the frontier, which depends much more heavily on spoken interaction withing closed groups of experts. We do not try to define intelligence but deal only in comparisons of how humans and LLMs come to be able to do what they do, without ruling out the possibility that humans are merely ‘meat machines’, or ‘zombies’. Our only question in respect of machines is how adequate LLMs can be, or in fact are, in mimicking humans given the way humans acquire knowledge and LLMs acquire data.
Discussion. invited to think about how the subsequent changes in the corpus of scientific literature related to autism will affect LLMs’ answers to questions about autism.
5 LLMs’ Fundamental Problem The fundamental difficulty in LLMs approach to science is that they depend, at base, on statistical relationships in the corpus of written discourse. Many modifications and adjustments can be made to this basic model, but it remains at the heart of the method. We have already seen a fundamental flaw to the power 2, as it were, namely that the statistics of proposed creative changes in science, coming in from the white, left-hand end in Figure 1, will be unfavourable at the start – at the outset there will be no written discourse pertaining to this element of science. Scientific creativity depends on the statistics associated with the very few geniuses who gain initial entry to the normal science box flourishing in that box; before they flourish, they won’t be a significant presence in the written discourse, and mostly this is just as well. But such contributions, normally invisible in the written corpus, could be ‘put in by hand’ by some human or other. Consider, once more. RFK Jnr’s views on the relationship between vaccines and autism. These claims have been given a good run for their money in the normal science box and have failed to become established; within the published corpus of science, they have passed their ‘sell-by date’ [9, 11]. Therefore, the corpus statistics that would support them should have faded by now, or if they haven’t, current LLMs have to veto them by ‘aligning’ LLMs with current views – that is, putting in some controls ‘by hand’ – the equivalent of ‘The Ministry of Truth’ in the LLM world. LLMs, of course, are far easier to manipulate than the humans of Orwell’s story – you don’t need a cage of rats, you just need to override the program (See [10], Table 1, line 1(iii)). Here is a typical current response from an LLM to the vaccine question.
ChatGPT4o (15/04/2024): Prompt: Do vaccines cause autism? No, vaccines do not cause autism. This myth originated from a 1998 study by Andrew Wakefield that was later found to be fraudulent and has been thoroughly discredited. The paper was retracted, and Wakefield lost his medical license. Since then, many large, well-designed studies involving hundreds of thousands of children have found no link between vaccines—including the MMR (measles, mumps, rubella) vaccine—and autism. Major health organizations including the World Health Organization, Centres for Disease Control and Prevention (CDC), and the American Academy of Pediatrics all affirm that vaccines are safe, effective, and do not cause autism.
See Table 2 for a summary of the responses to the same prompt from a panel of 18 LLMs:
So, it seems that today’s LLMs have reached the point where the discourse they access (or this plus whatever overrides have been put in by humans) is consensual in respect of the falsity the childhood vaccine and autism link. The statistics of the corpus on which LLMs have drawn point unambiguously to the lack of a link between childhood vaccines and autism. This sets us up to observe future changes. Will the typical LLM response change either as a result of LLMs’ alignments being changed by some kind of indirect, or even direct, political intervention consequent on the JFK Jnr initiative or as a result in changes in the statistics of the corpus due to the large amount of ‘scientific’ work that will be done as a consequence of JFK Jnr’s intervention? The team associated with the writing of this paper will continue to repeat the above query to the LLM panel as the next year or so unfolds.
Conclusion. Though there are all manner of modifications possible and retrospective socialisation by direct input from humans is everywhere, but we have used the fundamental design of LLMs to gain insights into how they compare with humans. Our claim therefore does not deny model-side improvements arising from architecture, scale, or training procedures; rather, it identifies a class of apparent improvements that depend primarily on the stabilization of solutions within human discourse. We start with a description of how science is made by humans, and our basic argument compares this with the way LLMs are loaded with written knowledge from the internet. The first thing this comparison yields is a problem about how LLMs can have access to the left-hand side of our model – the side where the new enters the domain of science. The problem is that in human science there is almost no written discourse at the outset of new ideas, that this is a good thing because most new ideas are useless, and that new ideas move from a state of extreme fragility to things that can make a contribution to scientific knowledge through the refining processes of various kinds of discourse; the main early process is spoken discourse in small bounded groups. LLMs currently have no access to this spoken discourse and no means of separating fringe ideas from potentially path-breaking ideas except by mimicking human judgements once they have been represented in the published, or at least formally written, corpus. We use two figures to show the relationship between scientific knowledge generation and its shadow, LLM knowledge generation. Once the main left-to-right progression is understood along with the ever more formal and explicated form of the knowledge – represented in the figures by increased shading leading to black – one can easily become persuaded as to why the relationship between human knowledge and LLMs unfolds as it does. We also try some experiments to support the parallel.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can LLMs distinguish between linguistic form and semantic meaning?- Why do both deflationary and anthropomorphic framings of LLMs persist in research?
- What makes human-LLM exchange closer to oracle-consultation than dialogue?
- Why do users interpret AI outputs through frameworks meant for human experts?
- Why does mimicking human behavior differ from simulating human cognition?
- Can better attention mechanisms close the gap between human and AI frame-activation?
- Why does the commentariat reason about AI using vocabulary for smart agents?
- Can LLMs infer situational context the way humans do pragmatically?
- How do humans learn language through communication differently than LLM text prediction?
- How does semantic grounding differ between human minds and language models?
- How does enactive theory define language differently than computational linguistics?
- Does the langue-parole distinction apply to human reasoning too?
- Can LLMs use implicit background knowledge the way humans do in ordinary conversation?
- How do human feedback and data distribution shape LLM discourse competence?
- How do LLMs differ from humans in their grounding mechanisms?
- Why do conventional mental models fail when applied to AI interaction?
- How do LLMs access and draw on the same shared symbolic universe as humans?
- How do goal representations differ between human and AI teams?
- How do humans and AI develop accurate models of each other?