Intellectual Automation in Science: Software, AI, and the Future of Scientific Work
AI and Science| Hugh Parsons
Software has been automating components of the scientific process for decades. Tools like BLAST for sequence alignment, Gaussian for quantum chemistry, and SHELX for crystallography each took a specific, well-defined task and made it faster, cheaper, or both [1]–[3]. Scientists adopted them readily because they fit into existing workflows: the software handled repetitive intellectual labour, allowing the scientist to dedicate resources elsewhere. Large language models (LLMs) are different. By approximating the reasoning patterns present in their training data, LLMs operate on the kinds of loosely structured, judgment-laden tasks that were previously the exclusive domain of human scientists — reviewing literature, generating hypotheses, writing manuscripts, designing analyses [4]. For the first time, this makes them plausible candidates for integrating multiple stages of the scientific process into closed-loop automated systems, a goal that AI researchers have pursued since the 1960s without practical success [5].
The pace and opacity of these developments make them difficult to evaluate. Claims range from imminent "scientific superintelligence" to dismissals that the technology is "just marketing" [6], [7]. Neither extreme is well-supported, but distinguishing between them requires context that most working scientists lack time to assemble.
In this article, we look at the history of software and AI in science — what was adopted, what failed, and why — to identify a consistent pattern that can inform expectations about the current wave: tools that automate well-defined computation get adopted universally; tools that claim to automate scientific judgment do not. Every successful tool raised the bar for human contribution while reshaping who thrived and who was displaced. Understanding this pattern matters for scientists at any career stage.
The Arc of Specific Automation
Since the 1960s, AI researchers have built systems intended to automate scientific discovery itself. DENDRAL generated candidate molecular structures from mass spectrometry data. BACON rediscovered empirical laws by searching for numerical regularities. LHASA planned organic synthesis routes through retrosynthetic reasoning [8]–[10]. The concept behind LHASA — retrosynthetic analysis, Corey's formalisation of the practice of reasoning backwards from a target molecule — genuinely transformed how chemists think [10]. But that conceptual contribution spread through teaching and textbooks, not through the software. LHASA itself was expensive and brittle, requiring every retrosynthetic transform to be hand-coded by expert chemists. Modern tools like SciFinder approach the same problem through database search, a fundamentally different and simpler approach. The pattern was repeated among early AI systems across scientific domains. The intellectual contributions survived and shaped future research, but the implementations did not become the broadly adopted tools we are familiar with today.
The software that transformed scientific practice was seemingly more mundane, focused on solving specific, well-defined problems. BLAST, a sequence-matching algorithm, became one of the most highly cited papers ever due to its utility [1]. SHELX automated crystallographic structure solution [3]. Gaussian made quantum chemistry computationally accessible, creating a new kind of scientist — the computational chemist — while imposing new expectations on experimentalists, who increasingly faced reviewer demands for computational support they hadn't been trained to provide [2]. Open-source ecosystems like SciPy democratised access to tools that had previously required expensive commercial licenses [11]. Non-LLM deep learning extended this trajectory: AlphaFold predicts protein structures with near-experimental accuracy across 200 million proteins, but it automates one specific task and has not made experimental structural biology obsolete [12]. In each case, automation raised the bar for human contribution. Once structure solution became routine, you needed to do something with the structure to publish. Once gene identification became trivial, the thesis had to be about gene function. The crystallography case illustrates the human cost most clearly: solving a crystal structure had been a PhD-level achievement, sometimes an entire thesis. By the 2000s, routine structures could be determined in an afternoon. Crystallography shifted from a specialist discipline to a service function in many departments [13], with consequences for researchers whose primary expertise lay in the craft of structure solution. While these tools did not make scientists redundant, they made specific contributions redundant — and the people whose careers depended on those contributions had to adapt or move on.
Software tools succeeded partly because they were modular — scientists composed BLAST, Gaussian, and R flexibly within their own workflows — whereas early AI systems bit off more than they could chew by tackling more general problems. However, we should note that these software tools still carried their own risks. SPSS made statistical analysis frictionless, enabling the undisciplined data exploration that contributed to the replication crisis [14]. For engineers, finite element analysis software can produce precise, authoritative-looking outputs from poorly specified inputs — a problem serious enough that modelling errors have contributed to real structural failures [15]. Reduced cognitive friction and unexamined authoritative outputs: both dynamics are now repeating with LLMs.
LLMs and Agent Systems
LLMs operate through natural language rather than rigid input specifications, which in theory lets them integrate across the scientific process in ways no previous tool could attempt. So far, scientists are using them much more narrowly. Surveys indicate that usage is dominated by writing, editing, and translation. In one survey of over 2,000 researchers, 93% of those using LLMs said they helped with writing or reviewing — with the largest productivity gains going to non-native English speakers, who post 43–89% more papers after adoption [4], [16]. Code generation is the second major use case. Hypothesis generation and experimental design, though most hyped, appear least prevalent [16]. Scientists are adopting LLMs for well-defined peripheral tasks first, just as they did with every previous generation of tools.
Already, these patterns are creating quality problems. An estimated 7–17% of peer reviews at top machine learning conferences in 2024 were substantially LLM-generated, concentrated among low-confidence reviewers submitting close to deadlines [17]. These reviews tend toward bland positivity, compressing score distributions and making it harder to distinguish strong from mediocre work [18]. For 2026, major AI and ML venues have modified their policies in order to deal with this dynamic.
Perhaps the most ambitious frontier is the full automation of the research loop. In 2023, ChemCrow and Coscientist demonstrated LLM-coordinated execution of known chemical syntheses within weeks of GPT-4's release [19], [20]. Recent projects like Denario can produce end-to-end research papers across astrophysics, biology, and materials science, with roughly 10% of outputs deemed interesting by domain experts [21]. ScienceClaw × Infinite (2026) tries to go further, deploying decentralised agent swarms that coordinate without central planning across 300+ scientific tools [22]. Research in this area continues every day, but fundamental limitations persist. LLMs cannot build up an internalised understanding over the course of a project the way a human researcher does, and they reason poorly under quantitative uncertainty. As it stands today, they can be prone to producing polished outputs with shallow judgment, and struggle to tell a genuine surprise from an error in their own pipeline [21], [23].
We should also be aware of the money in this space. FutureHouse raised $70M and Sakana AI recently raised $135M [24], [25]. Commercial ventures are betting on closing the physical loop too. Lila Sciences ($550M raised, "AI Science Factories"), Periodic Labs ($300M seed), and ChemLex ($71M raised) are building autonomous laboratories [6], [26], [27]. But as of early 2026, MIT Technology Review reports that there has been no breakthrough — "no discovery of new miracle materials or even slightly better ones" — despite enormous funding [6].
If we extrapolate the current trajectory, the near-term future of AI in science is likely characterised less by autonomous discovery than by a widening asymmetry between production and verification. The cost of generating hypotheses, candidate analyses, and draft interpretations is falling rapidly; the cost of confirming whether any of it is true is not. Physical experiments still take the time they take. Peer review, already strained, faces a growing volume of outputs that look increasingly competent on the surface. The binding constraint in most fields will migrate from "can we generate ideas?" toward "can we test them, and can we tell the good ones from the plausible-sounding ones?"
Agent systems will likely improve at well-defined subtasks — literature synthesis, standard statistical analyses, routine computational screening — before they make meaningful progress on open-ended judgment. For a period that could last years, we should expect a landscape where significant components of research are highly automatable but the integration across them still requires a human who understands the whole picture. The autonomous lab ventures may eventually close the physical loop for narrow, well-characterised domains, but genuinely exploratory experimental science — where you don't know what you're looking for until you find it — remains difficult to automate for the structural reasons outlined above. What emerges is not the replacement of the scientist but a reshaping of what the scientist's time is actually spent on: less production, more evaluation; less execution, more judgment about what to execute.
Impact on Scientists and Science
With this history in view, we can begin to reason about the future scientists might face. At the level of plain LLMs, we don't need to speculate much. The adoption was fast enough to observe real consequences. Scientists are using them overwhelmingly for writing, editing, and code generation, with hypothesis generation and experimental design lagging far behind [16]. A 2025 study found that scientists who use AI publish three times more and receive nearly five times more citations, but collectively AI adoption shrinks the range of topics studied and reduces engagement between scientists [28].
What's worth dwelling on is the nature of the friction being removed. Previous tools automated computation, but you still had to decide what to compute. LLMs automate the production of plausible reasoning itself. When you can generate a passable interpretation of your results in seconds, the temptation is to iterate on AI outputs rather than sit with your data and develop your own understanding. The mode of work shifts from thinking then producing to generating then curating, and we don't fully understand what that does to the depth of scientific understanding over time. We can, however, see early symptoms: the degradation of peer review quality at major conferences [17], [18], the influx of polished papers with little scientific substance [4], and the pressure every researcher feels to adopt these tools simply to keep pace with those who already have. These aren't unrelated problems — they're all consequences of the same reduction in the cost of producing work that looks like thinking.
This matters for anyone early in their career, because the formative experience of actually thinking through a problem — slowly, and sometimes painfully — is what builds the judgment that no tool currently replicates. Just as engineers trained entirely on simulation software struggled to develop the intuition to catch errors, scientists who outsource their reasoning to LLMs from the start risk never building the internal model they need to evaluate what the tools give them. This holds whether you're using an LLM to draft your discussion section or an agent system to run an entire research pipeline. The closed-loop systems extend the range of what can be automated, but the bottleneck remains the same: can you think well enough to know when the output is wrong?
So how do you cultivate that? Not by rejecting the tools — that's neither practical nor desirable. The pressure to publish, to keep up, to remain competitive in a system shaped by metrics and scarcity is real. You cannot simply opt out. But you can be deliberate about where you invest your cognitive effort. Build deep reasoning in your core area — the kind that comes from working through problems yourself, failing, and understanding why. Build broad knowledge across adjacent fields, because the connections between disparate areas are where the most important scientific insights tend to emerge, and they're precisely what current AI systems are worst at making reliably. These two investments — depth and breadth — give you the foundation to independently verify what the tools produce, and the confidence to trust your own judgment when it conflicts with a fluent but shallow AI output.
The more speculative question is what might happen as systems move towards fuller automation. Projects like Denario aim to replicate many capabilities of a junior researcher, and the funding pouring into autonomous laboratories suggests that closed-loop science is being taken seriously as a commercial proposition [6], [26], [27]. We don't know where these systems will plateau, but the uncertainty matters less than it seems. Whether they stall as unreliable brainstorming tools or mature into competent executors of routine research, the bottleneck shifts the same way: toward the judgment that decides what's worth investigating, whether the results are trustworthy, and what they mean. The earlier advice holds regardless.
More broadly, science itself isn't going away. Modern progress is built on innovation, and there are early signs in software engineering that AI-driven productivity gains may expand total activity rather than simply displacing workers — a pattern economists call the Jevons paradox [29], [30]. Whether science sees a contraction or an expansion in the coming decade is genuinely uncertain. But in either case, the scientists who will be most valuable are those who have invested in the ability to think independently, verify critically, and connect ideas that a model trained on existing literature would never think to combine.
[1] Q. He, Y. Liang, J. Jian, A. Rajagopalan, D. Kassab, S. Duron, and Y. Yang, "Mechanical behavior and piezoelectric potential of 3D-printed bouligand structures with Rochelle salt," in Proc. ASME Int. Manuf. Sci. Eng. Conf. (MSEC2026), State College, PA, USA, Jun. 2026.
[2] A. Andrusyk, "Piezoelectric effect in Rochelle salt," 2011. [Online]. Available: https://doi.org/10.13140/2.1.4242.0168.
[3] A. M. Helmenstine, "How to make Rochelle salt (sodium potassium tartrate tetrahydrate)," Science Notes, [Online]. Available: https://sciencenotes.org/how-to-make-rochelle-salt-sodium-potassium-tartrate-tetrahydrate/.
[4] C. A. Beevers and W. Hughes, "The crystal structure of Rochelle salt (sodium potassium tartrate tetrahydrate NaKC₄H₄O₆·4H₂O)," Proc. Royal Soc. A, vol. 177, pp. 251-259, 1941.
[5] F. Mo, R. H. Mathiesen, J. A. Beukes, and K. M. Vu, "Rochelle salt- a structural reinvestigation with improved tools. I. The high-temperature paraelectric phase at 308 K," Acta Crystallogr. B, vol. 71, no. 1, pp. 17-26, Feb. 2015. doi: 10.1107/S2052520614024438.
[6] D. Damjanovic, "Contributions to the piezoelectric effect in ferroelectric single crystals and ceramics," J. Am. Ceram. Soc., vol. 88, no. 10, pp. 2663-2676, 2005. doi: 10.1111/j.1551-2916.2005.00671.x.
[7] R. Sadanaga, "The crystal structure of potassium sodium dl-tartrate tetrahydrate, KNaC₄H₄O₆·4H₂O," Acta Crystallogr., vol. 3, no. 6, pp. 416-423, 1950.
[8] R. Styrkowiec, "The interaction between moving domain walls in Rochelle salt crystals," Phys. Status Solidi A, vol. 87, no. 2, pp. K135-K138, 1985. doi: 10.1002/pssa.2210870246.
[9] M. Carpentieri and G. Finocchio, "Spintronic oscillators based on spin-transfer torque and spin-orbit torque," in Handbook of Surface Science, Amsterdam, Netherlands: Elsevier, 2015, pp. 297-334. doi: 10.1016/b978-0-444-62634-9.00007-2.
[10] A. Akan and L. F. Chaparro, "Frequency analysis: The Fourier transform," in Signals and Systems with MATLAB Applications, Amsterdam, Netherlands: Elsevier, 2024, pp. 355-448. doi: 10.1016/b978-0-44-315709-7.00015-x.
Hugh Parsons recently submitted his MSc thesis in Computer Science at the University of Auckland. His research interests sit at the intersection of machine learning and AI for scientific discovery, with a focus on representation learning. Outside of research, he can hopefully be found outside.