Research Article  ·  Artificial Intelligence

Large Language Models as Research Infrastructure: Why the Next Scientific Breakthrough May Depend on Human–AI Collaboration

For two decades, the defining challenge of biomedical research has not been the absence of ideas — it has been the inability to process the data those ideas produce. A new class of tools is beginning to change that. Not by replacing scientists, but by acting as something closer to research infrastructure.

Published · June 2026 Reading time · ~12 min Unlisted · Internal

The conversation about large language models has, until recently, orbited productivity. Can they write? Can they code? Can they replace a customer service agent? These are reasonable questions. But they miss the more consequential shift happening at a quieter frequency — the emergence of LLMs as a structural layer in scientific research, particularly in medicine and the life sciences.

This article does not argue that language models are ready to do science. They are not. It argues something more specific and, in the long run, more important: that LLMs are beginning to function as a new kind of research infrastructure — the kind that determines not just what scientists can do, but who gets to do serious science at all.

The Problem Is Not Ignorance. It Is Volume.

Consider the informational landscape of modern neurology. A researcher studying the early progression of Alzheimer's disease must navigate simultaneously across neuroimaging findings, genetic risk profiles, wearable-derived behavioural data, digital cognitive assessments, longitudinal biomarker studies, and a literature base that grows by thousands of papers per year. No individual — regardless of discipline or intellect — can hold all of it coherently.

This is not a problem of intelligence. It is a problem of volume. The knowledge required for frontier biomedical research has outpaced the bandwidth of individual researchers, and most institutions respond by building larger teams. Biochemists collaborate with bioinformaticians who collaborate with statisticians who collaborate with neurologists. This works — for large institutions with access to all of those people.

"The knowledge required for frontier biomedical research has outpaced the bandwidth of individual researchers. LLMs are not the solution to that problem. They are, potentially, a new instrument for navigating it."

For everyone else, the gap is structural. A small research team working on a novel hypothesis about cognitive decline does not have a bioinformatician on call. It does not have a literature analyst who can surface every relevant study from the last eight years. It runs on the same curiosity that animates all good science, constrained by a fundamental asymmetry in access to analytical capacity.

What LLMs Can Actually Do in a Scientific Context

The risk of overselling artificial intelligence in medicine is well-documented. Every decade produces its own wave of optimism — expert systems in the 1980s, neural networks in the 2000s — followed by a quieter period of recalibration. The appropriate response to this history is not scepticism about AI in general but precision about which specific capabilities are genuinely novel and which are familiar claims in new clothing.

With that caution in place: language models exhibit at least three capabilities that have no clear precedent in earlier AI tools, and all three are directly relevant to scientific research.

Three Genuinely Novel Capabilities

1. Cross-domain synthesis. LLMs can hold and operate across multiple scientific disciplines simultaneously — connecting findings in genetics, neuropsychology, and wearable sensing in ways that previously required a multidisciplinary team.

2. Literature navigation at scale. When integrated with retrieval systems, they can surface, relate, and summarise bodies of literature that would take a human researcher months to review.

3. Hypothesis articulation. They can receive partial observations and suggest formalised hypotheses — not as conclusions, but as structured starting points for experimental design.

None of these replace human judgement. A language model cannot design a clinical trial with appropriate controls, cannot evaluate the ethical dimensions of an intervention, and cannot catch the smell of something being subtly wrong in the way an experienced scientist can. What they can do is compress the time between a nascent research intuition and a structured, testable proposition.

The Democratisation Argument

The most important implication of LLMs as research infrastructure may not be what they do for established institutions. It may be what they make possible for small teams.

A large pharmaceutical company has, among its assets, hundreds of researchers, dedicated bioinformatics units, clinical data management teams, and decades of proprietary knowledge encoded in institutional memory. A two-person research group has a shared interest in a problem and, if they are fortunate, access to a dataset.

The traditional response to this asymmetry has been academic collaboration — joining forces with university departments, applying for infrastructure grants, or forming consortia. These mechanisms work, but they operate on timescales that are mismatched with the pace at which promising hypotheses need to be either pursued or abandoned.

"LLMs will not close the resource gap between a startup and a pharmaceutical company. But they may compress a specific kind of gap — the gap between having a good question and being able to think rigorously about it."

Language models do not solve the funding problem. They do not grant access to cohort data, ethics approvals, or laboratory equipment. But they can provide something that has historically required institutional scale: sustained analytical companionship across disciplines. A small team working on a neurodegenerative disease hypothesis can, with the right tooling, engage with the literature, stress-test assumptions, and iterate on experimental designs in a way that was not previously accessible outside of well-resourced environments.

Digital Biomarkers and the Coming Data Wave

Nowhere is the need for this kind of infrastructure more acute than in the emerging field of digital biomarkers.

Traditional biomarkers — blood proteins, imaging findings, genetic variants — are collected at discrete moments: a hospital visit, a scheduled assessment. Digital biomarkers, by contrast, are generated continuously: from smartphone accelerometers, sleep monitors, typing patterns, speech cadence, gait analysis, and any number of other passive sensors that individuals carry or wear in the course of ordinary life. The data they produce is not a snapshot. It is a stream.

The promise of continuous digital biomarkers is substantial. Neurodegenerative conditions, for instance, may manifest in behavioural and functional signatures years before they appear in clinical presentation. Early detection — meaningful early detection — may depend not on finding a single definitive marker but on identifying patterns in longitudinal multidimensional data that human observation alone is too slow to catch.

Illustrative Example

Consider a hypothetical research platform collecting longitudinal digital biomarkers related to cognition, behaviour, and lifestyle in individuals at risk of neurodegenerative disease. Such a platform would generate data at a density and granularity that no clinical team, reviewing case notes at a scheduled appointment, could fully analyse. The analytical infrastructure required to find meaningful signal in that data — to distinguish adaptive behavioural change from compensatory cognitive decline — is precisely the kind of infrastructure that language models, embedded in the right workflow, could help provide.

This is not science fiction. The architectural questions are real and immediate: What does it mean to build a research pipeline around continuous data? How do you maintain rigour when the variable space is vast and the ground truth is delayed by years? And crucially — how do you do this with a small team?

The Question of Model Quality and Specificity

It would be convenient to treat all large language models as interchangeable. The evidence suggests otherwise.

Models differ significantly in their ability to reason about specialised domains, maintain factual accuracy across long contexts, and navigate the specific conventions of scientific literature. A model that performs well on general reasoning benchmarks may behave quite differently when asked to analyse a differential diagnosis or trace the conceptual lineage of a methodological choice in longitudinal study design.

This matters because the selection of a language model is not merely a procurement decision. It is, in a research context, an epistemological one. A model with well-documented tendencies to confabulate, or one that lacks adequate training data in a specific medical domain, is not simply less useful — it may actively mislead.

The implication is that researchers need to evaluate models with the same methodological discipline they apply to any other analytical tool. This is not the norm today. It should become one.

National AI Ecosystems and Biomedical Research

A parallel development deserves attention, though it is less often connected to the scientific research context in which it is most relevant.

Several countries are investing significantly in developing sovereign language model capabilities — models trained on local language data, aligned to local regulatory and cultural norms, and not dependent on external infrastructure. The stated motivations are primarily economic and strategic: reducing dependence on foreign technology providers, enabling domestic industrial applications, preserving linguistic and cultural heritage in the age of machine-mediated communication.

These are legitimate motivations. But they have an underappreciated corollary.

An Overlooked Corollary

Countries investing in sovereign AI capabilities may unintentionally create a new foundation for scientific research. Locally trained language models with strong multilingual capabilities — particularly in languages historically underrepresented in global scientific publishing — could become strategic assets not only for industry but also for biomedical research conducted in those languages and cultural contexts.

Clinical research produces documentation in local languages. Patient records, case notes, qualitative interview transcripts, and observational data are often not English-language texts that global models have seen in abundance. A model with genuine competence in Slovak, Czech, Hungarian, or Finnish medical language is not simply a convenience — it is access to a body of evidence that would otherwise remain analytically inaccessible.

This suggests an opportunity for alignment between national AI infrastructure investments and the needs of biomedical research institutions that may not yet have articulated it clearly. The collaboration between research organisations and teams building language model capabilities is not just commercially interesting. It is, in a specific and meaningful sense, scientifically interesting.

The Problem of Trust and Transparency

No honest account of LLMs in scientific research can omit the problems. They are real, and in a medical context, they carry consequences.

Language models produce fluent text. Fluency is not accuracy. A model asked about the neurobiological mechanisms of a rare syndrome may produce a coherent, well-structured, and completely incorrect account — and do so with no visible hesitation. For a researcher who lacks domain expertise in that specific area, the error may not be apparent.

This is not a reason to avoid language models in research. It is a reason to treat them as instruments rather than authorities. The same logic applies to any analytical tool: statistical software can produce a precisely wrong regression coefficient; imaging analysis algorithms can produce plausible-looking artefacts. The appropriate response is methodological discipline, not abstinence.

What methodological discipline looks like in the context of LLM use in research is an open question — one that the scientific community has not yet answered systematically. Reasonable starting principles include: verifying specific factual claims independently; being explicit in publications about where AI tools were used and in what capacity; and treating LLM outputs as hypotheses to be tested rather than conclusions to be reported.

Human–AI Collaboration as a Research Modality

The framing that has dominated public discourse — AI replacing human scientists — is almost certainly wrong in both directions. It overstates what AI can do and understates what experienced human researchers provide.

The more useful frame is collaboration. Not in the aspirational, diffuse sense in which the word is often used in technology contexts, but in a specific, operational sense: a human researcher who has deep domain expertise and a language model that has broad cross-domain knowledge, working together on a structured problem, each compensating for the limitations of the other.

This kind of collaboration requires design. It does not happen by simply giving a researcher access to a chat interface. It requires workflows that support it, tools that make the interaction traceable, and norms that govern how AI contributions are distinguished from human contributions in the research record.

We do not yet have those norms. Building them — before they are urgently needed — is one of the more important tasks facing the scientific community in the immediate term.


"The future of medical AI may not belong to the largest models. It may belong to the best collaborations."

What will distinguish the research teams that use this technology well from those that do not is unlikely to be access to the most capable model. It will be the quality of the judgment surrounding it — the ability to know when to trust, when to verify, when to discard, and when to follow an unexpected output into territory that turns out to be genuinely new.

That is, in the end, a description of what good scientific thinking has always required. The instrument changes. The thinking does not.


Selected References

  1. Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature Medicine, 29, 1930–1940. doi:10.1038/s41591-023-02448-8
  2. Moor, M., Banerjee, O., Abad, Z. S. H., Krumholz, H. M., Leskovec, J., Topol, E. J., & Rajpurkar, P. (2023). Foundation models for generalist medical artificial intelligence. Nature, 616, 259–265. doi:10.1038/s41586-023-05881-4
  3. Topol, E. J. (2019). High-performance medicine: the convergence of human and artificial intelligence. Nature Medicine, 25, 44–56. doi:10.1038/s41591-018-0300-7
  4. Makin, J. G., Moses, D. A., & Chang, E. F. (2020). Machine translation of cortical activity to text with an encoder–decoder framework. Nature Neuroscience, 23, 575–582. doi:10.1038/s41593-020-0608-8
  5. Bokura, H., Kobayashi, S., & Yamaguchi, S. (2005). Electrophysiological correlates for response inhibition in patients with mild cognitive impairment. European Journal of Neuroscience, 22(6), 1525–1532. PubMed:16190898
  6. Bhatt, P., & Bhatt, D. L. (2022). Artificial intelligence in cardiology: opportunities and challenges. Nature Reviews Cardiology, 19, 738–740. doi:10.1038/s41569-022-00698-6
  7. Wainberg, M., Merico, D., Delong, A., & Frey, B. J. (2018). Deep learning in biomedicine. Nature Biotechnology, 36, 829–838. doi:10.1038/nbt.4233
  8. European Commission. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council — Artificial Intelligence Act. Official Journal of the European Union. EUR-Lex 32024R1689