AI systems now function as autonomous research collaborators. These "AI Scientists," built with large language models and integrated tools, manage research from experimental design and strategy to workflow execution, adapting based on feedback.
A "verification crisis" emerged in agentic AI for scientific discovery, marking a gap between AI-generated results and the systems' ability to provide verifiable proof.
Frameworks like Chain-of-Evidence, Audit-Closed protocols, FEV, and structural FDR enforcement were proposed to establish validation standards for trustworthy agentic science.
Manifold Agentic Reasoning was introduced as a geometric framework, extending agentic reasoning to Riemannian manifolds for dynamic, constrained environments in scientific and embodied fields.
This framework addresses four production flaws in AI agents: silent hallucinations, reasoning drift, tool misuse, and black-box evaluation, seeking solutions beyond prompt-chained templates.
In simulated evaluations, Manifold Agentic Reasoning outperformed baseline systems, achieving higher recovery success, reduced shape and pattern errors, and fewer invalid transitions on a tissue manipulation benchmark.
Manifold Agentic Reasoning was extended to graph-agentic manifold reasoning, where node states are represented on manifolds and neighbor information is aggregated via logarithmic maps and attention.
Artificial intelligence and automation are converging in enzyme engineering, leading to increasingly autonomous Design-Build-Test-Learn (DBTL) systems, progressing from semi-automated to high-autonomy workflows.
AI-guided prediction combined with automated experimentation and active learning accelerates enzyme optimization, supporting the development of improved enzymes.
Researchers proposed the "FIRE" criteria (forward-looking, individual, reasoning-intensive, experimental) to identify decisions unsuitable for AI delegation, distinguishing them from tasks AI can handle.
A study found that when GPT-4 suggests references, it systematically favors highly cited papers, reinforcing the Matthew effect, with variations across scientific fields.
Research on GPT-4's reference generation identified biases toward recent papers, shorter titles, and smaller author teams, causing its suggestions to differ from human citation patterns.
GPT-4's bibliographic suggestions showed semantic alignment with focal paper content, comparable to human-generated references, indicating accurate relation of prior work.
GPT-4 generated references reproduced local citation-network structures similar to human-curated bibliographies, suggesting its ability to capture literature organization.
GPT-4's reference suggestions reduced author self-citations, a departure from human citation practices.
GPT-4 can generate content-relevant bibliographic suggestions using only its internal knowledge, without external databases, with implications for automated research assistance.
Large language models provide high-dimensional representational spaces, offering a new scale to examine the neural basis of linguistic processing.
The model-brain alignment framework evaluates the biological plausibility of language theories, interpreting representational correspondence considering behavioral, temporal, causal, and biological constraints.
Researchers developed AutoSupervision to evaluate if scientific manuscript revisions address reviewer feedback, using peer-review records from 56,000 Nature Communications articles.
Quick Hits:
- GPT-4 amplifies dominant citation patterns.
- LLMs achieved varying performance in AutoSupervision tasks.
- A practitioner guide details graph-based workflows for generative AI systems.
Sources
- A New Paradigm: Agentic AI for Scientific Discovery
Artificial intelligence in science is undergoing a foundational change. Rather than serving as a passive analytical instrument — classifying images, predicting structures, spotting patterns — AI systems are beginning to act as autonomous research collaborators. These systems, built on large language models and tool-integrated architectures, can reason about experimental design, formulate strategies, execute multi-step workflows, and refine their approaches from empirical feedback. Often called “AI Scientists,” they participate across the full research lifecycle, from the seed of a hypothesis…
- Artificial intelligence and automation in enzyme engineering: evolution, advances, and future perspectives
Natural enzymes often fail to meet industrial demands for catalytic efficiency, stability, and substrate specificity, creating a critical bottleneck in biomanufacturing. This review examines how artificial intelligence (AI) and automation are reshaping enzyme engineering from empirical trial‑and‑error toward data-driven, closed-loop design. We trace AI development from feature-engineered machine learning to supervised deep learning and self-supervised protein language models, and automation from standalone task execution to cascade integration and biofoundry-enabled build-test workflows.…
- Manifold Agentic Reasoning: Extending Agentic POMDPs and Post-Training Reasoning to Riemannian State and Reasoning Spaces
Abstract Agentic reasoning systems increasingly interact with environments whose states are only partially observed, dynamically evolving, and constrained by physical, biological, or logical structure. Existing agentic reasoning frameworks often model internal reasoning, tool use, and post-training adaptation using flat latent representations and struggle in curved manifold space environments. However, many scientific and embodied domains naturally lie on curved state spaces, including tissue geometry, developmental trajectories, protein conformations, robotic configuration spaces, and…
- AI and actor-specific decisions
Artificial intelligence (AI) is increasingly seen as potentially replacing humans in decision-making and problem-solving across many domains. AI is effective for many well-specified decisions. But we argue that AI cannot deal with what we call “actor-specificity.” Actor-specific decisions and problems are (a) forward-looking, (b) individual and idiosyncratic, (c) reasoning-intensive, and (d) experimental—requiring intervention in the world to facilitate “counter-to-data” reasoning. These four criteria, captured by the “FIRE” acronym, function as exclusion criteria: they identify when…
- HOW DEEP DO LARGE LANGUAGE MODELS INTERNALIZE SCIENTIFIC LITERATURE AND CITATION PRACTICES?
Abstract The spread of scientific knowledge depends on how researchers discover and cite prior work. Large language models (LLMs) now add a new layer to this process, but their alignment with human citation practices across domains remains unclear. Here, we compare human citations with GPT-4ogenerated reference suggestions produced from paper metadata and abstracts. Analyzing 274, 951 generated references for 10, 000 focal papers, we find that LLMs systematically reinforce the Matthew effect by favoring highly cited papers, with field-specific variation in the rate at which generated…
- Linguistics and human brain: a perspective of computational neuroscience
Elucidating the language-brain relationship requires bridging the methodological gap between linguistics' abstract theoretical frameworks and neuroscience's empirical neural data. As an interdisciplinary cornerstone, computational neuroscience formalizes language's hierarchical and dynamic structures into testable neural representation models through modeling, simulation, and data analysis, enabling computational dialogue between linguistic hypotheses and neural mechanisms. Recent advances in deep learning, particularly large language models (LLMs), have further advanced this inquiry: their…
- AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying whether reviewer feedback leads to meaningful and evidence-supported manuscript improvements. We introduce AutoSupervision, which evaluates whether scientific manuscript revisions genuinely address reviewer concerns through grounded evidence. AutoSupervision leverages transparent peer-review records as a natural source of supervision, where reviewer comments specify…