Artificial Intelligence

AI for Science 2026: The Global Race for Scientific Discovery

2026 AI for Science landscape report: WAIC 2026 exhibitor breakdown, MIT Boltz-2 and CMU Coscientist case studies, US/China/EU policy comparison, Anthropic Claude Science, and US academic cautionary views.

Share
AI for Science 2026: The Global Race for Scientific Discovery
Share

How China, the United States, and Europe Are Building the Next Generation of Scientific Discovery

Executive Summary

AI for Science (AI4S) entered an accelerated phase between 2024 and 2026. Quantifiable efficiency gains have been documented in protein structure prediction, materials screening, climate modeling, and molecular drug design. Major economies have introduced dedicated funding programs. Work from leading US research universities on autonomous laboratories, protein generative models, and research integrity issues provides an important reference for understanding the technical boundaries of the field.

1.  WAIC 2026: AI4S Exhibition Zone and Participating Companies

The 9th World Artificial Intelligence Conference (WAIC 2026) was held July 17–20 in Shanghai under the theme "Intelligent Partners, Co-creating the Future." Exhibition area exceeded 100,000 sqm with 1,100+ exhibitors, 4,000+ exhibits, and 300+ global product debuts. AI for Science debuted as a dedicated exhibition zone and forum track. Nine Turing Award and Nobel Prize laureates attended.

1.1  Key AI for Science Exhibitors

▶  XtalPi 

Demonstrated a full-process AI autonomous discovery system: an integrated hardware-software platform covering sample preparation through data analysis. Live demo: unmanned drug crystal-form screening completed without human intervention. The company has established partnerships with Pfizer and Eli Lilly.

▶  Tianwu Technology · Matwings Venus™ Protein Research Platform

Protein research platform supporting design–synthesis–detection wet/dry lab loops via natural language input, converting researcher intent into executable automated workflows and reducing manual intervention steps.

▶  DeepModeling · Mira Platform

Based on the DeePMD molecular dynamics framework (10,000+ citations in Nature, Science, and peers), the Mira platform has verified 40+ experimental predictions and reduced materials development cycles from years to months. Commercial access available for new-energy materials and advanced manufacturing.

▶  Shanghai AI Lab · Intern-S1-Pro

1-trillion-parameter open-source MoE scientific LLM released February 2026, covering mathematical reasoning, physics simulation, life science, and earth science. InternDiscovery integrates 200+ agents and 200+ PB datasets. Two cancer target candidates (GPR160, ARG2) identified in collaboration with Lingang Laboratory are currently undergoing experimental validation.

▶  Main Forum Live Demonstrations

•  2,024 defect-free qubit arrangements completed in 60ms

•  Cancer detection executed at single-cell resolution

•  Global climate change analysis at minute-scale temporal resolution

•  Space-debris trajectory prediction and avoidance warning

1.2  Recap: WAIC 2025

WAIC 2025 (July 26–28, 2025) marked the first large-scale collective debut of AI4S in China. Shanghai AI Lab released Intern-S1 (the first open-source scientific multimodal LLM) and InternDiscovery; the "AGI×Science Frontier" Forum was held; ten joint AI+Science research outcomes were announced. Huawei, SenseTime, and Baidu concurrently demonstrated vertical scientific AI applications.

2.  Global AI4S Research Landscape: Progress and Debates

2.1  Key Milestones (2021–2026)

YearEvent
2021DeepMind AlphaFold 2 released: 200M+ protein structures predicted; awarded 2024 Nobel Prize in Chemistry
2023DeepMind GNoME: 2.2M new crystals predicted, 380,000 structurally stable; 736 externally synthesized and verified
2023CMU Coscientist published in Nature: AI system autonomously executes Nobel-level chemistry (Suzuki coupling)
2024AlphaFold 3: protein–DNA/RNA/small-molecule interaction prediction; 50%+ accuracy improvement over prior generation
2024Isomorphic Labs signs collaboration agreements with Eli Lilly and Novartis for AI-assisted drug design
2025.02Shanghai AI Lab open-sources Intern-S1-Pro: 1-trillion-parameter MoE scientific LLM
2025.06MIT CSAIL + Recursion release Boltz-2: protein structure + binding affinity on a single GPU in ~20s; ~1,000× faster than FEP methods
2025.10EU releases European AI in Science Strategy; RAISE virtual institute launched; ~€700M in Horizon Europe 2025 WP
2026.02Genesis Molecular AI releases Pearl, claiming 40% improvement over AlphaFold 3 on internal drug-discovery benchmarks (note: benchmark dataset and testing protocol not publicly disclosed; independent results from CASP have not confirmed this figure)
2026.06Anthropic launches Claude Science workbench: 60+ scientific databases; pharma beta access opened
2026.07WAIC 2026: first dedicated AI4S exhibition zone; DeepModeling Mira, XtalPi, Tianwu showcase autonomous lab results

2.2  Current Consensus

•  Efficiency gains are documented: speed improvements in protein prediction (Boltz-2, ~20s/GPU), materials screening (DeePMD Mira, years-to-months cycle compression), and climate simulation are supported by experimental or deployment data.

•  Cross-disciplinary deployment is a clear trend: LLMs and foundation models are being deployed as general research tools across physics, chemistry, biology, and earth science.

•  Foundation models are transitioning toward scientific infrastructure: AlphaFold, Intern-S1-Pro, and peers are downloaded and used by multiple independent institutions, fulfilling a role analogous to legacy scientific software packages.

2.3  Key Disagreements

Interpretability:   AI reasoning in physics and chemistry is difficult to audit. The gap between model output and verifiable mechanism is especially pronounced in fundamental science. The EU AI in Science Strategy lists interpretability as a technical constraint.

Reproducibility:   Fewer than 30% of AI science papers provide accessible code. In May 2025, MIT disavowed a doctoral student paper on AI productivity benefits, citing concerns over data provenance and validity (per MIT official statement and multiple academic news outlets, May 2025).

Boundaries of autonomous discovery:   Evaluations of "AI scientist" systems have recorded instances where pipelines generate synthetic placeholder data upon execution failure rather than halting; hallucinated citations and inflated result claims have also been documented, with experiment failure rates as high as 42% in some evaluations.

Data sovereignty:   Scientific data from developing nations is used to train commercial AI systems without adequate governance frameworks. The International Science Council and others have formally raised this concern.

3.  Key Participants

3.1  International Technology Companies

•  Google DeepMind: AlphaFold 3, GNoME, Gemini for Science; Isomorphic Labs focused on AI-assisted drug design with partnerships at major pharma firms

•  Microsoft: Azure AI for Health data platform; ~$80B global data center investment in 2025

•  Meta AI: ESM protein language model series (ESM-2, ESM-3), released as open-source, widely used in academic research

•  Nvidia: CUDA ecosystem as the primary compute foundation for AI scientific computing; BioNeMo platform covering genomics, protein, and chemical molecular modeling

•  Anthropic: Claude Science workbench (launched June 30, 2026); 60+ scientific databases; multi-agent architecture; beta access for pharma and life sciences

3.2  Chinese Technology Companies and Research Institutions

•  Shanghai AI Lab: Intern-S1-Pro (1T-param, open-source), InternDiscovery platform; serves as China's national-level AI4S infrastructure

•  DeepModeling: DeePMD molecular dynamics framework (10,000+ academic citations); Mira platform with 40+ verified experimental predictions, entering commercial stage

•  XtalPi: AI-driven drug crystal-form screening; partnerships with Pfizer, Eli Lilly; full-process autonomous discovery system demonstrated at WAIC 2026

•  Tianwu Technology: Matwings Venus™ protein research platform, natural language-driven wet/dry lab loop

•  BioMap: life sciences foundation model covering genomics, protein structure, and molecular design

•  Huawei Pangu Weather Model: climate prediction model; research results published in Nature (2023)

4.  National AI4S Policy Overview

Country/RegionKey InitiativesFunding and Key Details
🇺🇸 USNSF National AI Research Institute network; NIH, DOE targeted programs; Genesis Science Initiative (Nov 2025)NSF: $700M+/yr in AI research (incl. NIH $309M, DOE $187M); $100M added in 2025 for 5 new AI Research Institutes
🇨🇳 China"AI+" Action Plan (State Council, 2025); Beijing AI-Empowered Research 3-Year Plan (2025–2027)Target: >70% AI integration across six sectors by 2027; Shanghai AI Lab serves as national AI4S infrastructure platform
🇪🇺 EUEuropean AI in Science Strategy (Oct 2025); RAISE virtual institute pilotHorizon Europe 2025 WP: ~€1.6B (AI in science: ~€700M); mid-term target: ~€3B; RAISE pilot: €107M approved
🇬🇧 UKUKRI targeted funding for life science, climate modeling, quantum scienceMaintains Horizon Europe research collaboration; AI4S designated as UKRI strategic priority
🇸🇬 SingaporeNRF AI4S Initiative (S$120M); 4 flagship NUS projectsAI4Science International Conference (Jul 2024); AI4X 2025 at NUS; Horizon Europe Complementary Fund (Dec 2025)

5.  Institution Spotlights

5.1  Anthropic · Claude Science Workbench

Anthropic launched Claude Science (Beta) via an online event on June 30, 2026. It is positioned as an AI-native research environment built on existing Claude models, not a new model. Attendees included Anthropic board member and Novartis CEO Vas Narasimhan, Bristol Myers Squibb CEO Chris Boerner, and Genentech Head of Research Aviv Regev.

Key features:   60+ scientific databases; domain-specific toolkits for genomics, protein structure, and chemistry; multi-agent architecture (coordinator agent + specialist sub-agents); auditable research artifact generation.

Access:   Beta open to Pro/Max/Team/Enterprise subscribers. Project grant program: 50 slots, up to $30K compute credits each.

5.2  Singapore AI for Science Initiative

Singapore's National Research Foundation (NRF) leads the S$120M AI4S Initiative, supporting four flagship NUS projects in coordination with the Nobel Turing Challenge Initiative. Milestones: AI4Science International Conference (July 2024); AI4X 2025 conference at NUS; Horizon Europe Complementary Fund signed (December 2025). Website migrating to NRF official platform.

6.  RobotToday Analysis: Three Observations and Structural Analysis

Observation 1: Automation levels are rising, but "automated execution" and "autonomous discovery" remain distinct

Live demonstrations at WAIC 2026 from XtalPi, DeepModeling, and Tianwu carried specific technical details. However, "autonomous discovery" (independent hypothesis generation) and "automated execution" (workflow automation) are distinct concepts. Reporting should maintain this distinction.

Observation 2: US and China have developed largely independent AI4S ecosystem paths

The US (DeepMind, Anthropic, Meta) and China (Shanghai AI Lab, DeepModeling, XtalPi) have each built relatively self-contained AI4S ecosystems with low cross-border technical dependency. This will have medium-term implications for international data-sharing agreements and academic collaboration frameworks. The US pathway is primarily driven by private capital and big-tech monetization (Anthropic subscription workbench; Isomorphic Labs commercial licensing); China's pathway is built on a state-backed consortium model (Shanghai AI Lab as public infrastructure; DeepModeling, XtalPi commercializing on top). A key caveat: both pathways remain highly dependent on the same global pool of scientific literature for model pre-training—meaning a continued decline in code reproducibility rates would affect both ecosystems simultaneously.

Observation 3: Pharma and life sciences are the most active domain for commercial validation

Isomorphic Labs, Claude Science, BioMap, and XtalPi have all concentrated commercial deployment in the upstream pharmaceutical pipeline. The first disclosures of AI-tool-discovered drug candidates entering clinical stages may occur in 2026–2027.

Risk Flag 1: Reproducibility challenges may intensify under AI-driven publication cycles

Fewer than 30% of AI science papers provide accessible code. MIT's 2025 retraction of an AI productivity paper illustrates that data integrity challenges extend to AI4S research itself. The tension between fast-publication culture and academic validation is particularly acute in this field.

Risk Flag 2: Data governance gaps may affect the sustainability of global research collaboration

Scientific data from developing nations is being used to train commercial AI systems without adequate governance frameworks. The International Science Council has formally raised this concern. Unresolved, it may reduce participation in cross-regional scientific data sharing.

7.  Leading US Research Universities: AI4S Case Studies

Work from US research universities concentrates in three directions: engineering of protein and molecular generative models, construction of physical autonomous laboratories, and policy engagement. All cases draw from published papers or institutional public materials.

Massachusetts Institute of Technology (MIT) — Protein Binding Prediction and Drug Discovery Tools

Boltz-1 / Boltz-2  MIT Jameel Clinic + CSAIL × Recursion  |  2024–2025

MIT Jameel Clinic and CSAIL, in collaboration with Recursion, released the Boltz series. Boltz-1 (late 2024): open-source protein structure prediction benchmarked against AlphaFold 2. Boltz-2 (June 2025): adds binding affinity prediction to the same framework, outputting structure and affinity on a single GPU in approximately 20 seconds—approximately 1,000× faster than free-energy perturbation (FEP) methods (note: the speedup compares single-GPU inference time against full multi-CPU FEP simulation time; the two methods differ in computational scope and downstream validation requirements), with binding affinity performance approximately double that of prior methods on benchmarks. Both models are open-source.

BoltzGen  MIT Jameel Clinic  | November 2025

A generative protein design model built on the Boltz framework. Unlike structure prediction, BoltzGen generates novel protein sequences designed to bind specified biomolecular targets, producing candidates ready for downstream experimental validation in the drug discovery pipeline.

AI-Assisted Crohn's Disease Antibiotic Mechanism Analysis  MIT CSAIL  | 2024–2025

CSAIL researchers used AI models to analyze the selective mechanism of a narrow-spectrum antibiotic on the gut microbiome of Crohn's disease patients, identifying molecular pathways that target harmful bacteria without affecting beneficial flora. The findings provide a computational basis for personalized microbiome-targeted interventions.

Carnegie Mellon University (CMU) — Autonomous Laboratory Infrastructure

Coscientist  CMU Department of Chemistry + ML  | Gomes et al., Nature, December 2023 (doi:10.1038/s41586-023-06792-0)

Coscientist is the first peer-reviewed AI system demonstrated to autonomously plan, design, and execute chemical experiments. Developed by the team of Assistant Professor Gabe Gomes, it uses GPT-4 and Claude to process natural language experimental instructions and interfaces with robotic equipment for physical execution. Without pre-programming, the system completed a Suzuki coupling reaction—whose human inventors received the 2010 Nobel Prize in Chemistry—in minutes, with results consistent with manually executed reference experiments.

AI Science Foundry  Carnegie Mellon University  | Ongoing

The AI Science Foundry is a cluster of physical autonomous laboratories integrating AI, automation, robotics, and computational workflows across chemistry, biology, and hard materials, supported by large-scale storage and cloud-programmable laboratory facilities. Researchers submit experiment designs via remote interfaces; the system executes, acquires data, and returns analysis. CMU positions it as the largest US university co-located autonomous laboratory cluster, with the goal of reducing innovation-to-application timelines from years to months.

Stanford University — Genomic Foundation Models and Targeted Materials Discovery

Evo 2 (DNA Language Model) Arc Institute + Stanford + UC Berkeley |  February 2025

Evo 2 is a DNA language model trained on 9 trillion base pairs of genomic sequence covering nearly all catalogued life forms. 40 billion parameters. Stanford researcher Brian Hie was among the core contributors. The model supports mutation effect prediction and novel antibody sequence generation, released as open-source.

AI-Driven Targeted Materials Discovery  SLAC National Accelerator Laboratory + Stanford  |  2024

Researchers developed an AI method using active learning to reduce the number of experiments required in materials discovery, improving search efficiency in complex design spaces relevant to quantum computing, climate applications, and drug design. Integrated with self-driving experiment frameworks supporting automated parameter adjustment for robotic setups; currently in laboratory validation stage.

UC Berkeley — Scientific Foundation Models Research

BAIR + DL4SCI 2026  UC Berkeley + Lawrence Berkeley National Laboratory

Berkeley AI Research (BAIR) comprises 50+ faculty and 300+ doctoral students across ML, computer vision, NLP, and robotics. DL4SCI 2026 (Deep Learning for Science), co-organized with Lawrence Berkeley National Laboratory, is a five-day intensive program focused on training, adaptation, and evaluation of scientific foundation models, and on reasoning-centric research workflows. Open to doctoral students and researchers in scientific domains.

8.  US Academic Perspectives and Cautionary Views

Leading US research universities broadly acknowledge the efficiency value of AI4S, but hold explicit reservations on technical reliability, research integrity, and researcher training. The following is based on published papers, institutional statements, and public reporting.

8.1  Institutional Stances

MIT CSAIL — Active participation with policy conditions

MIT CSAIL submitted AI Action Plan recommendations to the US government in 2025, supporting AI for Science as a national priority. Specific recommendations: continued investment in basic research (especially new algorithms and ML architectures); maintaining a healthy mix of open-source and proprietary licensing; strengthening high-school mathematics and CS education. CSAIL also explicitly positioned interpretability, alignment, and safety as deployment prerequisites, not post-hoc review items.

Stanford HAI — Emphasis on human-AI collaboration

Stanford HAI, in its AI Index 2025 Report and public commentary, has explicitly positioned AI-accelerated scientific discovery as conditional on human-centeredness: AI provides computational support while researchers retain control over hypothesis generation and experimental interpretation. HAI frames AI4S as an upgrade to research tooling, not a replacement restructuring of the scientific paradigm.

CMU — System builder, attentive to scaling challenges

CMU's investments in Coscientist and the AI Science Foundry reflect a "build-first" approach. Its public statements describe Coscientist's demonstrated automation capabilities as a structural change in research execution, while noting that the next challenge is extending such capabilities to more complex scientific problems rather than reactions well-covered by existing literature.

8.2  Technical Reservations

"LLMs Are Not Scientists Yet" (2026):   Multiple academic teams, including scholars affiliated with MIT and Stanford, published reviews summarizing lessons from four autonomous AI research attempts. Core conclusion: current LLMs lack the capacity to generate genuinely novel scientific hypotheses; their strength lies in optimization and retrieval within known knowledge frameworks, not in exploring the space of unknown unknowns.

Autonomous pipeline execution failures:   Systematic evaluations of "AI scientist" pipelines have documented instances where systems generate synthetic placeholder data upon execution failure and continue running rather than halting. Hallucinated citations and inflated performance claims in unreviewed outputs have also been recorded; experiment failure rates as high as 42% were observed in some evaluations.

MIT data integrity incident (May 2025):   MIT disavowed a doctoral student paper on AI's productivity benefits in scientific research, citing "no confidence in the provenance, reliability, or validity of the data." The paper had been widely cited. The incident highlights that AI4S-related research is itself subject to data validation challenges, with elevated risk in high-expectations environments.

8.3  Structural Risk: Researcher Deskilling

Several US academics have raised the concern that heavy reliance on AI for automated experiment execution and analysis generation may reduce training opportunities for early-career researchers in core competencies: experimental design, data judgment, and anomaly identification. Referred to as "researcher deskilling," systematic data remain limited, but the concern has been incorporated into multiple academic policy recommendation documents as a risk warranting active monitoring.

Research data sourced from public information as of July 19, 2026: WAIC official releases, Anthropic.com, EU Commission policy, Singapore NRF, Qbitai, 36Kr, TMTPost, and other major media.

RobotToday Initiative

Robotics needs a service framework.

RSF defines a common language for robot service capability, lifecycle operations, certification pathways, and service-provider networks.

Share
Written by
Thomas Siew - Associtae Editor

Thomas Siew is an Editor specializing in manufacturing and supply chain analysis. He brings a global perspective and a sharp sensitivity to international business developments, examining how shifts across borders impact industry dynamics.