In August 2024, Sakana AI shocked academia with a paper generation cost of $15/paper; in March 2026, this paper was officially published in Nature, and concurrently AI Scientist-v2 was upgraded to "Workshop-level automated scientific discovery." Full-pipeline AI-driven research—from topic selection to peer review—is moving from proof of concept to engineering reality—and all this comes from a company founded by the original authors of the Transformer paper.
Sakana AI's founding team includes multiple original authors of the Transformer architecture—the 2017 paper ("Attention Is All You Need") laid the technical foundation for all large language models today. Their next bet: if the Transformer enabled AI to understand and generate language, then the next step should be enabling AI to understand and generate scientific knowledge itself.
The core idea of the AI Scientist project is strikingly direct: build an end-to-end research agent pipeline, covering the complete closed loop from topic selection → literature review → experimental design → code implementation → result analysis → paper writing → peer review.
The first version released in August 2024 (arXiv: 2408.06292) caused a sensation in the AI community, with the core data point being: the generation cost per paper is approximately $15, including the full pipeline of topic selection, experiments, writing, and automated peer review. While the quality of the generated papers did not match top-tier human research, it reached a level submittable to mid-tier conferences.
If the Transformer enabled AI to understand and generate language, then the next step should be enabling AI to understand and generatescientific knowledge itself.
The core bet of Sakana AI's founding teamIn March 2026, AI Scientist reached two milestone events.
First, the original paper was officially published in Nature (sakana.ai/ai-scientist-nature). This is not only a recognition of the technology itself, but also marks "AI-automated research" entering the mainstream scientific discourse from a fringe topic. Nature's editors and reviewers recognized the academic value of this direction—even while reserving judgment on its current capabilities.
Second, AI Scientist-v2 was released (HuggingFace trending paper), upgraded to "Workshop-level automated scientific discovery." The core improvement in v2 is the introduction of Agentic Tree Search—instead of linearly executing the "topic selection → experiment → writing" pipeline as in v1, it expands multiple parallel exploration paths at each stage and selects the optimal branch through an evaluation function. This is essentially transferring AlphaGo's Monte Carlo tree search concepts to the research process.
Concurrently, research in related directions is also emerging rapidly:
EvoScientist: A multi-agent evolutionary AI scientist that evolves stronger research capabilities through collaboration and competition among multiple agents. Towards a Medical AI Scientist: Extending the AI Scientist framework to medical research, targeting higher-risk, higher-value application scenarios. Deep Research of Deep Research: A systematic review from Transformer to Agent, from AI to AI for Science.
The most disruptive number for AI Scientist isn't paper quality, but thecost structure.
The production function of traditional academic research is: 1 PhD student × 4-6 years × $30,000-$50,000/year salary = several papers. Adding advisor time, computing resources, and lab equipment, the full cost of one high-quality paper can reach $100,000-$500,000. AI Scientist's production function is: 1 API call × $15 = 1 mid-quality paper.
Even though v2's "Workshop-level" still falls short of top-conference papers, the cost difference is 4-5 orders of magnitude. This means:
The breadth of experimental exploration will be fundamentally transformed. Today a research team might only deeply explore 3-5 directions; AI Scientist can simultaneously explore 3,000-5,000. The value of failed experiments will be rediscovered. In human research, failed experiments are often discarded; AI Scientist can systematically record and analyze all failure paths, building a knowledge graph of "what doesn't work." Democratization of research will truly happen. Resource-constrained institutions (universities in developing countries, small corporate R&D departments) will gain exploration capabilities previously available only to MIT/Stanford/Google DeepMind.
The greatest leverage of research automation lies not in paper quality, but inexploration breadth anddemocratization—when the marginal cost of exploring a direction drops from $100k to $15, the scarcest resource in research is no longer funding, butthe ability to ask the right questions.
In January 2026, Google invested in Sakana AI, with the collaboration focusing on two directions.
AI-Scientist: The core research automation product, leveraging Google's computing infrastructure to scale up experiments. ALE-Agent: Sakana's other core project, aimed at "achieving reliable AI deployment and adoption in foundational industries"—particularly financial institutions and government departments with extremely high requirements for security and data control.
Google's investment logic is clear: if AI Scientist can truly accelerate research output, then the company with the largest computing resources and most complete data infrastructure will gain the greatest leverage. Research automation is not about replacing Google DeepMind's researchers, but about amplifying each researcher's output 10x.
First, the quality ceiling. Currently, papers generated by AI Scientist still fall far below top-tier human research in originality and depth. Workshop-level ≠ top-conference-level, let alone Nature/Science-level. The core bottleneck: AI excels at combinatorial innovation within known frameworks, but struggles to propose entirely new research paradigms.
Second, review credibility. AI Scientist uses LLMs for automated peer review, but LLMs' reviewing capability is itself an unsolved problem. When AI is both the paper author and the reviewer, how can review independence be guaranteed?
Third, research ethics. If a lab uses AI Scientist to mass-generate papers, the impact on the academic evaluation system (which centers on publication counts) would be devastating. Paper volume inflation would render existing academic assessment mechanisms ineffective.
Fourth, experimental validation. AI Scientist currently performs well mainly in computational experiment domains (ML experiments, numerical simulations), but physical experiments, biological experiments, and other domains requiring real-world interaction remain blind spots for automation. The Medical AI Scientist attempt is precisely challenging this boundary.
A judgment framework for AI practitioners and corporate managers:
Short-term (6-12 months): The first scenario where AI Scientist lands will be the preliminary research phase of corporate R&D departments—using AI to batch-explore the feasibility of technical directions and quickly screen out directions worth deeper human investigation. This isn't replacing R&D teams, but providing them with a "navigation system."
Medium-term (12-24 months): As v2's Agentic Tree Search capabilities mature, AI Scientist will have substantive impact in AI4S fields like drug discovery, materials science, and climate modeling—domains where the search space is enormous but the evaluation function is relatively well-defined.
Long-term (>2 years): The production function of research will be permanently rewritten. The Transformer authors founding Sakana to research "AI doing research" is itself the best metaphor for this trend: the next breakthrough technology will likely be discovered by AI itself.
Researchers who treat AI Scientist as a tool will be 5-10x faster than those who "oppose AI writing papers." Nature has already accepted an AI-autonomous paper, meaning academic peer review standards have already changed—this fact is irreversible.
AI Scientist's methodology can be directly transferred to industrial R&D—applying the same agent framework to "hypothesis generation / experiment execution / data analysis / report writing." Karpathy AutoResearch + AI Scientist v2 = the new infrastructure for industrial R&D.
Research automation is one of the most underrated AI application tracks. Paper output rate × 100, research barrier reduction × 1,000—meaning the rate of new scientific discoveries could see a qualitative leap in 2027-2028.
It took only 18 months for research automation to go from concept to Nature publication. AI Scientist v2's "Workshop-level" means not just single-paper output, butentire research directions can be taken over by AI. The irony: this comes from a company founded by the original authors of the Transformer paper—the people who invented the Transformer are now having it write papers. This alone is a marker of the era.