BAIR + RDI · Web Agent Benchmarks
Berkeley BAIR + RDI
BAIR, home to Dawn Song, Pieter Abbeel, and Sergey Levine, joined forces with Dawn Song's RDI (Research Center for Decentralized Intelligence) to establish the gold standard for evaluating whether AI agents "can operate real websites like humans"—the WebArena benchmark series. It features real websites, long-horizon tasks, and programmatic verification: human completion rate is ~78%, GPT-4 in 2023 managed only 14%, and by 2025 frontier systems have caught up to 60%+. They also explore TinyAgent: distilling tool-calling capabilities into a 7B small model so agents can run off-cloud. RDI hosts the annual LLM Agents MOOC, Hackathon, and Agentic AI Summit, serving as a vital hub for the global agent research community.
WebArena / ICLR 2024TinyAgentAgentic AI Summit
Sky Computing Lab · Inference Infra
Berkeley Sky Computing
Created vLLM—the world's most widely used open-source LLM inference engine. Its core innovation, PagedAttention, applies the OS virtual memory paging concept to KV cache management, boosting GPU memory utilization from 30–60% to 90%+ and improving serving throughput by 2–4x. It now serves as the inference backbone for companies like Anyscale, Perplexity, and Alibaba Cloud. The lab, led by Ion Stoica (creator of Spark/Ray/Databricks), extends its lineage into a larger proposition called "Sky Computing"—unified compute scheduling across clouds and data centers. The companion project SkyPilot already supports multi-cloud AI task scheduling across AWS/GCP/Azure.
vLLM · 50k+ ★PagedAttentionSkyPilot
OpenHands · Software Engineering Agent
CMU OpenHands
Launched in 2024 by Professor Graham Neubig (originally OpenDevin), it places AI agents inside Docker sandboxes to execute Bash commands, browse the web, run tests, and submit PRs just like human engineers. With 31,000+ GitHub Stars and 180+ community contributors, it consistently ranks in the global top tier on the authoritative SWE-bench (2,294 bug-fix tasks from real projects like Django/Flask/NumPy). Neubig co-founded All Hands AI to productize it—a prime example of an academic lab incubating an agent company within 12 months.
31k+ ★SWE-bench #1→ All Hands AI
Simon Initiative + LearnLab · AI Tutoring
CMU Simon Initiative
Originating from Herbert Simon's 1956 cognitive information processing theory, this is the birthplace of Intelligent Tutoring Systems (ITS). With 30 years of learning behavior data and knowledge tracing models, AI tutors can diagnose "why a student is wrong" rather than simply judging "right or wrong." A 2025 randomized controlled trial conducted with Stanford found that human-AI collaborative tutoring significantly outperforms pure AI tutoring—AI excels at precise practice feedback, while humans excel at emotional connection and higher-order guidance. The new CogGen framework is attempting to use LLMs to automatically generate cognitive models, compressing ITS development cycles from months to weeks.
ITS Birthplace30 Years of DataCogGen 2025
RSL + AI Center · Embodied AI · Zurich
ETH Zurich RSL / AI Center
The Robotic Systems Lab (RSL), led by Marco Hutter, is the world's most widely cited and commercially successful academic team in legged robotics—its ANYmal quadruped robot already operates in nuclear power plants, mines, and offshore platforms, and the spin-off ANYbotics is valued at over $200M. They pioneered the Sim-to-Real transfer solution in 2017, and the Isaac Lab co-built with Nvidia has become the de facto standard for global legged robot training. The ETH AI Center (37 core professors) also features a medical AI track (Gunnar Rätsch, ICU foundation models) and serves as a core advisor on EU AI Act technical standards.
ANYmal / ANYboticsIsaac LabEU AI Act Advisor
DeepMind + Google Research · Frontier Science
Google DeepMind / Research
The heaviest signal "off the main stage" during I/O 2026 was the AI for Science matrix: DeepMind + Google Research published two papers in Nature on the same day in May—Co-Scientist and ERA—accompanied by the launch of the Gemini for Science trio, connecting 30+ life sciences databases. Co-Scientist's multi-agent "hypothesis tournament" has already provided verifiable scientific hypotheses in Stanford's liver fibrosis research and MIT+Harvard's ALS research. The infrastructure side is equally dense: DiLoCo cross-data-center training, Gemma 4 open-source in four sizes, and AlphaEvolve already deployed at enterprises like BASF and Klarna—this is a corporate lab that treats "academic output" as a core competitive advantage.
Co-Scientist / NatureAlphaEvolveDiLoCo · Gemma 4
Rajpurkar Lab · Generalist Medical AI
Harvard Rajpurkar Lab
Pranav Rajpurkar (author of CheXNet, student of Andrew Ng) at Harvard Medical School's Department of Biomedical Informatics is betting on "Generalist Medical AI"—not a patchwork of specialty diagnostic models, but a general medical brain capable of multimodal reasoning across CT imaging, clinical notes, lab data, and vital signs. The co-founded a2z Radiology AI has received FDA approval, becoming the first US CT triage AI product covering 20+ abnormal findings, representing the rare path of a "top academic lab going straight to FDA approval." They are also advancing the ClinicalBench evaluation framework in Nature Medicine, emphasizing the use of practicing physician review rather than standard NLP benchmarks to evaluate clinical AI.
a2z Radiology AI · FDAClinicalBenchMulti-Agent Consultation
Media Lab · Large Population Models
MIT AgentTorch
Professor Ramesh Raskar's team (core researcher Ayush Chopra) uses "LLM Archetypes" to bring million-scale agent social simulations to GPU-parallel scale—rather than calling an LLM for each agent individually, it clusters them by socioeconomic characteristics into several "archetypes," makes one call, and broadcasts the decision distribution in batch. It has been used for a 5-million-population digital twin in New Zealand (H5N1 avian flu, measles response) and a Minnesota COVID vaccine study, making it one of the few systems that truly combines LLM cognitive capabilities with large-scale social simulation.
LLM Archetype5M Population TwinAAMAS 2025
CSAIL · Agent Governance & Alignment
MIT CSAIL Agent
A massive research cluster of 100+ professors and 900+ graduate students has chosen two of the most overlooked yet fundamental paths in the agent era: the AI Agent Index (2025) found that 87% of agents lack safety documentation; EnCompass reduces agent search logic code by 80%; Palimpzest automates model-inference configuration selection. The Algorithmic Alignment Group's Open-Universe Assistance Games research extends the alignment problem from "closed action spaces" to the infinite action spaces that agents truly possess.
AI Agent IndexEnCompassAlignment Group
Computational Cognitive Science Lab
MIT CoCoSci
Josh Tenenbaum (2019 MacArthur Fellow) is the most empirically grounded dissenter against LLMs—his 2015 Science cover paper used probabilistic program induction to prove that children can learn a new concept from just one example, whereas deep learning requires thousands. The 2023 paper "From Word Models to World Models" argues that language is merely an entry point to understanding the world; physical, causal, and psychological intuitions cannot emerge from language statistics alone. CoCoSci is not anti-AI; its latest 2025–26 work is combining LLM-generated candidate programs with Bayesian hypothesis selection, serving as a reserve paradigm "if Scaling hits a wall."
Science Cover 2015From Word to WorldCBMM
Shaping the Future of Work Initiative
MIT Future of Work
Daron Acemoglu (2024 Nobel laureate in economics), David Autor, and Simon Johnson lead the Stone Center on Inequality, and their conclusions directly clash with Silicon Valley's optimistic narrative: employment rates for young workers aged 22–25 in AI-high-exposure occupations have relatively declined by 16%; about 80% of US workers have at least 10% of their work tasks affected by LLMs; the current automation-oriented AI path will increase capital returns and depress labor income, and the rate of "new task creation" has not yet been proven to keep pace with automation. Their core tool, the "task model of jobs," breaks down occupations into task-level analyses of technical and economic feasibility.
2024 Nobel Prize in EconomicsTask FrameworkVanishing Entry-Level
Laboratory for Financial Engineering
MIT Laboratory for Financial Engineering
Andrew Lo (who proposed the "Adaptive Markets Hypothesis" as an alternative to the Efficient Market Hypothesis) is betting that within 5 years, AI will be able to autonomously manage personal investment portfolios—not just chat-based advice, but complete democratization of wealth management that understands user risk preferences, dynamically rebalances, and identifies and intervenes in panic selling. The team also researches using LLMs to parse earnings reports and conference calls to extract sentiment factors (quantamental investing), as well as AI's role in systemic risk monitoring for financial regulation. Lo himself also serves as Chief Strategist at the AlphaSimplex hedge fund, a rare dual identity as both "scholar and real-money fund manager."
Adaptive Markets HypothesisQuantamental InvestingNeurofinance
Princeton Language and Intelligence
Princeton PLI
Founded in 2023 by Sanjeev Arora, Danqi Chen, and Karthik Narasimhan, it does not chase the agent application-layer arms race, but instead builds advantages at the foundational layer of language models—construction, alignment, understanding—that will determine the ceiling for future agents. Danqi Chen researches model knowledge editing and sources of hallucination; Narasimhan is one of the earliest researchers to combine NLP with reinforcement learning, focusing on tool use and long-horizon reasoning. In 2024, two new directions were added: AI² (AI Accelerated Invention) and NAM (Natural and Artificial Minds). Results like SimCSE and knowledge editing are widely cited, and PhD graduates consistently flow to Anthropic, OpenAI, and DeepMind.
Founded 2023SimCSETalent Pipeline
Center for Research on Foundation Models
Stanford CRFM
In 2021, Percy Liang led the publication of a 200-page report that named the "Foundation Models" paradigm—GPT-4, Claude, and Gemini can be discussed within the same narrative framework largely thanks to CRFM. Its core outputs, the HELM holistic evaluation framework (entering maintenance mode in June 2026) and the FMTI transparency index, have been cited by the EU AI Act implementation team as prototypes for regulatory language. In 2024, it launched the Marin project—completely open training of its own foundation model, with weights/data/code/logs all public, answering the question of whether "only evaluating others" has credibility.
Named Foundation ModelsHELM / FMTIMarin Open Model
Digital Economy Lab + SALT Lab
Stanford Dual Labs · Future of Work
Erik Brynjolfsson's Digital Economy Lab uses real enterprise data to measure AI's impact: after customer service teams adopted AI, productivity increased by 14%, with novices benefiting over twice as much as experts; Diyi Yang's SALT Lab conducted a four-quadrant survey of 1,500 workers + 52 AI experts, delineating green/red zones for "AI can do × workers are willing" and proposing the Human Agency Scale (H1–H5) to annotate the degree of necessary human intervention for each task. The two labs complement each other to form the most systematic academic answer to "how AI changes work."
Productivity +14%Four-Quadrant FrameworkHuman Agency Scale
Open Virtual Assistant Lab · STORM
Stanford OVAL
Professor Monica Lam's team built the deep research agent STORM: given a topic, it automatically simulates different expert perspectives asking each other questions, searching, and synthesizing, ultimately outputting a Wikipedia-level report with citations, entirely without human intervention. Since launch, approximately 1 million users have used it. The upgraded Co-STORM (EMNLP 2024) introduces human-AI collaborative interruption/confirmation mechanisms. Compared to similar deep research products from OpenAI/Google, STORM is fully open-source, reproducible, and modular—currently extending toward discovering implicit connections in biomedical literature.
13k+ ★Co-STORM1M Users
Regulation, Evaluation & Governance Lab
Stanford RegLab
Led by Daniel Ho, it simultaneously answers two questions—"how AI can help governments function" and "how society can govern AI"—a rare bidirectional research approach. The STARA tool helped San Francisco city attorneys clean up thousands of pages of outdated regulations; in collaboration with Princeton and Santa Clara County, AI was used to annotate racially restrictive covenants from 5 million property deeds within weeks (a task that would take humans decades); in partnership with the City of San Diego, computer vision was used to optimize road maintenance resource allocation, preventing the systematic neglect of low-income communities. Starting in 2025, in conjunction with Stanford HAI, it runs AI boot camps for government officials at all levels.
STARARacial Covenant DetectionAI Boot Camp
Scaling Intelligence Lab
Stanford Scaling Intelligence
Azalia Mirhoseini used reinforcement learning at Google Brain to train AlphaChip, which generated chip floorplans in 6 hours that comprehensively outperformed human engineers; it has been used in Google's TPU production design (in 2024, Nature published a correction regarding experimental conditions, but the controversy itself proves this field has entered serious scrutiny). The lab she founded in 2023 pursues a larger question: recursive self-improvement—AI designs better chips, chips train stronger AI, can this closed loop accelerate? She is also an early contributor to the Mixture-of-Experts (MoE) architecture, and in 2024 co-founded Ricursive Intelligence to continue this line of research.
AlphaChip / TPUMoE Early ContributorRicursive Intelligence
Allen School + Allen Institute for AI
UW Ai2 · OLMo
The University of Washington's Allen School and the nonprofit Ai2 (founded by Paul Allen in 2014) form the tightest academic-nonprofit AI collaboration in the US—joint faculty appointments, and the OLMo project is co-led by individuals holding both professorships and senior director roles at Ai2. OLMo is the world's most "truly open-source" LLM series: it opens not only the weights but also the training data, code, and logs, a fundamental distinction from Llama's weight-only openness; the companion OLMoTrace can trace model outputs back to specific training documents. The OMAI launched in August 2025 ($152M, NSF + NVIDIA) is the largest single open AI infrastructure investment by the US academic community.
OLMo True Open SourceOLMoTraceOMAI $152M