Skip to content
← DeepDive Labs & Builders · v1
← DeepDiveDD · 0101 · 2026-07-30
University AI Labs · Field Report
Quick Directory · 19 Labs →
SPECIAL ISSUE
UNIVERSITY FIELD REPORT · SPECIAL ISSUE

Top US University
AI Labs Panorama

14 labs, 6 universities, covering AI agent reasoning, AI infrastructure, healthcare, finance, education, labor economics, and beyond. These researchers don't chase Leaderboards — they define what AI truly can and cannot do.

Labs Covered
14
Labs
Universities
6
MIT · Stanford · Berkeley · CMU · Princeton · Harvard
Research Areas
6
Vertical Domains
Years Span
2021
→ 2025
Labs by Research Category COUNT
Agent & Reasoning 6
AI & Society 4
AI Infrastructure 2
Domain AI 2
§ 01 / Overview

The Real Researchers
Who Don't Chase Leaderboards

2023–2025 was a qualitative inflection point for AI research at top US universities. While OpenAI and Google competed on parameter scale, university labs did what companies wouldn't: measure what AI can actually do, what it cannot do — and what it is already doing to human society.

These 14 labs span six universities and six research directions: Can AI agents independently complete real tasks? How deep is LLM's impact on the labor market? How has inference efficiency become a new capability race? Can AI truly make medical diagnoses? Can chip design be taken over by AI? When will educational AI surpass real teachers?

None of the answers to these questions can be read from parameter counts or MMLU scores. They require RCT experimental design, neuroscience methods, historical archive analysis, or even line-by-line scanning of 5 million property deeds. This is the fundamental methodological divide between university AI research and industry AI research.

87% of AI agents have no safety documentation whatsoever. This isn't a product defect — it's a systematic omission.

§ 02 / Agents

What AI Agents
Can and Cannot Do

A01 2024 · AAMAS 2025
MIT Media Lab
Simulation LLM Agents

MIT AgentTorch: Simulating Entire Societies with Millions of AI Agents

A team led by Ayush Chopra and Ramesh Raskar developed AgentTorch, a GPU-parallelized million-agent simulation framework. The core innovation is LLM Archetypes — replacing traditional ABM behavioral rule sets with language models, allowing each digital person to make personalized decisions based on census profiles rather than following homogenized mathematical equations.

In 2024, the team built a digital twin of New Zealand's 5 million citizens, reproducing the effects of various COVID-19 intervention measures. This is currently the world's largest LLM-driven social simulation, and was accepted as an Oral paper at AAMAS 2025.

Key Finding

LLM-driven million-scale simulation outperforms traditional ABM in infectious disease modeling and can answer counterfactual questions like "what if a different policy had been used" — something traditional models cannot do.

A02 2024 · NAACL / EMNLP
Stanford OVAL
Research AI Multi-Agent

Stanford OVAL STORM: Letting AI Autonomously Write Complete Research Reports

Monica Lam's Open Virtual Assistant Lab (OVAL) developed the STORM system: AI generates Wikipedia-grade in-depth research reports through multi-perspective simulated conversations. The core innovation is a "pre-writing phase" — having AI first simulate multiple experts with different viewpoints asking each other questions, then building a research outline based on these questions, and finally generating a substantive long-form article.

STORM rapidly gained mainstream adoption after release: 13,000+ GitHub Stars, a million users, and spawned Co-STORM (a human-in-the-loop collaborative version, EMNLP 2024). This is a rare case of large-scale real-world deployment from an academic NLP lab.

Key Finding

Compared to having AI write reports directly, having AI first "simulate questioning" before writing significantly improves content coverage breadth and logical depth — validating that AI's "thinking warm-up" is more effective than direct generation.

A03 2024 · ICLR 2024
Berkeley BAIR + RDI
Benchmark Web Agent

Berkeley WebArena: Teaching AI Agents to Actually Use Browsers

BAIR (Pieter Abbeel, Dawn Song, Sergey Levine, et al.) and RDI jointly developed WebArena — a real-deployed open-source website environment (shopping, forums, code repositories, maps) where AI agents complete tasks on real web pages rather than synthetic environments. Published at ICLR 2024, it became the de facto benchmark for Web Agent research.

Key data comparison: Human completion rate on WebArena is approximately 78%. The best AI agents in 2024 achieved around 35-40%, a still-massive gap. This figure has circulated widely in the industry, becoming the core evidence in discussions about "AI agents being overhyped."

Task Success Rate — WebArena COMPLETION %
Human 78%
Best Agent '25 ~45%
Best Agent '24 ~36%
A04 2024 · SWE-bench #1
CMU NLP Group
Code Agent 31k Stars

CMU OpenHands: Making AI a True Software Engineer

Graham Neubig's NLP Group developed OpenHands (formerly OpenDevin), the AI agent platform closest to "truly writing code independently." The core design is a sandboxed Docker runtime — the agent can autonomously write, run, and debug code in an isolated environment, handling real issues from GitHub.

OpenHands has consistently ranked #1 globally on SWE-bench Full (a real GitHub issue-fixing benchmark), spawned the startup All Hands AI, and accumulated 31,000+ GitHub Stars. SWE-bench Multimodal was published at ICLR 2025, extending evaluation to multimodal software tasks.

Key Finding

SWE-bench proves there is a massive gulf between "completing real tasks on real codebases" and "passing coding interviews." The best Code Agents can resolve ~50% of real GitHub issues, but failure modes are heavily concentrated on understanding cross-file context and long-range dependencies.

A05 2025 · Safety Research
MIT CSAIL
Alignment Audit

MIT CSAIL Algorithmic Alignment: Mapping the Global Landscape and Alignment Boundaries of AI Agents

The MIT CSAIL Algorithmic Alignment group released the AI Agent Index (2025), systematically surveying 200+ deployed AI agents and revealing industry-wide systematic deficiencies in safety documentation, permission transparency, and user control. They also released EnCompass (a framework with 80% less code than baselines) and Palimpzest, a data science workflow acceleration tool.

Core Data (AI Agent Index)

87% of deployed AI agents have no safety documentation; 76% lack clear permission scope descriptions; only 8% provide interfaces for users to control agent behavior.

A06 2023–2025 · Full Stack
Princeton PLI
LLM Foundations Alignment

Princeton Language & Intelligence: Planting the Seeds for Agents at the Foundation Layer

PLI, co-led by Danqi Chen (pioneer in RLHF and in-context learning), Karthik Narasimhan (pioneer in language-grounded agents), and Sanjeev Arora (theoretical ML), is one of the highest-density institutions for LLM foundational research in the US. Research covers the full LLM lifecycle: construction (pre-training/post-training), alignment, and interpretation.

AI² (Algorithmic Interpretability Initiative) and NAM (Narratives and Media) are two large-scale interdisciplinary projects. PLI researchers have consecutively received NSF CAREER awards at NeurIPS and ICLR, making it a core node in university LLM research.

Young adults aged 22–25 entering the labor market experienced a 16% relative employment decline in AI-high-exposure occupations. This is not a prediction — it is real data from 2023.

§ 03 / Society

AI's Real
Impact on Society

B01 2023–2025 · Labor Econ
MIT Sloan + NBER
Nobel 2024 Labor

MIT Future of Work: AI Warnings from Three Nobel-Caliber Economists

The "Shaping the Future of Work" initiative, led by Daron Acemoglu (2024 Nobel Prize in Economics), David Autor (leading labor economist), and Simon Johnson, is the most systematic academic investigation of AI's labor impact in the US. The core proposition: the deployment path of AI technology is itself a political-economic choice, not technological determinism.

Acemoglu's flagship conclusion: we have overestimated AI's short-term impact on productivity — much "AI automation" merely outsources low-quality work to machines without genuinely improving output quality. The Stone Center's research on social inequality shows that AI benefits are highly concentrated among a small elite.

Key Data

Young adults aged 22–25 in AI-high-exposure occupations saw a relative employment decline of 16% (2023 actual data). Internship positions decreased noticeably over the same period, and entry-level white-collar jobs were the hardest-hit group.

B02 2024–2025 · RCT
Stanford HAI + SALT
Productivity AI Ethics

Stanford Dual Labs: Measuring AI's Real Impact on Work with Data

Erik Brynjolfsson (Digital Economy Lab) and Diyi Yang (SALT Lab) measure AI's impact on human work from two dimensions. Brynjolfsson's "Generative AI at Work" (in collaboration with MIT) is one of the most rigorous RCTs on AI productivity; Yang's SALT Lab focuses on the friction between AI and social values, particularly cross-cultural and cross-linguistic ethical boundaries.

AI Productivity Gains — Stratified by Worker Type PRODUCTIVITY GAIN
Overall Average +14%
Novice Workers +35%
Senior Workers +4%

Source: Brynjolfsson et al., "Generative AI at Work" (NBER Working Paper, 2023)

B03 1993–2025 · Education
CMU Simon Initiative
AI Tutor Learning Science

CMU Simon Initiative: 30 Years of Learning Science and the Ultimate Fusion with AI Tutors

The Simon Initiative and LearnLab, led by Ken Koedinger, are direct heirs to Herbert Simon's (Nobel Prize in Economics, Turing Award) cognitive science tradition. Their Knowledge Components, Knowledge Tracing, and CogGen framework constitute the most complete theoretical system in the AI tutoring field to date.

CMU's RCT study (2025) yielded a counterintuitive conclusion: Human + AI paired tutoring > AI tutoring alone, even under cost-controlled conditions. AI outperforms humans in personalized difficulty adjustment, but still requires human assistance in recognizing student emotional states and switching strategies at the right moment.

Key Finding (arXiv:2506.20600)

The CogGen framework reveals: the effectiveness of AI tutoring depends heavily on precise modeling of knowledge components. An incorrect KC definition can reduce AI tutoring effectiveness by over 40% — "having enough data" cannot substitute for "having the right cognitive model."

B04 2020–2025 · Gov AI
Stanford Law + HAI
Governance Policy

Stanford RegLab: Making AI Work for Government, Not Against It

The RegLab (Regulation, Evaluation, and Governance Lab), led by Daniel Ho, simultaneously answers two questions: How can government use AI to better execute its functions? How can society design mechanisms to control AI risks? STARA (Statutory Research Assistant) helped the San Francisco City Attorney clear thousands of pages of outdated regulations; an AI developed in collaboration with Princeton identified racially restrictive covenants from 5 million property deeds.

Landmark Project

5 million property deed scan: What would take humans decades, AI completed in weeks, identifying historical evidence of systematic racial discrimination in 20th-century America. RegLab uses randomized controlled trials (RCTs) rather than observational studies, giving conclusions greater credibility in legal and policy circles.

§ 04 / Infra

The Upstream Battlefield
of the AI Capability Race

C01 2023 · SOSP 2023
Berkeley Sky Lab
Inference 50k★

Berkeley Sky Computing: vLLM and Next-Generation AI Inference Infrastructure

vLLM, developed by Ion Stoica, Joseph Gonzalez, and Woosuk Kwon, is currently the world's most widely used open-source LLM inference engine. The core innovation, PagedAttention: borrowing the OS virtual memory paging mechanism, it partitions the KV cache into non-contiguous memory pages, boosting GPU utilization from 30-60% to 90%+, and increasing LLM serving throughput by 2-4x on the same hardware.

vLLM has become the de facto standard for AI infrastructure: 50,000+ GitHub Stars, 800+ contributors, production deployments at Anyscale, Perplexity AI, Baidu, Alibaba Cloud, and others; in July 2024 it became an official project of the Linux Foundation/PyTorch Foundation. SkyPilot provides multi-cloud AI job scheduling, enabling automatic switching and cost optimization across AWS/GCP/Azure.

vLLM vs Traditional Inference — GPU Memory Utilization UTILIZATION
vLLM (PagedAttn) 90%+
Traditional Engine 30-60%

Source: Kwon et al., "Efficient Memory Management for Large Language Model Serving" (SOSP 2023)

C02 2021 · Nature / MoE
Stanford + Ricursive
Chip Design RL / MoE

Stanford Scaling Intelligence: Using AI to Design Chips, Using Chips to Train Stronger AI

Azalia Mirhoseini used RL to train AI for chip design (AlphaChip) at Google DeepMind, publishing a landmark result in Nature in 2021: generating chip layouts in 6 hours, matching or exceeding human engineers on three key metrics, directly used in Google TPU and data center CPU design. In 2023 she joined Stanford to found the Scaling Intelligence Lab, targeting Recursive Self-Improvement.

Mirhoseini's other major contribution is the early development of Mixture-of-Experts (MoE) models — today nearly all frontier large models (GPT-4, Gemini, DeepSeek, Llama 3) use MoE architecture, a foundation she laid during her time at Google Brain. In 2024 she co-founded Ricursive Intelligence, focusing on the core closed loop of "using AI to design chips that train AI."

Recursive Closed Loop

AI designs better chips → chips train stronger AI → stronger AI designs better algorithms → algorithms produce stronger AI. The Scaling Intelligence Lab's mission is to study the feasibility and safety of this closed loop.

§ 05 / Domains

Healthcare & Finance
AI Reaches the Critical Threshold

D01 2021–2025 · FDA Cleared
Harvard Medical School
Medical AI Generalist

Harvard Rajpurkar Lab: Building AI Doctors That Can Truly Practice Medicine

The lab led by Pranav Rajpurkar (Department of Biomedical Informatics, Harvard Medical School) focuses on Generalist Medical AI — integrating CT images, medical records, lab data, and vital signs to perform multimodal reasoning like a physician and make comprehensive diagnoses. This is a systematic challenge to the "specialist model" paradigm.

The co-founded a2z Radiology AI developed the first FDA-cleared CT abdominopelvic multi-finding AI triage system in the US (De Novo authorization), covering 20+ abnormal findings. In 2025, they published ClinicalBench in Nature Medicine and a multi-agent medical system paper in Nature — proving that multiple specialist AI agents collaborating in consultation significantly outperform a single agent.

Latest Progress (arXiv:2509.03906)

Using reinforcement learning to train reasoning chains compatible with human cognition — making AI's diagnostic reasoning process expressible in medical terminology rather than black-box neural network activations. This is a key prerequisite for FDA approval and clinical adoption.

D02 2004–2025 · Quant + AI
MIT Sloan + CSAIL
Finance AI AMH

MIT Financial Engineering Lab: Andrew Lo's AI Trading Systems and Financial Democratization

Andrew Lo (MIT Sloan + CSAIL, Chief Investment Strategist at AlphaSimplex) is a pioneer in integrating behavioral economics, neuroscience, and quantitative trading. His Adaptive Market Hypothesis (AMH), proposed in 2004, reconstructs efficient market theory from an evolutionary biology perspective: market efficiency varies with the environment, and bubbles and crashes result from evolutionary adaptation lags.

Lo's prediction: within 5 years, AI will be able to autonomously manage personal investment portfolios without human financial advisor intervention — not a chatbot giving advice, but genuine goal understanding, dynamic rebalancing, behavioral correction, and personalized modeling. Neurofinance research (skin conductance, functional MRI) reveals the direct association between stress hormones (cortisol) and trading decision quality, providing a neuroscience basis for AI-assisted trading.

Quantamental Integration

LLMs parse 10-K reports, earnings calls, and news articles to extract sentiment signals predictive of asset prices, fused with traditional quantitative factors (price momentum, financial ratios). This is a two-way validation between university finance labs and hedge fund real capital.

§ 06 / Judgments

What Answers
These 14 Labs Provide

01

Agent Capabilities Are Far Below the Hype

WebArena data shows the best Web Agents achieve ~35-45% completion, versus 78% for humans. The best Code Agents on SWE-bench fix ~50% of real GitHub issues — failures concentrated on cross-file context understanding. Claims that "AI can replace programmers" remain overhyped at least through 2025.

Berkeley BAIR / CMU OpenHands
02

AI Productivity Gains Are Highly Unevenly Distributed

The average 14% productivity gain masks extreme stratification: novice workers gain 35%, senior workers only 4%. The biggest beneficiaries are those with the largest skill gaps — meaning AI is a "talent equalizer" in the short term, but also that entry-level jobs face the greatest substitution risk.

Stanford Digital Economy Lab · Brynjolfsson et al.
03

Inference Efficiency Has Become the New Main Battlefield

As the model capability race nears its ceiling, inference efficiency has become the new competitive dimension. vLLM doubles LLM throughput per unit of compute. On-device model research shows a 5.3x improvement in intelligence output per unit of energy from 2023–2025 (architecture contributing 3.1×, hardware 1.7×). A large volume of AI queries can be completed entirely on-device.

Berkeley Sky Computing · vLLM / PagedAttention
04

Medical AI Has Reached the Clinical Threshold

2024–2025 marks a qualitative shift in medical foundation models from "single-task classification" to "general medical reasoning." FDA De Novo clearance, Nature Medicine 2025, Nature 2025 multi-agent medical system paper — technology readiness is approaching the threshold for large-scale clinical trials. What's missing is the regulatory pathway and liability framework, not the technology itself.

Harvard Rajpurkar Lab · a2z Radiology AI
05

AI's Impact on Labor Has Already Happened

Acemoglu et al.'s data is not a prediction — it actually happened in 2023: young adults aged 22–25 experienced a 16% relative employment decline in AI-high-exposure occupations. This is not the natural displacement of technological progress, but the consequence of political-economic choices — who benefits and who bears the costs depends on policy design, not on technology itself.

MIT Future of Work · Acemoglu / Autor / Johnson
06

Universities Do Research Companies Won't

The 87% safety documentation gap in AI agents, the irreplaceability of humans in AI tutoring, the unfairness of AI productivity distribution — these findings all come from universities, not companies. The reason is simple: companies have no incentive to publish data unfavorable to themselves. This is the most irreplaceable value of university AI research in this era.

MIT CSAIL / CMU Simon Initiative / Stanford RegLab

The shared conclusion across these 14 labs: there is a systematic gap between AI's real capabilities and public expectations — the direction of the gap varies by scenario, but its source is consistent: a lack of rigorous measurement of real tasks in real environments. This is the core mission of university researchers, and the entire purpose of this report.

§ 07 / References

Complete Citation Sources

CODA

The value of these 14 labs does not lie in releasing a bigger model faster than OpenAI — their value lies in telling us what that model actually does, ten years after it's released.

University AI Labs Field Report · 2025

Update Log

First published 2026-07-30