Ph.D. Student · Emory University

Shanglin Wu (Jason)

I’m a second-year Ph.D. student in Computer Science and Informatics at Emory University, advised by Dr. Kai Shu. I received my B.S. in Artificial Intelligence from Yuanpei College, Peking University in 2025.

My research focuses on understanding the fundamental principles that govern multi-agent AI systems, particularly how autonomous agents collaborate, learn, and adapt through interaction. I am especially interested in how collective behaviors emerge from individual agents, how agents learn continually from experience, and how coordination failures and emergent risks arise as these systems scale. My goal is to develop a deeper understanding of the mechanisms underlying learning, collaboration, and safety in increasingly autonomous multi-agent systems.

Multi-Agent Systems Agent Memory Collective Intelligence AI Safety
Portrait of Shanglin Wu

News

  • 2026.09JARVIS-Bench is accepted to NeurIPS 2026 (E&D Track).
  • 2026.06Our paper Rethinking Memory Mechanisms of Foundation Agents in the Second Half: A Survey is accepted by TMLR.
  • 2026.06Memory in LLM-based Multi-agent Systems received the Best Survey Paper Award 🏆 at PAKDD 2026.
  • 2026.05Honored to receive the Excellence in Teaching Assistance Commendation (2025–26) from the Department of Computer Science.
  • 2026.04Joining Cisco Research as an AI/Intelligent Systems PhD Intern for Summer 2026.
  • 2026.02Memory in LLM-based Multi-agent Systems is accepted by PAKDD 2026.
  • 2025.07Completed my internship at Microsoft Research Asia.

Publications

05
Overview figure of the foundation agent memory survey
TMLRCore contributor

Rethinking Memory Mechanisms of Foundation Agents in the Second Half: A Survey

Wei-Chieh Huang, Weizhi Zhang, Yueqing Liang, …, Shanglin Wu, …, Kai Shu

Transactions on Machine Learning Research

arXiv

Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-world evaluation. As the field enters the “second half,” the central challenge becomes real utility in long-horizon, dynamic, and user-dependent settings such as agentic coding, deep research, and computer use, where LLM-based agents face context explosion beyond fixed context windows and must continuously accumulate, manage, and selectively reuse information across extended interactions. Memory, with hundreds of papers released in 2025, therefore emerges as the critical solution to fill this utility gap. Beyond passive storage, memory is increasingly the substrate through which agents self-evolve: short-term memory gates which experiences are perceived and abstracted during execution, while long-term memory consolidates them into reusable knowledge and skills, forming the loop through which agents improve from their own experience. In this survey, we provide a unified view of foundation agent memory along three dimensions: memory substrate (internal parametric state and external retrieval-augmented stores), cognitive mechanism (sensory, working, episodic, semantic, and procedural), and memory subject (user-centric personalization and agent-centric experience). We then analyze how memory is operated under single- and multi-agent topologies and highlight learning policies over memory operations, showing how memory management itself is becoming a trainable capability spanning reinforcement-learned context curation, experience consolidation at decision time, and the emerging ecosystem of portable, shareable agent skills. Finally, we review evaluation benchmarks and metrics for memory utility, and outline open challenges and future directions.

Overview figure of LLMA-Mem
Preprint · 2026

Scaling Teams or Scaling Time? Memory Enabled Lifelong Learning in LLM Multi-Agent Systems

Shanglin Wu, Yuyang Luo, Yueqing Liang, Kaiwen Shi, Yanfang Ye, Ali Payani, Kai Shu

arXiv

Large language model (LLM) multi-agent systems can scale along two distinct dimensions: by increasing the number of agents and by improving through accumulated experience over time. Although prior work has studied these dimensions separately, their interaction under realistic cost constraints remains unclear. In this paper, we introduce a conceptual scaling view of multi-agent systems that jointly considers team size and lifelong learning ability, and we study how memory design shares this landscape. To this end, we propose LLMA-Mem, a lifelong memory framework for LLM multi-agent systems under flexible memory topologies. We evaluate LLMA-Mem on MultiAgentBench across coding, research, and database environments. Empirically, LLMA-Mem consistently improves long-horizon performance over baselines while reducing cost. Our analysis further reveals a non-monotonic scaling landscape: larger teams do not always produce better long-term performance, and smaller teams can outperform larger ones when memory better supports the reuse of experience. These findings position memory design as a practical path for scaling multi-agent systems more effectively and more efficiently over time.

Other Publications

  1. JARVIS-Bench: Benchmarking Personal Intelligence Agents on Long-Horizon Real-User Daily Traces
    Weizhi Zhang, Wei-Chieh Huang, Yueqing Liang, …, Shanglin Wu, …, Philip S. Yu
    NeurIPS 2026Evaluations & Datasets Track OpenReview
  2. Improving Factuality in LLMs via Inference-Time Knowledge Graph Construction
    Shanglin Wu, Lihui Liu, Jinho D. Choi, Kai Shu
    Preprint · 2025 arXiv

    Large Language Models (LLMs) often struggle with producing factually consistent answers due to limitations in their parametric memory. Retrieval-Augmented Generation (RAG) paradigms mitigate this issue by incorporating external knowledge at inference time. However, such methods typically handle knowledge as unstructured text, which reduces retrieval accuracy, hinders compositional reasoning, and amplifies the influence of irrelevant information on the factual consistency of LLM outputs. To overcome these limitations, we propose a novel framework that dynamically constructs and expands knowledge graphs (KGs) during inference, integrating both internal knowledge extracted from LLMs and external knowledge retrieved from external sources. Our method begins by extracting a seed KG from the question via prompting, followed by iterative expansion using the LLM’s internal knowledge. The KG is then selectively refined through external retrieval, enhancing factual coverage and correcting inaccuracies. We evaluate our approach on three diverse Factual QA benchmarks, demonstrating consistent gains in factual accuracy over baselines. Our findings reveal that inference-time KG construction is a promising direction for enhancing LLM factuality in a structured, interpretable, and scalable manner.

Experience & Education

Research Experience

  • Cisco Research
    AI/Intelligent Systems PhD Intern
    May 2026 – Aug 2026
  • Microsoft Research Asia
    Research Intern · Beijing, China
    Mar 2025 – Jul 2025

Education

  • Emory University
    Ph.D. in Computer Science and Informatics
    2025 – Present
  • Peking University
    B.S. in Artificial Intelligence · Yuanpei College
    2021 – 2025

Honors & Service

Honors
Excellence in Teaching Assistance Commendation, Emory University, 2026
Awarded by the Department of Computer Science for exceptional dependability, proactivity, and initiative.
Program Committee
WSDM 2027
Reviewer
ICLR 2027IEEE BigData 2027NeurIPS 2026NeurIPS HAIC Workshop 2026