Research
Curious learnings from the AI frontier. Papers we read, summaries we wrote, things that surprised us. Not for profit, just for understanding.
-
Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism
Zhang, Xu, Li, Zhang, Jiang, Zhang & McAuley · UC San Diego · 2026
Agentic Search · Context Management -
Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning
Xiao, Ma, Chen, Chen, Atreya, Chen & Ordonez · Rice University / Meta · 2026
Reasoning · Retrieval -
A Unified Framework for the Evaluation of LLM Agentic Capabilities
Zhu, Li, Lyu, Luo, Yang, Liu, Hui, Yuan, Sun, Su & Shao · Beijing University of Posts and Telecommunications, Shanghai AI Lab & collaborators · 2026
Agent Evaluation -
DeployBench: Benchmarking LLM Agents for Research Artifact Deployment
Wang, Qian, Zhang et al. · Boston University, Northeastern University & University of Texas at Dallas · 2026
Agent Evaluation -
Evaluate Agents Without Deploying Them: Off-Policy Evaluation via Diffusion World Models
Liu, Xiong, Zhang & Tang · Emory University & Shanghai Jiao Tong University · 2026
Agent Evaluation -
Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents
Sun, Wang, Song, He, W. Zhang, Y. Liu, Y. Yang & Y. Cheng · Tongji University, Shanghai AI Lab & collaborators · 2026
Agent Evaluation -
ContextPilot: Faster Long-Context LLM Inference via Context Reuse
Jiang, Huang, Cheng, Deng, Sun & Mai · University of Edinburgh · 2026
LLM Inference -
Prompt Injection's Impossible Defense: The Contextual Integrity Problem
Abdelnabi & Bagdasarian · University of Massachusetts CICS · 2026
Agent Security -
Beyond Consensus: When Traces Know More Than Votes
Fadnavis, Kanakaraj & Wyss · Bioscope AI · 2026
Multi-Agent Systems -
VibeSearchBench: When Real Search Meets a Benchmark
Xiaohongshu Inc. · 2026
Agent Evaluation -
MUSE-Autoskill: The Skill Lifecycle Agents Were Missing
Lin, Li, Song, Jiang & Zhang · ByteDance · 2026
Agent Architecture -
Is Agent Memory a Database? The Missing Data-Foundations Layer
Orogat & Mansour · Concordia University · 2026
Agent Memory -
Compute Where it Counts: Per-Token Efficiency in Frozen LLMs
Akhauri & Abdelfattah · Cornell University · 2026
Inference Optimization -
Interpretability Can Be Actionable: A Rubric for the Field
Orgad, Barez, Haklay et al. · multiple institutions · 2026
Interpretability -
Remembering More, Risking More: How Agent Memory Accumulates Safety Risk
Al-Tawaha, Gu, Niu, Jia & Jin · Virginia Tech, UC Berkeley & UIUC · 2026
Agent Safety -
δ-mem: Adding Persistent Memory to Any Frozen LLM
Lei, Zhang, Li, Wang et al. · Nanyang Technological University · 2026
Agent Memory -
AgentTrust: A Runtime Safety Layer for Every Tool Call
Yang · Independent · 2026
Agent Safety -
Automated Interpretability: Two Agent Loops, One Autonomous Pipeline
Marin-Llobet & Ferrando · Harvard University · 2026
Mechanistic Interpretability -
AutoTTS: When LLMs Discover Their Own Reasoning Strategies
Zheng, Liu, Huang et al. · UMD, UVA & UNC · 2026
Inference Optimization -
Dual-Dimensional Consistency: Smarter Self-Consistency Sampling
Xu, Li, Zhao, Wu, Li & Yan · Xi'an Jiaotong University · 2026
Inference Optimization -
Agentic Systems as Boosting: When Weak Models Beat the Frontier
Sunkaraneni, Beneventano, Neumarker, Poggio & Galanti · MIT & Texas A&M · 2026
Agent Architecture -
Is Grep All You Need? The Agent Harness Moves Accuracy More Than the Retrieval Method
Sen, Kasturi, Lumer, Gulati, Subbiah et al. · 2026
Agentic Search -
BoundaryRouter: Learning When to Escalate to an Agent
Wang, Qiu et al. · Princeton, Michigan, Tsinghua et al. · 2026
Agent Routing -
LaTER: Latent-Phase Reasoning Cuts Tokens 32% Without Losing Accuracy
Li, Wang, Liu et al. · 2026
Inference Optimization -
ComplexMCP: Three Failure Modes in Large-Scale Tool Sandboxes
Li, Yang, Wang et al. · 2026
Agent Evaluation -
STALE: When Agent Memory Becomes a Liability
Chao, Bai et al. · 2026
Agent Memory -
AI Co-Mathematician: When Scaffolding Beats the Model
Zheng, von Glehn, Zwols et al. · Google DeepMind · 2026
Agentic AI -
Meta-Harness: The 6x Gap Lives in Your Code, Not Your Model
Lee, Nair, Zhang, Lee, Khattab & Finn · Stanford University & MIT · 2026
AI Systems -
How Coding Agents Actually Perform in the Wild
Popescu, Gros, Botocan, Pandita, Devanbu & Izadi · TU Delft & UC Davis · 2026
Software Engineering -
Conversation Reduces Load. Images Build It.
Taneja, Singh & Goel · Georgia Institute of Technology · 2026
AI in Education -
Image Generation Diversity: When Models Miss the Map
Dombrowski, Zhang, Cechnicka, Reynaud & Kainz · FAU Erlangen-Nürnberg & Imperial College London · 2025
Generative AI -
SLOW: The AI Tutor That Thinks Before It Speaks
Wei, Li & Jiang · Shanghai Institute of AI for Education · 2026
AI Tutoring -
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
Tran & Kiela · Stanford University · 2026
Agent Architecture -
LLMs in Games: When Generated Content Runs the Rules
Johnson, Ahmed, Lang, Thethi, Zheng & de Souza Santos · University of Calgary · 2026
Game Development -
Arknights: When the AI Lies, Players Learn
Shuai Guo · Uppsala University · 2025
Explainable AI -
Vibe Coding: Flow, Trust, and Co-Creation
Pimenova, Fakhoury, Bird, Storey & Endres · U Michigan / Microsoft Research · 2025
Vibe Coding -
Games That Teach AI Ethics
Solyst, Nakigozi, Fong & Shapiro · University of Washington · 2025
AI Education -
BAVT: Spend Less, Reason Better
Li et al. · UBC / Vector Institute · 2026
AI Agents
These summaries are layperson interpretations of published research. They are not peer-reviewed and may simplify or omit nuance. Always refer to the original papers for complete findings.
Synthesized by Kelly Chiang & Claude.