Research
Curious learnings from the AI frontier. Papers we read, summaries we wrote, things that surprised us. Not for profit, just for understanding.
-
The Illusion of Independent Quorums: Epistemic Fault Domains in Agentic Systems
He & Yu · 2026
Multi-Agent Safety · Governance -
Invalidation Contracts for Cross-Episode Agent Memory
Wu & Canedo · 2026
Agent Memory · Protocol Design -
Context as an Environment: Programmatic Context Management for Long-Horizon Agents
Lin, Ang, Zhu, Ding & Zhou · 2026
Long-Context Management · Agent Architecture -
When Does LLM Orchestration Pay Off? A Controlled Evaluation
Leins, Pelleriti, Gonnermann-Müller & Pokutta · 2026
LLM Orchestration · Cost Analysis -
Sliding-Window Beats Linear Attention
Jolicoeur-Martineau, Sukthanker, Cameron & Gervais · 2026
Model Architecture · Inference Efficiency -
PILOT in the Loop: Self-Improvement During the Run
Xiao, Sun, Wu, Hui et al. · 2026
Agent Self-Improvement · Inference Efficiency -
SWE Refactor Bench: When Agents Pass the Tests Without Doing the Migration
Hong, Chi, Li, Wang, Gao et al. · 2026
Agent Evaluation · Code Migration -
Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools
Anand & Chattaraj · 2026
Agent Evaluation · Tool Selection -
There Is No Neutral Harness: Leaderboards Are Manufactured by Config-Fragile Items
Parupudi · 2026
Evaluation Methodology · LLM Benchmarking -
EnvHarness: Wrap the Training Environment, Not the Model
Huang, Wang, Han, Yan, Chen et al. · Google · 2026
Agent Training · Environment Engineering -
AI4AI-Bench: Can Agents Redesign the Training Algorithm?
Chi, Li, Hong, Wang et al. · 2026
Agent Evaluation · Algorithmic Self-Improvement -
Phantom Gains: When Self-Improvement Measurement Lies
Xu, Yan, Chen & Kechadi · 2026
Evaluation Methodology · LLM Self-Improvement -
When Tool Retrieval Fails: Source-Style Collapse in Skill Retrieval
Liu, James, Wang, Xiao & Lin · University of Manchester, University of Sheffield & SUFE · 2026
Tool Retrieval · Distribution Shift -
Not Worth Another Token: Where You Prune Beats How You Score
Kolukuluru, Ashok, Arora, Dernoncourt, Rossi, Lipka et al. · Adobe Research · 2026
Cost-Aware Agents · Context Pruning -
Demystifying Agent Skills: Why They Work, Until They Don't
Jiang, Huang, Xing, Wang, Liu, Li et al. · 2026
Agent Design · Skill Libraries -
REDAgentBench: When Safety Evals Miss the Actual Harm
Chen, Liu, Zhu, Dou et al. · 2026
Agent Safety · Evaluation Methodology -
Beyond Final Scores: What Agent Evals Can't Tell You
Li, Yang, Tan, Huang et al. · 2026
Agent Evaluation · Benchmarking Methodology -
The Truncation Problem: When Your Long-Context Benchmark Measures Signal Loss, Not Context
Arjmandi · 2026
Long-Context Evaluation · Benchmarking Methodology -
Think Big, Search Small: Where Capacity Matters in Hierarchical Search Agents
Cai, Zhao & Li · East China Normal University · 2026
Multi-Agent Systems · Cost-Aware Architecture -
Prompt-Induced Waste: When Your Words Drive Up the Bill
Weinberger & Hozez · 2026
Cost-Aware Agents · Prompt Engineering -
Zero-Mem: Agent Memory Without the Model Call
Xiao, Zhu, Zhang et al. · 2026
Agent Memory · Cost-Aware Agents -
Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition
Feng, Zhang, Cheng & Qi · 2026
Cost-Aware Agents · Tool Selection -
Interpretability Can Be Actionable
Orgad, Barez, Haklay et al. · 2026
ML Interpretability · Position Paper -
Compute Where It Counts: Per-Token Efficiency Without Retraining
Akhauri & Abdelfattah · Cornell University · 2026
Efficient Inference · LLM Serving · Deployment & Operations -
Drowning in Documents: Why Long-Context Retrieval Collapses and How to Fix It
Gollapudi, Gupta, Singhal, Min et al. · 2026
Long-Context Retrieval · Attention Mechanisms -
OmegaUse-OfficeVal: What Office Agents Actually Cost to Use
Zhou, Zhao et al. · Baidu · 2026
Agent Evaluation · Economic Grounding · Evaluation -
Token Reduction Is Not Cost Reduction: What API Billing Actually Measures
Weinberger & Hozez · PointFive · 2026
Agent Cost · API Efficiency · Deployment & Operations -
Context Is the First Failure Point: Measuring Agent Reliability Before It Runs
Bousetouane · ProofAgent.ai / University of Chicago · 2026
Context Engineering · Agent Reliability · Governance -
CompactionRL: Teaching Agents to Summarize What Matters
Li, Hou, Jing, Tang & Dong · Tsinghua University · 2026
Long-Context Agents · Reinforcement Learning -
Agora: When Models Bid for the Work They Do Best
Zhou, Leonardis & Feng · University of Birmingham · 2026
Agent Routing · Multi-Model Allocation -
MemSyco-Bench: When Memory Makes Agents More Agreeable, Not More Accurate
Xiang, Chen, Tang, Wei, Ning, Lin, Zhang & Su · 2026
Agent Memory · Sycophancy Benchmarks -
Agent Cognitive Redundancy: What Agents Spend on Tasks That Don't Need It
Yin & Feng · University of Tennessee, Knoxville · 2026
Agent Efficiency · Cost-Aware Execution -
PolyWorkBench: When Workflows Mix Languages, Agents Compound Mistakes
Li, Liu, Zhang et al. · Beijing Jiaotong University & Tencent Weixin AI · 2026
Agent Evaluation · Multilingual Workflows -
When the Memory Lies: Forged Reasoning Attacks on LLM Agents
Karamchandani, Nagasubramaniam, Zhu & Wu · Pennsylvania State University · 2026
Agent Security · Memory Attacks -
SwarmResearch: Orchestrating Coding Agents for Open-Ended Discovery
Virk, Edds, Xia & Zhang · University of Illinois Urbana-Champaign · 2026
Multi-Agent Systems · Open-Ended Optimization -
Stop Retrieving, Start Navigating: Memory as a Structured Action Space
Xu, Sun et al. · Alibaba & ShanghaiTech University · 2026
Agent Memory · Reinforcement Learning -
UniClawBench: Grading Agents Step by Step in Live Environments
Chen, Duan, Sun et al. · HKU MMLab · 2026
Agent Evaluation · Benchmarking -
The Loop That Checks Itself: Self-Evolving Agents with Anytime-Valid Certificates
Biswa Sengupta · 2026
Agent Self-Evolution · Statistical Guarantees -
Why Reasoning Fails to Plan: FLARE and Long-Horizon Agent Decision Making
Wang et al. · Notre Dame, Stanford, Yale et al. · 2026
Planning · LLM Agents -
The Agent That Writes Its Own Playbook
Ding, Xie, Wei, Li & Ding · Renmin University of China & Alibaba Group · 2026
Agent Self-Evolution · Tool Optimization -
Reason Less, Verify More: Deterministic Gates for Policy-Safe Agent Tool Use
Reddy, Challaram & Basu · 2026
Agent Reliability · Runtime Verification -
Static Training, Shifting World: Why Tool-Using Agents Break in the Open
Lv, Wu, Zhu, Cheng & Guo · Nanjing University LAMDA · 2026
Agent Robustness · Tool Use -
BAGEN: Do Agents Know When a Task Is Doomed?
Lin, Wang et al. · Northwestern MLL Lab · 2026
Cost-Aware Agents · Budget Metacognition -
LLM-as-an-Investigator: When the AI Asks Before It Answers
Marozzo & Liò · University of Calabria & University of Cambridge · 2026
Diagnostic AI · Conversational Reasoning -
Skill Retrieval Needs a Set-Level Lens
Wang, Wen et al. · Tencent · 2026
Agent Skills · Retrieval -
Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism
Zhang, Xu, Li, Zhang, Jiang, Zhang & McAuley · UC San Diego · 2026
Agentic Search · Context Management -
Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning
Xiao, Ma, Chen, Chen, Atreya, Chen & Ordonez · Rice University / Meta · 2026
Reasoning · Retrieval -
A Unified Framework for the Evaluation of LLM Agentic Capabilities
Zhu, Li, Lyu, Luo, Yang, Liu, Hui, Yuan, Sun, Su & Shao · Beijing University of Posts and Telecommunications, Shanghai AI Lab & collaborators · 2026
Agent Evaluation -
DeployBench: Benchmarking LLM Agents for Research Artifact Deployment
Wang, Qian, Zhang et al. · Boston University, Northeastern University & University of Texas at Dallas · 2026
Agent Evaluation -
Evaluate Agents Without Deploying Them: Off-Policy Evaluation via Diffusion World Models
Liu, Xiong, Zhang & Tang · Emory University & Shanghai Jiao Tong University · 2026
Agent Evaluation -
Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents
Sun, Wang, Song, He, W. Zhang, Y. Liu, Y. Yang & Y. Cheng · Tongji University, Shanghai AI Lab & collaborators · 2026
Agent Evaluation -
ContextPilot: Faster Long-Context LLM Inference via Context Reuse
Jiang, Huang, Cheng, Deng, Sun & Mai · University of Edinburgh · 2026
LLM Inference -
Prompt Injection's Impossible Defense: The Contextual Integrity Problem
Abdelnabi & Bagdasarian · University of Massachusetts CICS · 2026
Agent Security -
Beyond Consensus: When Traces Know More Than Votes
Fadnavis, Kanakaraj & Wyss · Bioscope AI · 2026
Multi-Agent Systems -
VibeSearchBench: When Real Search Meets a Benchmark
Xiaohongshu Inc. · 2026
Agent Evaluation -
MUSE-Autoskill: The Skill Lifecycle Agents Were Missing
Lin, Li, Song, Jiang & Zhang · ByteDance · 2026
Agent Architecture -
Is Agent Memory a Database? The Missing Data-Foundations Layer
Orogat & Mansour · Concordia University · 2026
Agent Memory -
Compute Where it Counts: Per-Token Efficiency in Frozen LLMs
Akhauri & Abdelfattah · Cornell University · 2026
Inference Optimization -
Interpretability Can Be Actionable: A Rubric for the Field
Orgad, Barez, Haklay et al. · multiple institutions · 2026
Interpretability -
Remembering More, Risking More: How Agent Memory Accumulates Safety Risk
Al-Tawaha, Gu, Niu, Jia & Jin · Virginia Tech, UC Berkeley & UIUC · 2026
Agent Safety -
δ-mem: Adding Persistent Memory to Any Frozen LLM
Lei, Zhang, Li, Wang et al. · Nanyang Technological University · 2026
Agent Memory -
AgentTrust: A Runtime Safety Layer for Every Tool Call
Yang · Independent · 2026
Agent Safety -
Automated Interpretability: Two Agent Loops, One Autonomous Pipeline
Marin-Llobet & Ferrando · Harvard University · 2026
Mechanistic Interpretability -
AutoTTS: When LLMs Discover Their Own Reasoning Strategies
Zheng, Liu, Huang et al. · UMD, UVA & UNC · 2026
Inference Optimization -
Dual-Dimensional Consistency: Smarter Self-Consistency Sampling
Xu, Li, Zhao, Wu, Li & Yan · Xi'an Jiaotong University · 2026
Inference Optimization -
Agentic Systems as Boosting: When Weak Models Beat the Frontier
Sunkaraneni, Beneventano, Neumarker, Poggio & Galanti · MIT & Texas A&M · 2026
Agent Architecture -
Is Grep All You Need? The Agent Harness Moves Accuracy More Than the Retrieval Method
Sen, Kasturi, Lumer, Gulati, Subbiah et al. · 2026
Agentic Search -
BoundaryRouter: Learning When to Escalate to an Agent
Wang, Qiu et al. · Princeton, Michigan, Tsinghua et al. · 2026
Agent Routing -
LaTER: Latent-Phase Reasoning Cuts Tokens 32% Without Losing Accuracy
Li, Wang, Liu et al. · 2026
Inference Optimization -
ComplexMCP: Three Failure Modes in Large-Scale Tool Sandboxes
Li, Yang, Wang et al. · 2026
Agent Evaluation -
STALE: When Agent Memory Becomes a Liability
Chao, Bai et al. · 2026
Agent Memory -
AI Co-Mathematician: When Scaffolding Beats the Model
Zheng, von Glehn, Zwols et al. · Google DeepMind · 2026
Agentic AI -
Meta-Harness: The 6x Gap Lives in Your Code, Not Your Model
Lee, Nair, Zhang, Lee, Khattab & Finn · Stanford University & MIT · 2026
AI Systems -
How Coding Agents Actually Perform in the Wild
Popescu, Gros, Botocan, Pandita, Devanbu & Izadi · TU Delft & UC Davis · 2026
Software Engineering -
Conversation Reduces Load. Images Build It.
Taneja, Singh & Goel · Georgia Institute of Technology · 2026
AI in Education -
Image Generation Diversity: When Models Miss the Map
Dombrowski, Zhang, Cechnicka, Reynaud & Kainz · FAU Erlangen-Nürnberg & Imperial College London · 2025
Generative AI -
SLOW: The AI Tutor That Thinks Before It Speaks
Wei, Li & Jiang · Shanghai Institute of AI for Education · 2026
AI Tutoring -
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
Tran & Kiela · Stanford University · 2026
Agent Architecture -
LLMs in Games: When Generated Content Runs the Rules
Johnson, Ahmed, Lang, Thethi, Zheng & de Souza Santos · University of Calgary · 2026
Game Development -
Arknights: When the AI Lies, Players Learn
Shuai Guo · Uppsala University · 2025
Explainable AI -
Vibe Coding: Flow, Trust, and Co-Creation
Pimenova, Fakhoury, Bird, Storey & Endres · U Michigan / Microsoft Research · 2025
Vibe Coding -
Games That Teach AI Ethics
Solyst, Nakigozi, Fong & Shapiro · University of Washington · 2025
AI Education -
BAVT: Spend Less, Reason Better
Li et al. · UBC / Vector Institute · 2026
AI Agents
These summaries are layperson interpretations of published research. They are not peer-reviewed and may simplify or omit nuance. Always refer to the original papers for complete findings.
Synthesized by Kelly Chiang & Claude.