Ai agent papers guide
๐ฌ A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | ็ฒพ้ 23,000+ AI Agent ๆ่ฝๅบ๏ผ่ฆ็8ๅคง็คพไผ็งๅญฆๅญฆ็ง็ๅฎ่ฏ็ ็ฉถใCoPaper.AI 20ๅ้ๅฎๆไธ็ฏๅฏๅค็ฐ็่ง่ๅฎ่ฏ่ฎบๆ๏ผๅนถๆฏๆ็จๆทไธไผ Skillsใ-- Maintained by CoPaper.AI from Stanford REAP.
npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill ai-agent-papers-guideAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
What its author says it does
Copied from the file, not written here
Curated 2024-2026 AI agent research papers collection
SKILL.md
4.7 KB, as published. Nobody here has run it
AI Agent Papers Guide (2024-2026)
Overview
A focused collection of AI agent research papers from 2024-2026, tracking the latest developments in LLM-based agent systems. Unlike broader collections, this focuses on recent breakthroughs โ new architectures, benchmarks, multi-agent coordination, and real-world applications. Updated frequently as the field evolves rapidly.
Paper Categories
Recent AI Agent Research
โโโ Agent Architectures
โ โโโ Planning (o1-style reasoning, search-augmented)
โ โโโ Memory (long-term, episodic, working)
โ โโโ Tool use (function calling, code execution)
โโโ Multi-Agent Systems
โ โโโ Collaboration (task decomposition, debate)
โ โโโ Competition (red team, adversarial)
โ โโโ Emergence (self-organization, culture)
โโโ Evaluation
โ โโโ Benchmarks (SWE-bench, WebArena, GAIA)
โ โโโ Safety (jailbreak, misuse, alignment)
โ โโโ Reliability (error recovery, hallucination)
โโโ Applications
โ โโโ Software engineering (coding agents)
โ โโโ Scientific research (lab automation)
โ โโโ Web automation (browsing, form-filling)
โ โโโ Enterprise (workflow, data analysis)
โโโ Infrastructure
โโโ Frameworks (LangGraph, CrewAI, AutoGen)
โโโ Protocols (MCP, A2A, tool standards)
โโโ Deployment (scaling, monitoring, cost)
Highlighted Papers (2024-2025)
| Paper | Venue | Key Contribution |
|---|---|---|
| SWE-agent | ICLR 2025 | Agent interface design for SE |
| OpenHands | 2024 | Open platform for coding agents |
| AgentBench | ICLR 2024 | Multi-environment agent benchmark |
| GAIA | ICLR 2024 | General AI assistant benchmark |
| Voyager | NeurIPS 2024 | Lifelong learning in Minecraft |
| OS-Copilot | 2024 | Self-improving computer agent |
| AutoGen | 2024 | Multi-agent conversation framework |
| Agent-FLAN | ACL 2024 | Agent fine-tuning methodology |
Tracking New Papers
import arxiv
from datetime import datetime, timedelta
def find_recent_agent_papers(days=14):
"""Find cutting-edge agent papers."""
queries = [
"ti:agent AND (ti:LLM OR ti:language model)",
"abs:autonomous agent AND abs:tool use AND abs:2024",
"ti:multi-agent AND abs:large language",
"abs:coding agent OR abs:software agent",
]
seen = set()
papers = []
for q in queries:
search = arxiv.Search(
query=q, max_results=15,
sort_by=arxiv.SortCriterion.SubmittedDate,
)
for r in search.results():
if r.entry_id not in seen:
seen.add(r.entry_id)
papers.append({
"title": r.title,
"date": r.published.strftime("%Y-%m-%d"),
"url": r.entry_id,
})
papers.sort(key=lambda x: x["date"], reverse=True)
for p in papers[:20]:
print(f"[{p['date']}] {p['title']}")
print(f" {p['url']}")
find_recent_agent_papers()
Framework Comparison
frameworks = {
"LangGraph": {
"paradigm": "Graph-based workflows",
"persistence": "Built-in checkpointing",
"multi_agent": "Yes",
"language": "Python/JS",
},
"CrewAI": {
"paradigm": "Role-based agents",
"persistence": "Memory module",
"multi_agent": "Yes (crew)",
"language": "Python",
},
"AutoGen": {
"paradigm": "Conversational agents",
"persistence": "Chat history",
"multi_agent": "Yes (group chat)",
"language": "Python/.NET",
},
"OpenHands": {
"paradigm": "Computer use agent",
"persistence": "Workspace state",
"multi_agent": "No",
"language": "Python",
},
}
for name, info in frameworks.items():
print(f"\n{name}:")
for k, v in info.items():
print(f" {k}: {v}")
Use Cases
- Literature tracking: Stay current on agent research
- Framework selection: Compare agent development tools
- Research planning: Identify open problems and trends
- Course material: Teach cutting-edge agent systems
- Benchmark tracking: Compare agent capabilities