Topic
#benchmark
Loot, blog posts and adjacent themes connected to this topic. Follow the tag to keep it in your orbit.
1Shown loot
3Shown articles
8Linked neighbor tags
Topic paths
If you want to go deeper, the adjacent tags are the fastest way to compare and branch into related workflows.
Loot
More from this topic
Blog
Related reads
Wissen & Lernen
PlanBench-XL tests whether agents can recover when tool paths break
PlanBench-XL is a June 2026 arXiv benchmark for long-horizon LLM tool-use agents, with 327 retail tasks, 1,665 tools, retrieval-limited visi…
Wissen & Lernen
MosaicLeaks shows how research-agent search queries can leak private data
MosaicLeaks is a new benchmark for deep-research agents that shows how external web queries can expose private enterprise facts through the …
Wissen & Lernen
Agents' Last Exam tests AI agents on real professional workflows
Agents' Last Exam is a new Berkeley-led benchmark for computer-use AI agents, with long-horizon professional tasks, verifiable outcomes, pub…
