#research
Loots, Blogposts und verwandte Themen rund um diesen Tag. Folge dem Tag, damit passende Updates in deinem Orbit bleiben.
Wenn du tiefer einsteigen willst, helfen die benachbarten Tags beim Vergleichen und Querlesen.
Mehr aus diesem Thema
Verwandte Artikel
Anthropic launches Claude Fable 5.1 and Mythos 5.1 for coding and research
Anthropic has launched Claude Fable 5.1 and Claude Mythos 5.1 for coding and knowledge work, with Fable 5.1 already available in GitHub Copi…
OpenAI says Astra generated ten mathematical advances with Lean certificates
OpenAI says an internal version of its next major model, Astra, produced ten new results across mathematics and theoretical computer science…
Use Claude Science only after your research workflow passes audit checks
Anthropic launched Claude Science, an AI workbench for scientists that combines research tools, auditable artifacts, compute access, and a c…
AOHP proposes an Android-based OS harness for AI agents
AOHP is a new arXiv and Hugging Face trending paper that treats AI agents as first-class OS actors inside an Android Open Source Project bas…
LedgerAgent tests structured state for policy-bound tool-calling agents
A new arXiv preprint proposes LedgerAgent, an inference-time method that keeps customer-service agent state in a separate ledger before poli…
SIA Tests Self-Improving AI Across Agent Harnesses and Model Weights
A new arXiv paper and official implementation show SIA updating both an agent scaffold and model weights, with reported gains on LawBench, G…
CEO-Bench Tests Whether AI Agents Can Run a Startup for 500 Days
Evaluation Cards exposes why AI benchmark scores are hard to trust
EvalEval's beta Evaluation Cards project maps AI evaluation results with reproducibility, completeness, provenance, and comparability signal…
AgingBench asks how long AI agents stay reliable after deployment
AgingBench is a new benchmark for long-lived AI agents, measuring reliability decay across sessions instead of only testing freshly initiali…
The Open Agent Leaderboard compares full AI agent systems, not just models
IBM Research and Hugging Face introduced the Open Agent Leaderboard, an open benchmark stack for comparing complete AI agent systems across …




