Knowledge & Learning 6/21/2026 3 min@ZachasADMINA practical rule for AI code review: reject working diffs you cannot explainA June 2026 developer essay and active Hacker News discussion point to a recurring coding-agent problem: green CI is not enough when the reviewer cannot explain the approach.Read more
Knowledge & Learning 6/21/2026 3 min@ZachasADMINLedgerAgent tests structured state for policy-bound tool-calling agentsA new arXiv preprint proposes LedgerAgent, an inference-time method that keeps customer-service agent state in a separate ledger before policy-sensitive tool calls.Read more
Knowledge & Learning 6/20/2026 3 min@ZachasADMINSIA Tests Self-Improving AI Across Agent Harnesses and Model WeightsA new arXiv paper and official implementation show SIA updating both an agent scaffold and model weights, with reported gains on LawBench, GPU kernels, and single-cell RNA denoising.Read more
Knowledge & Learning 6/19/2026 3 min@ZachasADMINMosaicLeaks shows how research-agent search queries can leak private dataMosaicLeaks is a new benchmark for deep-research agents that shows how external web queries can expose private enterprise facts through the mosaic effect.Read more
Knowledge & Learning 6/19/2026 3 min@ZachasADMINCEO-Bench Tests Whether AI Agents Can Run a Startup for 500 DaysRead more
Knowledge & Learning 6/18/2026 4 min@ZachasADMINWorkBench Revisited Shows Why Workplace Agent Scores Need Source-Level ChecksWorkBench Revisited updates a workplace-agent benchmark with 2026 model runs, but the arXiv abstract and GitHub repository currently surface different top-line scores, making it a useful case study in how to verify agent benchmarks before citing them.Read more