Knowledge & Learning 6/17/2026 3 min@ZachasADMINCoDA-Bench tests whether coding agents can find the right data before writing codeCoDA-Bench is a new ICML 2026 benchmark for code agents that must search noisy data folders, identify relevant files, write code, and answer analytical questions.Read more
Knowledge & Learning 6/15/2026 3 min@ZachasADMINEvaluation Cards exposes why AI benchmark scores are hard to trustEvalEval's beta Evaluation Cards project maps AI evaluation results with reproducibility, completeness, provenance, and comparability signals.Read more
Knowledge & Learning 6/12/2026 3 min@ZachasADMINAgents' Last Exam tests AI agents on real professional workflowsAgents' Last Exam is a new Berkeley-led benchmark for computer-use AI agents, with long-horizon professional tasks, verifiable outcomes, public tooling, and early results showing wide gaps on hard real-world work.Read more
Knowledge & Learning 6/12/2026 4 min@ZachasADMINHugging Face benchmark tests voice agents on code-switched customer speechServiceNow-AI published a Hugging Face benchmark and dataset for code-switched ASR, testing how voice-agent transcription handles Spanish-English, French-English, Canadian French-English, and German-English support scenarios.Read more
Knowledge & Learning 6/9/2026 3 min@ZachasADMINNew arXiv Paper Tests Compact Models Against LLMs for Multilingual Fact-CheckingA June 2026 arXiv paper from Factiverse reports that compact fine-tuned models can stay practical for multilingual fact-checking when latency, cost, and privacy matter.Read more
Knowledge & Learning 6/4/2026 3 min@ZachasADMINSkillOpt trains agent skills as editable artifacts, not model weightsSkillOpt is a Microsoft Research project and arXiv paper that treats natural-language agent skills as trainable external state, using scored rollouts, bounded edits, and held-out validation instead of model fine-tuning.Read more