Supabase开源了评估框架supabase/evals,用于测试AI代理构建Supabase项目的表现。它运行真实任务并评分,发现代理在声明式模式、新库发现和技能激活上存在不足,但技能加载能提升表现。结果已公开,未来将扩展场景并改进评分。
QCon AI Boston 2026 focused on the operational challenges of deploying AI agents, emphasizing the need for robust production infrastructure. Key themes included improving context management,...
Mallika Rao discusses the hidden risk of evaluation debt in production AI systems, drawing on her experience at Twitter, Walmart, and Netflix. She explains why traditional metrics fail modern...
In this podcast Shane Hastie, Lead Editor for Culture & Methods spoke to Sam Bhagwat, co-founder and CEO of Mastra, about building and sustaining open source communities, the emerging discipline...
微软推出了开源工具包Evals for Agent Interop,旨在帮助开发者评估AI代理在数字工作场景中的互操作性。该工具包提供场景、数据集和评估框架,系统性地评估AI代理在企业工作流中的表现,尤其是在复杂任务和应用集成方面。开发者可进行定制化测试,以提升代理的性能和可靠性。
Hugging Face推出Community Evals功能,允许在Hub上创建基准数据集排行榜并自动收集评估结果。该系统基于Git基础设施,确保提交的透明性、可版本化和可重复性。用户可通过拉取请求提交评估结果,提升评估的一致性和可追溯性,目前处于测试阶段。
Most organisations have run an AI pilot. Far fewer have deployed an AI product. There is a fundamental gap between AI experimentation and production.
LangSmith推出Align Evals功能,帮助用户校准评估者以更好地匹配人类偏好。该功能允许用户迭代评估提示,比较人类评分与LLM生成的分数,并保存基线对比。用户可以通过选择评估标准、创建示例数据、手动评分和测试提示来逐步提升评估者的表现,未来还将推出分析工具和自动提示优化功能。
本研究提出了xai_evals框架,旨在评估机器学习模型后置局部解释方法的可靠性,以提高AI系统的可解释性和信任度。
完成下面两步后,将自动完成登录并继续当前操作。