Meta has detailed MTIA 300, its first in-house accelerator optimized for training ranking and recommendation models. By Matt Foster
Amazon Web Services has extended Amazon Bedrock AgentCore with runtime instances, a new compute option that gives AI agents persistent infrastructure purpose-built for complex long-running...
Arun Joseph shares real-world insights on scaling enterprise agentic platforms like Deutsche Telekom’s LMOS. He discusses bridging organizational fault lines, replacing tool sprawl with core...
Diogo指出Kaplan等人的Scaling Law存在技术缺陷,导致“参数越大越好”的错误结论。DeepMind的Chinchilla论文于2022年纠正了这一问题,提出了20:1的Token/参数最优比,促使行业转向更合理的训练策略。Meta的LLaMA系列和OpenAI的GPT-4等模型遵循这一原则,推动了AI训练方法的演变。虽然Diogo的文章有价值,但并非新发现,而是对已被纠正的技术偏差的回顾。
Oracle halved the Always Free Ampere A1 compute allowance from 4 OCPUs and 24 GB RAM to 2 OCPUs and 12 GB RAM with no public announcement. Support agents gave conflicting answers on whether PAYG...
Apple chose Google Cloud to run Private Cloud Compute outside its own data centers for the first time, using NVIDIA Blackwell GPUs, Intel TDX, and Google's Titan chip. Apple maintains an...
Roofline模型用于判断算子是计算密集型还是内存密集型,算术强度(AI)是关键指标,定义为浮点运算数与内存搬运字节数之比。通过双对数图,Roofline展示了性能与算术强度的关系,分为带宽限制区和算力限制区。优化策略包括提高算术强度,采用算子融合和tiling等方法,以减少内存访问并提升性能。
本文介绍了NVIDIA的两个调优工具:Nsight Systems和Nsight Compute。Nsight Systems用于分析程序的时间线,找出占用时间最多的kernel;而Nsight Compute则深入分析单个kernel的性能瓶颈。优化流程是先用Systems定位热点,再用Compute分析原因,如访存延迟或计算瓶颈。在没有完整工具链时,可使用CUDA event进行轻量测量。最终目标是提升GPU性能,优化计算效率。
本文讨论了GPU kernel的调试与数值正确性,主要包括内存/竞态错误和数值错误两类。使用compute-sanitizer工具检查内存问题,并通过与高精度参考实现对比来验证数值正确性。强调浮点运算的非结合性可能导致结果微小差异,需用容差比较。总结了常见的kernel错误及调试方法,确保正确性是关键。
AI’s next breakthrough may not be a smarter model but a cheaper token.
By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha ChandrashekarAs a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned...
PostgreSQL 14 unified query-id computation across all subsystems, but defaulting to always-on would tax every backend.
Paper Compute公司专注于为AI代理构建基础设施,提供开源工具以增强生产环境中的可控性和可见性。其产品包括记录代理活动的Tapes和确保代理在受限环境中运行的StereOS。创始人Brian Douglas指出,随着AI代理的普及,理解其行为和决策变得至关重要,未来将专注于优化AI系统的管理和运行。
Australia has an opportunity to become an Asia–Pacific AI hub, unlocking economic growth and productivity—though a number of constraints need to be addressed.
Uber launches IngestionNext, a streaming-first data lake ingestion platform that reduces data latency from hours to minutes and cuts compute usage by 25%. Built on Kafka, Flink, and Apache Hudi,...
2026年3月,中东冲突导致全球氦气供应链崩溃,氦气价格飙升,严重影响半导体制造和人工智能产业,尤其是东亚地区的芯片生产面临挑战,能源成本上升,AI基础设施建设的金融风险加剧,可能引发经济危机。
AI基础设施设计应从“后果”出发,通过分层结构将“意图”与“资源后果”结合。控制面表达意图,治理面限制后果,以确保成本和风险可控。未来,基础设施将随着上下文等新变量的引入而不断演进。
谷歌推出了Private AI Compute系统,利用Gemini云模型处理AI请求,同时保护用户数据隐私。该技术通过多层保护和AMD硬件的可信执行环境(TEE)加密和隔离内存,防止数据泄露。系统仅在满足用户查询时保留数据,确保安全性,并提升Pixel 10手机的智能建议功能,反映了行业对隐私保护AI系统的趋势。
Power and cooling equipment are the backbones of data center infrastructure. Innovations and on-time supply of this technology will become increasingly relevant as the demand for data centers grows.
We are thrilled to announce the newest member of our JupyterLite kernel ecosystem: Xeus-Octave. Xeus-Octave allows you to run GNU Octave code directly on your browser. GNU Octave is a free and...
完成下面两步后,将自动完成登录并继续当前操作。