小红花·文摘
  • 首页
  • AI Tokens🪙
  • 排行榜🏆
  • 直播
  • FAQ

Meta has detailed MTIA 300, its first in-house accelerator optimized for training ranking and recommendation models. By Matt Foster

Meta Expands its Custom Silicon Strategy from Compute into Networking

InfoQ InfoQ · 2026-08-28T07:43:00Z

Amazon Web Services has extended Amazon Bedrock AgentCore with runtime instances, a new compute option that gives AI agents persistent infrastructure purpose-built for complex long-running...

Multi Agent Collaboration Gets Persistent Compute in Bedrock AgentCore

InfoQ InfoQ · 2026-08-19T08:00:00Z

Arun Joseph shares real-world insights on scaling enterprise agentic platforms like Deutsche Telekom’s LMOS. He discusses bridging organizational fault lines, replacing tool sprawl with core...

Presentation: Architecting AI Systems for the Messy Reality of Enterprises: Why Agentic Compute is the Missing Layer

InfoQ InfoQ · 2026-08-03T08:08:00Z
从Kaplan到Test-Time Compute:Scaling Law的真实演变与中文媒体的叙事偏差 - 张善友

Diogo指出Kaplan等人的Scaling Law存在技术缺陷,导致“参数越大越好”的错误结论。DeepMind的Chinchilla论文于2022年纠正了这一问题,提出了20:1的Token/参数最优比,促使行业转向更合理的训练策略。Meta的LLaMA系列和OpenAI的GPT-4等模型遵循这一原则,推动了AI训练方法的演变。虽然Diogo的文章有价值,但并非新发现,而是对已被纠正的技术偏差的回顾。

从Kaplan到Test-Time Compute:Scaling Law的真实演变与中文媒体的叙事偏差 - 张善友

张善友 张善友 · 2026-07-06T13:24:00Z

Oracle halved the Always Free Ampere A1 compute allowance from 4 OCPUs and 24 GB RAM to 2 OCPUs and 12 GB RAM with no public announcement. Support agents gave conflicting answers on whether PAYG...

Oracle Quietly Halves Free Tier Ampere A1 Compute Limits with No Public Announcement

InfoQ InfoQ · 2026-07-03T10:13:00Z

Apple chose Google Cloud to run Private Cloud Compute outside its own data centers for the first time, using NVIDIA Blackwell GPUs, Intel TDX, and Google's Titan chip. Apple maintains an...

Apple Extends Private Cloud Compute to Google Cloud for the First Time

InfoQ InfoQ · 2026-07-02T10:04:00Z

Roofline模型用于判断算子是计算密集型还是内存密集型,算术强度(AI)是关键指标,定义为浮点运算数与内存搬运字节数之比。通过双对数图,Roofline展示了性能与算术强度的关系,分为带宽限制区和算力限制区。优化策略包括提高算术强度,采用算子融合和tiling等方法,以减少内存访问并提升性能。

【GPU 算子工程】Roofline 模型:判断算子是 compute-bound 还是 memory-bound

土法炼钢兴趣小组的博客 土法炼钢兴趣小组的博客 · 2026-06-28T00:00:00Z

本文介绍了NVIDIA的两个调优工具:Nsight Systems和Nsight Compute。Nsight Systems用于分析程序的时间线,找出占用时间最多的kernel;而Nsight Compute则深入分析单个kernel的性能瓶颈。优化流程是先用Systems定位热点,再用Compute分析原因,如访存延迟或计算瓶颈。在没有完整工具链时,可使用CUDA event进行轻量测量。最终目标是提升GPU性能,优化计算效率。

【GPU 算子工程】Nsight 调优工作流:Compute 与 Systems 怎么读

土法炼钢兴趣小组的博客 土法炼钢兴趣小组的博客 · 2026-06-28T00:00:00Z

本文讨论了GPU kernel的调试与数值正确性,主要包括内存/竞态错误和数值错误两类。使用compute-sanitizer工具检查内存问题,并通过与高精度参考实现对比来验证数值正确性。强调浮点运算的非结合性可能导致结果微小差异,需用容差比较。总结了常见的kernel错误及调试方法,确保正确性是关键。

【GPU 算子工程】调试与数值正确性:compute-sanitizer 与对齐测试

土法炼钢兴趣小组的博客 土法炼钢兴趣小组的博客 · 2026-06-28T00:00:00Z

AI’s next breakthrough may not be a smarter model but a cheaper token.

Frontiers of compute: The technologies to reduce AI inference costs

McKinsey Insights & Publications McKinsey Insights & Publications · 2026-06-25T00:00:00Z

By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha ChandrashekarAs a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned...

How Netflix Simplified Batch Compute with Kueue

Netflix TechBlog Netflix TechBlog · 2026-06-22T21:35:01Z

PostgreSQL 14 unified query-id computation across all subsystems, but defaulting to always-on would tax every backend.

Christophe Pettus: All Your GUCs in a Row: compute_query_id

Planet PostgreSQL Planet PostgreSQL · 2026-05-30T01:00:00Z
GitHub资深人士Brian Douglas创立Paper Compute以改善AI代理基础设施

Paper Compute公司专注于为AI代理构建基础设施,提供开源工具以增强生产环境中的可控性和可见性。其产品包括记录代理活动的Tapes和确保代理在受限环境中运行的StereOS。创始人Brian Douglas指出,随着AI代理的普及,理解其行为和决策变得至关重要,未来将专注于优化AI系统的管理和运行。

GitHub资深人士Brian Douglas创立Paper Compute以改善AI代理基础设施

The New Stack The New Stack · 2026-04-27T12:00:00Z

Australia has an opportunity to become an Asia–Pacific AI hub, unlocking economic growth and productivity—though a number of constraints need to be addressed.

Australia’s AI moment: Building Asia–Pacific’s compute hub

McKinsey Insights & Publications McKinsey Insights & Publications · 2026-04-15T00:00:00Z

Uber launches IngestionNext, a streaming-first data lake ingestion platform that reduces data latency from hours to minutes and cuts compute usage by 25%. Built on Kafka, Flink, and Apache Hudi,...

Uber Launches IngestionNext: Streaming-First Data Lake Cuts Latency and Compute by 25%

InfoQ InfoQ · 2026-03-25T14:05:00Z
The 2026 Middle East Crisis, the Helium Cutoff, and the Chain Implosion of the AI Compute Bubble

2026年3月,中东冲突导致全球氦气供应链崩溃,氦气价格飙升,严重影响半导体制造和人工智能产业,尤其是东亚地区的芯片生产面临挑战,能源成本上升,AI基础设施建设的金融风险加剧,可能引发经济危机。

The 2026 Middle East Crisis, the Helium Cutoff, and the Chain Implosion of the AI Compute Bubble

Lv. MAX Lv. MAX · 2026-03-21T00:00:00Z
Why Start with Compute Governance, Not API Design

AI基础设施设计应从“后果”出发,通过分层结构将“意图”与“资源后果”结合。控制面表达意图,治理面限制后果,以确保成本和风险可控。未来,基础设施将随着上下文等新变量的引入而不断演进。

Why Start with Compute Governance, Not API Design

云原生 云原生 · 2026-01-18T05:45:13Z
Private AI Compute实现谷歌推理,采用硬件隔离和短暂数据设计

谷歌推出了Private AI Compute系统,利用Gemini云模型处理AI请求,同时保护用户数据隐私。该技术通过多层保护和AMD硬件的可信执行环境(TEE)加密和隔离内存,防止数据泄露。系统仅在满足用户查询时保留数据,确保安全性,并提升Pixel 10手机的智能建议功能,反映了行业对隐私保护AI系统的趋势。

Private AI Compute实现谷歌推理,采用硬件隔离和短暂数据设计

InfoQ InfoQ · 2025-11-30T15:34:00Z

Power and cooling equipment are the backbones of data center infrastructure. Innovations and on-time supply of this technology will become increasingly relevant as the demand for data centers grows.

Beyond compute: Infrastructure that powers and cools AI data centers

McKinsey Insights & Publications McKinsey Insights & Publications · 2025-10-29T00:00:00Z

We are thrilled to announce the newest member of our JupyterLite kernel ecosystem: Xeus-Octave. Xeus-Octave allows you to run GNU Octave code directly on your browser. GNU Octave is a free and...

GNU Octave Meets JupyterLite: Compute Anywhere, Anytime!

Jupyter Blog Jupyter Blog · 2025-10-16T15:04:48Z
  • <<
  • <
  • 1 (current)
  • 2
  • 3
  • >
  • >>
👤 个人中心
在公众号发送验证码完成验证
登录验证
在本设备完成一次验证即可继续使用

完成下面两步后,将自动完成登录并继续当前操作。

1 关注公众号
小红花技术领袖公众号二维码
小红花技术领袖
如果当前 App 无法识别二维码,请在微信搜索并关注该公众号
2 发送验证码
在公众号对话中发送下面 4 位验证码
友情链接: MOGE.AI 九胧科技 1tok 菜鸟教程 Remio.AI DeekSeek连连 53AI 神龙海外代理IP IPIPGO全球代理IP 东波哥的博客 匡优考试在线考试系统 开源服务指南 蓝莺IM Solo 独立开发者社区 AI酷站导航 极客Fun 我爱水煮鱼 周报生成器 He3.app 简单简历 白鲸出海 T沙龙 职友集 TechParty 蟒周刊 Best AI Music Generator 模力方舟 Gitee AI

小红花技术领袖俱乐部
小红花·文摘:汇聚分发优质内容
小红花技术领袖俱乐部
Copyright © 2021-
粤ICP备2022094092号-1
公众号 小红花技术领袖俱乐部公众号二维码
视频号 小红花技术领袖俱乐部视频号二维码