小红花·文摘
  • 首页
  • 广场
  • 排行榜🏆
  • 直播
  • FAQ
Dify.AI
The rise of the agent runtime: The compute platform behind production agents

The fast pace of AI research means organizations now have a wide range of models to choose from that can The post The rise of the agent runtime: The compute platform behind production agents...

The rise of the agent runtime: The compute platform behind production agents

The New Stack
The New Stack · 2026-07-21T14:00:00Z
从Kaplan到Test-Time Compute:Scaling Law的真实演变与中文媒体的叙事偏差 - 张善友

Diogo指出Kaplan等人的Scaling Law存在技术缺陷,导致“参数越大越好”的错误结论。DeepMind的Chinchilla论文于2022年纠正了这一问题,提出了20:1的Token/参数最优比,促使行业转向更合理的训练策略。Meta的LLaMA系列和OpenAI的GPT-4等模型遵循这一原则,推动了AI训练方法的演变。虽然Diogo的文章有价值,但并非新发现,而是对已被纠正的技术偏差的回顾。

从Kaplan到Test-Time Compute:Scaling Law的真实演变与中文媒体的叙事偏差 - 张善友

张善友
张善友 · 2026-07-06T13:24:00Z

Oracle halved the Always Free Ampere A1 compute allowance from 4 OCPUs and 24 GB RAM to 2 OCPUs and 12 GB RAM with no public announcement. Support agents gave conflicting answers on whether PAYG...

Oracle Quietly Halves Free Tier Ampere A1 Compute Limits with No Public Announcement

InfoQ
InfoQ · 2026-07-03T10:13:00Z

Apple chose Google Cloud to run Private Cloud Compute outside its own data centers for the first time, using NVIDIA Blackwell GPUs, Intel TDX, and Google's Titan chip. Apple maintains an...

Apple Extends Private Cloud Compute to Google Cloud for the First Time

InfoQ
InfoQ · 2026-07-02T10:04:00Z

Roofline模型用于判断算子是计算密集型还是内存密集型,算术强度(AI)是关键指标,定义为浮点运算数与内存搬运字节数之比。通过双对数图,Roofline展示了性能与算术强度的关系,分为带宽限制区和算力限制区。优化策略包括提高算术强度,采用算子融合和tiling等方法,以减少内存访问并提升性能。

【GPU 算子工程】Roofline 模型:判断算子是 compute-bound 还是 memory-bound

土法炼钢兴趣小组的博客
土法炼钢兴趣小组的博客 · 2026-06-28T00:00:00Z

本文介绍了NVIDIA的两个调优工具:Nsight Systems和Nsight Compute。Nsight Systems用于分析程序的时间线,找出占用时间最多的kernel;而Nsight Compute则深入分析单个kernel的性能瓶颈。优化流程是先用Systems定位热点,再用Compute分析原因,如访存延迟或计算瓶颈。在没有完整工具链时,可使用CUDA event进行轻量测量。最终目标是提升GPU性能,优化计算效率。

【GPU 算子工程】Nsight 调优工作流:Compute 与 Systems 怎么读

土法炼钢兴趣小组的博客
土法炼钢兴趣小组的博客 · 2026-06-28T00:00:00Z

本文讨论了GPU kernel的调试与数值正确性,主要包括内存/竞态错误和数值错误两类。使用compute-sanitizer工具检查内存问题,并通过与高精度参考实现对比来验证数值正确性。强调浮点运算的非结合性可能导致结果微小差异,需用容差比较。总结了常见的kernel错误及调试方法,确保正确性是关键。

【GPU 算子工程】调试与数值正确性:compute-sanitizer 与对齐测试

土法炼钢兴趣小组的博客
土法炼钢兴趣小组的博客 · 2026-06-28T00:00:00Z

AI’s next breakthrough may not be a smarter model but a cheaper token.

Frontiers of compute: The technologies to reduce AI inference costs

McKinsey Insights & Publications
McKinsey Insights & Publications · 2026-06-25T00:00:00Z

By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha ChandrashekarAs a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned...

How Netflix Simplified Batch Compute with Kueue

Netflix TechBlog
Netflix TechBlog · 2026-06-22T21:35:01Z

PostgreSQL 14 unified query-id computation across all subsystems, but defaulting to always-on would tax every backend.

Christophe Pettus: All Your GUCs in a Row: compute_query_id

Planet PostgreSQL
Planet PostgreSQL · 2026-05-30T01:00:00Z
GitHub资深人士Brian Douglas创立Paper Compute以改善AI代理基础设施

Paper Compute公司专注于为AI代理构建基础设施,提供开源工具以增强生产环境中的可控性和可见性。其产品包括记录代理活动的Tapes和确保代理在受限环境中运行的StereOS。创始人Brian Douglas指出,随着AI代理的普及,理解其行为和决策变得至关重要,未来将专注于优化AI系统的管理和运行。

GitHub资深人士Brian Douglas创立Paper Compute以改善AI代理基础设施

The New Stack
The New Stack · 2026-04-27T12:00:00Z

Australia has an opportunity to become an Asia–Pacific AI hub, unlocking economic growth and productivity—though a number of constraints need to be addressed.

Australia’s AI moment: Building Asia–Pacific’s compute hub

McKinsey Insights & Publications
McKinsey Insights & Publications · 2026-04-15T00:00:00Z

Uber launches IngestionNext, a streaming-first data lake ingestion platform that reduces data latency from hours to minutes and cuts compute usage by 25%. Built on Kafka, Flink, and Apache Hudi,...

Uber Launches IngestionNext: Streaming-First Data Lake Cuts Latency and Compute by 25%

InfoQ
InfoQ · 2026-03-25T14:05:00Z
Why Start with Compute Governance, Not API Design

AI基础设施设计应从“后果”出发,通过分层结构将“意图”与“资源后果”结合。控制面表达意图,治理面限制后果,以确保成本和风险可控。未来,基础设施将随着上下文等新变量的引入而不断演进。

Why Start with Compute Governance, Not API Design

云原生
云原生 · 2026-01-18T05:45:13Z
Private AI Compute实现谷歌推理,采用硬件隔离和短暂数据设计

谷歌推出了Private AI Compute系统,利用Gemini云模型处理AI请求,同时保护用户数据隐私。该技术通过多层保护和AMD硬件的可信执行环境(TEE)加密和隔离内存,防止数据泄露。系统仅在满足用户查询时保留数据,确保安全性,并提升Pixel 10手机的智能建议功能,反映了行业对隐私保护AI系统的趋势。

Private AI Compute实现谷歌推理,采用硬件隔离和短暂数据设计

InfoQ
InfoQ · 2025-11-30T15:34:00Z

Power and cooling equipment are the backbones of data center infrastructure. Innovations and on-time supply of this technology will become increasingly relevant as the demand for data centers grows.

Beyond compute: Infrastructure that powers and cools AI data centers

McKinsey Insights & Publications
McKinsey Insights & Publications · 2025-10-29T00:00:00Z

We are thrilled to announce the newest member of our JupyterLite kernel ecosystem: Xeus-Octave. Xeus-Octave allows you to run GNU Octave code directly on your browser. GNU Octave is a free and...

GNU Octave Meets JupyterLite: Compute Anywhere, Anytime!

Jupyter Blog
Jupyter Blog · 2025-10-16T15:04:48Z
服务器渲染基准测试:Fluid Compute与Cloudflare Workers

独立开发者Theo Browne发布了Fluid compute与Cloudflare Workers的服务器端渲染性能基准测试,结果显示Fluid compute在计算密集型任务中比Cloudflare Workers快1.2到5倍,且响应时间更一致。Fluid compute支持标准Node.js和Python,并在同一云区域部署以减少网络延迟。

服务器渲染基准测试:Fluid Compute与Cloudflare Workers

Vercel News
Vercel News · 2025-10-09T13:00:00Z
Modular:SF Compute与Modular合作革新AI推理经济

Modular与SF Compute联合推出大型推理批处理API,旨在降低AI生态系统中的计算成本。该API支持20多种先进模型,提供高达80%的成本节约,优化AI推理的经济结构,推动AI创新。

Modular:SF Compute与Modular合作革新AI推理经济

Modular Blog
Modular Blog · 2025-07-31T00:00:00Z
Fluid Compute的主动CPU定价降低了费用

Vercel Functions在Fluid Compute上实施了主动CPU定价,仅在CPU使用时收费,从而降低了空闲成本。定价基于主动CPU时间、配置内存和调用次数,适用于所有Hobby、Pro和新Enterprise团队。

Fluid Compute的主动CPU定价降低了费用

Vercel News
Vercel News · 2025-06-25T13:00:00Z
  • <<
  • <
  • 1 (current)
  • 2
  • 3
  • >
  • >>
👤 个人中心
在公众号发送验证码完成验证
登录验证
在本设备完成一次验证即可继续使用

完成下面两步后,将自动完成登录并继续当前操作。

1 关注公众号
小红花技术领袖公众号二维码
小红花技术领袖
如果当前 App 无法识别二维码,请在微信搜索并关注该公众号
2 发送验证码
在公众号对话中发送下面 4 位验证码
友情链接: MOGE.AI 九胧科技 模力方舟 Gitee AI 菜鸟教程 Remio.AI DeekSeek连连 53AI 神龙海外代理IP IPIPGO全球代理IP 东波哥的博客 匡优考试在线考试系统 开源服务指南 蓝莺IM Solo 独立开发者社区 AI酷站导航 极客Fun 我爱水煮鱼 周报生成器 He3.app 简单简历 白鲸出海 T沙龙 职友集 TechParty 蟒周刊 Best AI Music Generator

小红花技术领袖俱乐部
小红花·文摘:汇聚分发优质内容
小红花技术领袖俱乐部
Copyright © 2021-
粤ICP备2022094092号-1
公众号 小红花技术领袖俱乐部公众号二维码
视频号 小红花技术领袖俱乐部视频号二维码