➡️
继续阅读
-
GKE Pod Snapshots Cut Model Load Times, and Move the Work to Snapshot Lifecycle Management
Google has published benchmarks for GKE Pod snapshots, reporting up to 89% lo...
-
Fast & Efficient LLM Inference with vLLM-II
本文介绍vLLM推理优化:通过连续批处理让GPU持续忙碌,用PagedAttention分页管理KV缓存减少碎片,并利用前缀缓存跳过重复计算。还介绍Gui...
-
Presentation: Adaptive Recommenders in the Real World: Inference, Evals, and System Design
Mallika Rao explains that the true complexity of adaptive recommendation syst...
-
Fast & Efficient LLM Inference with vLLM-I
文章介绍vLLM推理优化课程:先讲解KVCache原理,32K上下文约占16GB显存;再介绍LLM Compressor通过量化与稀疏化降低显存搬运,如W...
-
【Rust日报】2026-09-28 PolyXOR128:带 Lean 证明的高速 128 位通用哈希
本期介绍四个 Rust 项目:PolyXOR128 是带 Lean 形式化证明的 128 位通用哈希,碰撞概率上界为 (n/4096+3)/2^128,吞...
-
发现频道:10款大家发现的好评软件[2026年第39期]
小众软件论坛发现频道公布近10日热门排行榜,上榜项目包括:Piko(开源PikPak客户端)、Mimi(实时翻译字幕)、Bref(多端播放器)、Arrow...