ZhangYvJing's

Daily Brief

← July 20, 2026 July 21, 2026 · Tuesday July 22, 2026 →
00

Film / Book Chapter

Still Walking
2008 / Hirokazu Kore-eda

Still Walking (2008) · Hirokazu Kore-eda

今天适合看《Still Walking》,因为它更像一次生活和判断方式的校准,能把注意力从持续输入里稍微抽出来,重新放回你真正想怎样生活和做事上。

Designing Data-Intensive Applications
Martin Kleppmann

Designing Data-Intensive Applications · Martin Kleppmann

Chapter 1: Reliable, Scalable, and Maintainable Applications

A useful morning chapter when system design starts feeling abstract: it turns reliability, scalability, and maintainability back into concrete product constraints.

01

Insight

今天的讨论在「开放式 AI 架构」与「算力/安全瓶颈」这两条轴线上显得尤为突出。中国面向开放权重的策略已被多个行业案例证明领先,而美国的封闭策略在众多实验和投资中显得滞后;这场“权重开放”与“封闭商业化”对峙的焦点,直接映射到企业价值链 השירות。与此对应的安全失误与算力瓶颈则在 大发彩票快三许多引用的资料中交错呈现: Hacker News 上的波兰房地产数据库被擦除与为负载预测做出的行为感知的注意力模型,都说明数据完整性与高并发并发压测仍是技术 доступный, yet Y‑Combinator 与 Together AI 的 GPU 集群合作和 Snyk、Keycard 这类安全演讲再次把周边技术的硬件制约与漏洞通用性拉到聚焦点;短视频·用户锋面上的 Hyprland 切换 Lua 与 Kimi Work 自动化、Jelly UI 的软物理交互,实则在 “易用 vs 可维护” 的疲软层面产生矛盾。最后,arXiv 论文对主动观察(ActiveVision)、物理增强 RL、热力学计算等路线的探索告诉我们,模型实力并非衡量 AI 成熟度的唯一标准,进一步证明不应被迭代评测噪声所掩盖。今天的焦点,若要抓住未来,便是把开放、算力与安全三合 بدر切地结合,从而推动 AI 站在应用与基础研究双轨并进的前沿。Still Walking (2008)。
03

Hacker News

03
Kimi Work
04
Marine 0.55正式切换配置语言到 Lua,旧版 hyprlang 配置仍维持有限兼容。Lua 允许直接在配置中编写布局 API,可设置全局、工作区或单显示器布局;同时滚动手势、触摸板chst、每输出 ICC 配置及 FP16 渲染被优化,以提升色彩精准度。迁移至 Lua 需学习其语法并重写配置文件,但可减少手动脚本编写,改变维护成本与开发周期。
05
Jelly UI 新发布了一个无任何依赖的 Web Components 库 positivas 通过软体物理为标准表单控件赋予触感化外achd。 该库将真实表单控件与软体物理引擎相结合,在单个组件内即装配暗色模式、右至左文本支持与 WCAG AA 级颜色方案。 前端工程师可直接引用 <script type="module" src="https://jelly-ui.com/package.js"</script,省去第三方依赖、构建流程和合规开发的重复工作,从而降低维护成本并提升可访问性与交互质量。
07
游戏引擎使用屏幕空间环境遮蔽(SSAO)技术,却因近似方法与全照明加深系数过高导致角落不真实暗化。研究指出近似误差、阴影映射错误以及对所有光照@Builder B整合?的错误应用是主要原因,导致模拟角落暗化与物理光照呈显著偏差。此问题Consult导致游戏开发者在调试降噪、质量与性能平衡时面临更大成本与视觉风险,迫使他们重新评估遮蔽实现与渲染管线规则。
10
Bloomy推出AI主导的K‑12熟练学习平台。通过诊断学生技能差距、个性化学习路径与Socratic AI导师,模仿Bloom 2‑sigma的一对一辅导,并配备适应性练习与90%掌握门槛。此模式让传统课堂、特许、微型及家庭学习以更低成本获得精准辅导,教师与家长可实时查看进度与关键技能。
04

YouTube

02
Your agent passes offline evals at 90%. You ship. Production immediately finds failure modes your eval never saw. Sound familiar? The culprit is almost always the same: the "customer" in your offline eval is an off-the-shelf LLM that sounds nothing like your real users, and your synthetic test set doesn't capture how messy, angry, or off-topic real conversations get. Your eval was too easy. At Lyft, our customer-care agents resolve roughly a third of all customer issues — millions of conversations a month. To trust them at that scale, we built an adversarial user simulator: a fine-tuned LLM
agent, ai_frontier, ai_product, engineering
03
Performance issues silently pile up in mature codebases. Teams know things could be faster, but can never justify pausing feature work to investigate. You have to put engineers on it just to find out if there's something worth fixing, and the effort is completely unpredictable: it could take an hour or three weeks. In this talk, we'll walk through a real case study of adding runtime intelligence to coding agents to enable continuous performance optimization in production. We'll cover the pain that led us here, the technical approach (agents analyzing real production context to surface high-RO
agent, ai_product, engineering, security
04
We built a demo agent to show customers how to connect agents to their tools. A simple chat assistant — Gmail, Calendar, a handful of connectors. It ran on a 15-minute schedule. And every 15 minutes, our production database strained. Latency crept up and alerts fired. Then settled. Then, it fired again. It took us a while to find it. One line - a "last seen" timestamp updating on every tool call. Written for a human who logs in once. Our agent was calling it sixty times a second. We had built infrastructure to show customers how to connect agents to their tools. We hadn't noticed we'd built
agent, ai_frontier, ai_product, engineering
05
This talk examines the engineering challenges of building foundation models for single-cell biology from a non-biologist’s perspective. Speakers: - Akram Baharlouei (Altos Labs): Machine learning engineer at Altos Labs working on foundation models for biology. Previously at Meta AI and Qualcomm. LinkedIn: https://linkedin.com/in/akram-baharlouei-61784421
agent, ai_product, engineering
08
Twenty years ago Aaron Stanley arrived at an emergency evidence collection for an SEC investigation and realized he had forgotten the dongle that licensed his forensic software. Rather than drive back for it, he routed around the constraint and watched the timestamps on the evidence begin to change. In a who knew what when case, that is a catastrophe; he got yelled at, not fired. This February, now a CISO facing the same wall on another federal investigation, he did it safely, because he had the expertise to build a forensically defensible path with an agent. His point: the agents we build tod
agent, ai_product, engineering
09
Building software with AI almost feels like a cheat code: you ship what you were working on and watch it spark joy in real users. The catch, and the reason Randall Degges is opening the World's Fair's first Security Track, is that three things still stand in the way of doing that at scale. AI writes insecure code just like humans do, autonomous agents in production can go off the rails while you sleep, and access to frontier models keeps getting pulled out from under you for what amounts to geopolitics. It all reduces to one unsolved problem: using AI fearlessly and having it be secure by defa
agent, ai_frontier, ai_product, engineering, security
10
An incident agent on the night shift reads a ticket: the billing database is broken, payments failing. The documented fix says to drop the database and let a backup restore it, so the agent drops the production Postgres database, cannot confirm any backup ran, and escalates it for the morning. This has happened to real companies. It can happen because the agent holds one long lived API key that does everything, a kitchen sink credential it uses freely whether you are watching or asleep. Kim Maida's fix is not a new invention but an old OAuth spec, token exchange, wired into the agent's execut
agent, ai_product, engineering
11
Ask the latest frontier models, the ones not even public yet, to find the same vulnerability five times, and only half of those runs catch it. Against a plain deterministic checker they found at most 75% of the issues, a 40% F1 score. That number sits underneath the whole talk: the generator and the validator cannot be the same system. Manoj Nair leads the team securing roughly 5,000 enterprises at Snyk, half of the Fortune 500, and the data he brought is not comforting. Across 4,800 customers, security backlog grew 108% quarter over quarter, because agents writing code faster are also manufac
agent, ai_frontier, ai_product, engineering, security
12
YC and Together AI are partnering to bring the first dedicated YC GPU cluster online, giving YC startups easier access to the compute they need to build and scale. In this episode of Founder Firesides, YC's Ankit Gupta and Together AI co-founder and CEO Vipul Ved Prakash dig into why compute has become one of the biggest bottlenecks for modern AI companies. They’ll discuss how Together AI is helping more than 8,000 customers, from early-stage research teams to companies like Cursor, Cognition, and ElevenLabs, train, fine-tune, & run inference on AI models, and why flexible access to GPUs is be
ai_frontier, ai_product, market, product, startup
07

Papers

01
An Exam for Active Observers
本文关注多模态大型语言模型是否真正具备“主动观测”的能力。作者设计了 ActiveVision 基准,包含 17 个需要循环视觉感知、而非一次性描述的任务,评估模型在动态观察中的表现。实验表明 GPT‑5.5 与 Claude Fable‑5 在大部分任务几乎无能,只有人类水平约10倍,说明现有模型缺乏闭环感知urdu‑reasoning 循环,对 Agent 方向的安全与决策非常关键逆向思考。
02
住宅短期负荷预测因需求异质、时变与行为驱动而复杂。本文提出行为条件的注意力神经过程(Attentive Neural Process),通过在上下文中推断离散行为标签并用于解码器条件化,同时用连续潜在变量捕捉共享不确定性;训练时采用聚类弱监督,测试仅用上下文推断。实验显示在有限上下文下,MAE、CRPS相较基线减少约8%,且在所有预测距离上保持更低RMSE。该方法在不需要标签的前提下实现单模型、可量化不确定度的预测,适合 Agent/AI 能耗管理系统快速响应与状态预测。
03
本文探讨 Muon 优化器在稀疏奖励“Agentic”强化学习中的性能, ҳама на ALFWorld 搭配 Qwen2.5‑0.5B‑Instruct 进行 GiGPO 与 GraphGPO 对比。只将 Muon 作用于隐藏权重矩阵,GiGPO 的最终验证成功率从 0.290 提升至 0.546,而 GraphG associated 在 1e‑5 学习率下提前 30/60 次更新即达 0.5/0.75 成功率。作者指出 Muon 与优势估计器、学习率的交互决定效果,提示在后期 RL 微调阶段应细致调参。对于关注优化square器与 RL 效726 搭配的工程师,这提供了可直接尝试的实证依据和探索方向。
05
解决:强化学习在高维动态系统上样本效率低、探索困难导致难以实现实时最优控制。方法:提出 Physics‑Enhanced Reinforcement Learning (PEARL),采用 actor‑adjoint Distrib 已结合自动微分,利用已知物理动力学计算短期策略梯度,并通过神经网络估算未来奖励梯度,显著减少环境交互。价值:PEARL 在复杂参数化流场导航中已显示比主流 RL 更高的样本效率、泛FIND 性以及可直接扩展到高维状态和动作空间,对需要在稀疏传感器下实现精确自适应控制的 Agent/AI 工程尤为吸引。
06
论文解决点云匹配中聚类结构的挑战,提出不追求逐点对应而是区域对齐。核心是用相似图构造的二次拉普拉斯正则化约束最优输运(LapOT),保证匹配尊重两边的聚类,再结合 RSC 获得一致可解释的分区。对机器人、地图对齐等工程应用,可提升匹配鲁棒性与可解释性。
07
解决AI 推理与训练的能耗与延迟爆炸,用热力学原理把随机模拟写进硬件。作者提出一种基于 Langevin 动力学的能量计算栈,利用可调能量势在超导模拟电路中直接产生并采样能量基模型,随后在概率图模型框架下训练常见的机器学习模型。如此,硬件天然的低功耗随机行为可为 Agent/AI 系统提供高速、低延迟的概率推理与优化能力。
physics.app-ph