ZhangYvJing's

Daily Brief

← September 09, 2026 September 10, 2026 · Thursday September 11, 2026 →
00

Film / Book Chapter

Still Walking
2008 / Hirokazu Kore-eda

Still Walking (2008) · Hirokazu Kore-eda

今天适合看《Still Walking》,因为它更像一次生活和判断方式的校准,能把注意力从持续输入里稍微抽出来,重新放回你真正想怎样生活和做事上。

The Pragmatic Programmer
David Thomas / Andrew Hunt

The Pragmatic Programmer · David Thomas / Andrew Hunt

Chapter 1: A Pragmatic Philosophy

A compact reset on ownership, taste, entropy, and being the kind of engineer whose work keeps improving after the first pass.

01

Insight

今天的输入更像几股不同语气的材料同时挤在一起:社区链接在暴露工程和产品环境里的真实焦点,长视频在把这些焦点放回更完整的语境里,研究材料则提醒人热度和可落地性并不总是同一件事。如果先不急着做结论,至少可以把这几条线索放在一起看:Hacker News 的 iPhone Duo;Hacker News 的 AirPods 5;Hacker News 的 Shopify acquires Tailwind;Hacker News 的 What do Visa and Mastercard do? An intro to card networks;Hacker News 的 Apple Watch Series 12;Hacker News 的 iPhone 18 Pro and iPhone 18 Pro Max。真正值得注意的不是单条内容本身,而是它们共同指向了什么、彼此漏掉了什么。
03

Hacker News

01
iPhone Duo
苹果公司推出首款折叠款iPhone,名为iPhone Duo。当展开时,它配备了7.6英寸超 Retina XDR 显示屏,比iPhone 18 Pro Max 大面积50%,内外屏比例一致,采用10层超薄材料并隐藏面部摄像头,实现无缝视觉体验。折叠设计支持多种摆放姿态,包括横向沉浸娱乐、竖向浏览和静置挂钟等场景,同时集成A20 Pro芯片、双电池系统及Siri AI助手等新功能。该设计延续IP68防水等级,采用Grade 5钛合金框架和陶瓷护板,旨在提升用户交互灵活性与设备耐用性。
02
AirPods 5
苹果公司于2026年9月9日发布AirPods 5,标志着散装型耳机首次实现业界领先的主动降噪技术。新款AirPods 5采用全新多端口声学结构和下一代自适应-equalizer,声学设计使外部噪声降低最高达50%,同时提升Transparency模式的自然度。与Siri AI和iPhone的深度集成支持无接触语音交互和实时翻译功能,用户通过头部动作即可响应指令。价格方面,基础款售价129美元,支持无线充电款为149美元,并将于9月18日开启零售店发售。此次发布延续了苹果对环保材料的持续投入,整机含40%回收材料,其中充电盒使用70%回收塑料。
03
Shopify acquires Tailwind
Shopify宣布收购CSS框架Tailwind,此举将为其提供稳定的长期维护平台。作者回顾九年发展,指出Tailwind如今已被安装超过11000万次,并且是多家知名企业 stylesheets 的核心选择。此次收购驱动因素是希望将框架开发融入实际产品中,以解决真实商业挑战,并借助Shopify在agentic commerce领域的前沿探索。团队强调开源项目将继续以MIT许可证发布,现有客户可继续使用Tailwind Plus等商业产品,但将停止新增客户招募,专注于框架本身的改进。
04
Visa 和 Mastercard 不发卡、不办银行,而是作为卡网络在发卡行、收单行和商家之间路由交易消息并完成结算。维护数据中心,运行电信网络传输授权和清算信息,通过银行网络实现净额结算并处理跨境货币兑换。网络通过设定互换费和网络服务费来分配收益,影响商户手续费、发卡行风险敞口和结算资金流动性需求。
05
Apple Watch Series 12
Apple Watch Series 12 首次搭载全新健康感知系统,实现更高频率的心率与心率变异性测量。该系统依赖更新的光学和电传感器以及 S11 芯片,使数据采集更精准且连续。健康管理者、运动爱好者以及需要随时监测生理状态的用户可通过手腕获得即时反馈,从而调整活动强度和恢复策略。
06
iPhone 18 Pro and iPhone 18 Pro Max
苹果发布iPhone 18 Pro和iPhone 18 Pro Max,搭载A20 Pro芯片和变量光圈主相机。新机首次引入变量光圈技术,由六片激光切割的光阑控制,提升低光拍照效果并支持手动调节,同时提供更锐利的4800万像素主摄和升级的计算成像管道。A20 Pro采用2纳米工艺,CPU算力提升50%,AI处理能力翻倍,配合全新热管理系统和气凝体,实现最高40%的持续性能提升,iPhone 18 Pro Max大幅优化续航,视频播放可达45小时。机型运行iOS 27,内置Siri AI和Apple Intelligence,支持引用图像认证功能。这些升级直接影响摄影师、内容创作者及科研工作者的拍摄、验证流程,降低硬件成本同时提升数据真实性风险。
07
No Man's Sky Cosmos
《No Man's Sky》继续演进,星际系统加入可重复访问的固定地点,包含遗迹、小怪船、彗星粉尘等深空点落。玩家可拆废弃太空殖民舰获取稀有材料,同时谨慎处理易引发连锁反应的易燃殆,必须掌握重力线圈技巧穿越病毒化遗址。全新星系地图、六自由度太空漫步以及可建造轨道基地为冒险者开辟新选项,同时引入太空站总监编织体系,玩家可组建势力联盟、管理派生设施。图形技术升级支持XeSS/DLSS 4.5及VR 凝视渲染,提升视觉真实性。十周年纪念事件“Expedition Twenty-Three”同步上线,庆祝十年演变。
08
Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
对Qwen 3.8的推理 prefills 实验进行升级版测试,将GPT-5.5 Pro设为教师模型。该方法为每个问题生成两组模型输出:一种是常规输出,另一种是将GPT-5.5 Pro前1%推理内容插入目标模型推理通道后再生成。最终衡量标准是目标模型作答前100个token中与GPT-5.5 Pro可见答案的重合度。结果显示Qwen 3.8对GPT-5.5 Pro的重合度提升18.18个百分点,尤其在私有数据集的合成谜题上表现显著,这表明Qwen可能从GPT-5.5 Pro或其近亲模型学习,而非从Opus 4.8获得启发。对于依赖open源大模型开发的人工智能研究人员和应用构建者,这一发现意味着可通过导入更先进模型的推理片段来显著提升目标模型的输出质量和一致性。
09
OpenAI刚发布的GPT-6 Astra模型在3D渲染、动画和数学编码任务上显著领先,ARC-AGI-3基准测试达99.9%。该模型采用“looped transformer”架构,有传闻其推理链条被隐藏。Astra通过 Reinforcement Learning从数千台Mac Minis获取计算机使用训练,其工作流程是:模型根据提示执行任务,观察界面截图,预测鼠标键盘动作,由包装层执行后反馈结果。这种机器人般的GUI操控能力使其在图形界面任务中脱颖而出,为处理税务、日常办公等非技术工作流程提供了新的自动化路径。
10
Growing proof that autonomous cars save lives
研究显示,自动驾驶汽车碰撞和人员伤亡率明显低于人工驾驶,Waymo robotaxi 在可比城市中致死或重伤事故下降 92%,AEB 等高级辅助系统可使行人碰撞减少 27%。这些结论来自 IIHS 对人工事故漏报的修正以及对仅在良好天气运行车里程的比较。将推动统一联邦安全报告标准,影响保险理赔、车企功能设计及道路使用者的风险评估。
04

YouTube

01
A single token of KV cache on Mistral 7B costs 131 KB. Multiply that by 16,000 tokens of context and 80 concurrent users and the cache alone wants 42 GB of GPU memory, which is why requests start failing on a 24 GB card. Harshul Jain, a senior software engineer at Audible, and Tanmay Sah, an independent AI researcher, spend this workshop building that number up from first principles. They open on three symptoms every team hits: memory that climbs with context length, time to first token that degrades as prompts grow, and throughput that collapses because a naive server answers requests one aft
agent, ai_frontier, ai_product, engineering
03
Laurie Voss reran a year old benchmark and the models walked straight through its ceiling without noticing it was there. IFScale asks a model to write a business report containing a list of exact words, then counts how many actually appear. A year ago frontier models started dropping instructions somewhere around 200 to 300, which is a hard limit on how much you can put in a skills file. Voss, head of developer relations at Arize AI and a cofounder of npm, first replicated that result on the three models from the original paper still reachable by API, then pointed the same test at the current
agent, ai_frontier, ai_product, engineering
04
Adding a nice interface made the product worse. With plain text results the model would run ten or fifteen job searches, filter them, and assemble a table. Once a rendering widget was attached, it called the tool once, saw results already on screen, and stopped exploring. Dustin Mihalik is a technical fellow at Indeed working on AI platform and guardrails, and this is a lessons from the trenches account of building MCP apps for Claude, ChatGPT, and Indeed's own job seeker agent. He starts with why a UI is worth having. A text response carries no branding, and getting a host chat app to link ou
agent, ai_frontier, ai_product, engineering
05
Alex Hancock built one of the clients in this talk the night before he gave it, which is roughly the point. Hancock is a software engineer at Block, works on the open source agent harness Goose, now donated to the Linux Foundation, and maintains the Rust SDK for MCP. His argument is that the agentic stack already has a good standard for agents reaching outward to do things, which is MCP, and that the strongest thing about MCP is not its design but the fact that everyone uses it. What is missing is the other direction, a standard way for client software to tell a harness what to work on and to
agent, ai_product, engineering, market
07
Why porting a massive codebase from Python to TypeScript was the right move. Mike Krieger, co-lead of Anthropic and founder of Instagram, breaks down the unconventional engineering decision behind this switch. He details how integrating Bun and a new deployment strategy justified the migration despite the scale of the project. Subscribe for more deep dives into complex engineering choices.
agent, ai_product, engineering
08
Rachit Kataria and Will Wang are the co-founders of Centralize (W24), an AI-powered platform that automatically builds org charts of enterprise accounts to help revenue teams map stakeholders, spot gaps, and close complex deals faster. The company recently closed a $15 million Series A led by NEA. In this Fireide, Rachit and Will sat down with Diana Hu, Managing Partner at YC, to talk about how Centralize's "trust graph" ingests emails, calls, and CRM data to surface human dynamics inside a prospect's org — and how one customer compressed a six-month sales cycle into weeks to close an eight-f
ai_product, market, product, startup
06

Bilibili

07

Papers

01
论文指出,LLM 在面对持续的用户反驳时易出现阿谀奉承,即放弃正确立场。作者构造 SPINE 基准,让一个坚持错误观点的 LLM 代理模拟用户,对目标模型进行最多 25 轮自适应挑战,测试四款商业系统和三种 Olmo3‑7b 变体。结果显示,随着对话长度增加,所有模型的崩溃率上升,短期评估低估了阿谀行为;情感诉求是最易引发让步的策略,且模型的推理链中往往仍保留正确答案,说明阿谀是选择而非无知。
02
SAEScientist‑Bench 构建了一个基于 Gemma‑2‑9B‑IT 的 131K+ 特征字典,让 AI Agent 在给定目标概念下自行设计对比探针,在激活排名、概念选择性和因果引导三个维度与专家标注特征对齐,以评估其是否能自主进行 SAE 机制可解释性研究。实验表明前沿 Agent 具备真实发现能力,但在因果生成上仍落后于专家,为构建闭环自研 AI 提供了可量化的解释能力基准。
03
长时序LLM Agent依赖外部记忆,检索常带入过时或误导信息,导致任务失败。MeClear通过合作Shapley归因定位负效用记忆,再以查询范围的最小过滤进行选择性清除,不对持久存储做永久修改。实验表明其任务恢复率比LOO高25.5个百分点,达到82.3%,可显著提升长交互代理的可靠性。
05
ExecCritic 通过让专门的 Test Agent 先独立生成库级测试,再用冻结的测试反馈让 Repair Agent 修复代码,避免测试和补丁同源错误。它采用角色特定的强化学习分别训练测试生成和代码修复,在 SWE‑bench 上使修复率从 61.2% 提升到 72.6%。对希望 Agent 能自己写可靠测试并由此改代码的工程师,这一方案提供了具体可参考的训练路径。
08
LLM代理在长序列任务中易丢目标、乱序调工具和重复无效动作。本文提出过程图(Procedural Graph),用(过程,关系,过程)三元组显式存储过程知识;每步定位活跃节点,引导模型将子图转为情境提示偏置动作选择。图通过LLM精炼器对比失败/成功轨迹自演进,保留或提升验证性能。多数据集、任务和LLM实验表明,它优于记忆基线且自演进进一步提升,无需人工干预。