01
ZhangYvJing's
Daily Brief
00
Film / Book Chapter
Moneyball
Moneyball (2011) · Bennett Miller
今天适合看《点球成金》,因为它讲的不是体育,而是在旧评价体系失灵时,怎样靠更冷静的判断和证据重新决定什么值得下注。
The Effective Engineer
The Effective Engineer · Edmond Lau
Chapter 1: Focus on High-Leverage Activities
A direct chapter for choosing what to work on today: it keeps attention on compounding engineering output rather than just being busy.
01
Insight
今天的输入更像几股不同语气的材料同时挤在一起:社区链接在暴露工程和产品环境里的真实焦点,长视频在把这些焦点放回更完整的语境里,研究材料则提醒人热度和可落地性并不总是同一件事。如果先不急着做结论,至少可以把这几条线索放在一起看:Hacker News 的 Microsoft Comic Chat is now open source;Hacker News 的 Decoy Font;Hacker News 的 NotebookLM is now Gemini Notebook;Hacker News 的 Kimi K3: Open Frontier Intelligence;Hacker News 的 LM Studio Bionic: the AI agent for open models;Hacker News 的 M 3.9 Experimental Explosion – 147 Km ENE of Ponce Inlet, Florida。真正值得注意的不是单条内容本身,而是它们共同指向了什么、彼此漏掉了什么。
03
Hacker News
02
03
04
05
06
07
08
09
10
11
Moonshot AI launched Kimi K3, a frontier-class open-weights model with 2.8T parameters, 1M-token context window, and native multimodal input. It features novel Kimi Delta Attention (KDA) enabling up to 6.3x faster decoding and Attention Residuals for ~25% higher training efficiency. K3 is live on mu
04
YouTube
01
Lee Robinson discusses the future of Cursor and AI-native software development. Speaker: Lee Robinson — ML, Model Behavior, Cursor Model research and personality at Cursor. Previously Vercel. X: https://x.com/leerob LinkedIn: https://www.linkedin.com/in/leeerob/ GitHub: https://github.com/leerob Website: https://leerob.com Timestamps 0:37 - Introduction and recursive model improvement overview 1:55 - The two-loop training framework (inner and outer loops) 2:33 - Progress and success of Composer 2.5 4:31 - Improving the outer loop with user feedback 5:40 - Climbing the inner loop with high
02
Three agents click, type, and scroll through three different apps on one desktop at the same time, and the user's own mouse and keyboard never move. That's the live demo behind cua driver, a tool the team built in a single weekend after Codex shipped its own computer use model. Instead of taking over the hardware cursor, it talks straight to the accessibility layer underneath the operating system: UI Automation on Windows, AT SPI on Linux, AX on macOS. Those undocumented APIs let a click land on a background window or a keystroke reach a hidden one, so any number of agents can act without stea
03
Full episode: https://www.youtube.com/watch?v=OS1NZLgKM2c Me on twitter: https://x.com/dwarkesh_sp
04
Where the loss function for training LLMs comes from. Job opportunities aligned to this audience: https://3b1b.co/talent Early views and other perks for supporters: https://3b1b.co/support Home page: https://www.3blue1brown.com Manim animations by Aaron Gostein and Grant Sanderson NanoGPT animation by Clayton Rabideau 3d black-box model by Paul Dancstep Chess game and Lagrange Multiplier scenes by Nishad Deulkar Music by Vince Rubinetti Timestamps 0:00 - Language trees and zipping 3:02 - Recap optimal codes 5:20 - Defining cross-entropy 8:26 - Intuition and examples 12:59 - Application to l
05
Full episode: https://www.youtube.com/watch?v=OS1NZLgKM2c Me on twitter: https://x.com/dwarkesh_sp
06
Earlier this year, OpenAI ran Parameter Golf, a model-training competition that doubled as a hiring filter. Over 1,000 researchers competed to train the best small language model under a 16MB cap. The top contributor was the one candidate OpenAI couldn't hire. Our autonomous research agent Aiden finished with 7 merged records, more than twice as many as any other contributor, and ended up the most-cited participant in the community. This talk is about what those 22 days showed. I'll cover on high level how does it works and which of its ideas produced the records. But the part worth more than
07
Eve Bouffard from Y Combinator explores the concept of "Imagination Engineering"—the idea that with increasingly powerful AI models, the primary challenge for humans is no longer technical execution, but the ability to dream up bold, innovative ideas (0:13-0:59). Key themes and experiments shared: Thinking in Public: Inspired by Paul Graham, Eve describes an experiment where she shared her stream of consciousness in a dedicated Slack channel (Eve thoughts). She then used AI to aggregate these raw thoughts into a personal, interactive website (1:35-4:26). Software on Demand: Eve demonstrates
08
Garry Tan, President of Y Combinator, discusses how the rise of AI-native companies is revolutionizing organizational productivity, allowing lean teams to operate at a scale previously requiring hundreds or thousands of employees (0:52-1:07). Key Takeaways: The 400x Leverage: Tan reports a massive increase in output (approximately 400x) by shifting from writing individual lines of code to managing AI agents (2:32-2:35). Wiring the Work: He emphasizes that founders should treat AI not as a simple autocomplete tool, but as a workforce (3:58). He explains that core organizational components—lik
09
Andy Beam (CTO) and Rafa Gómez-Bombarelli (Co-founder & CSO of Physical Sciences) of Lila Sciences join us to talk about building scientific superintelligence. Andy makes the case that the internet is a spent resource ("we have but one internet. It's the fossil fuel. We fracked"), and that the next internet-scale dataset comes from running the scientific method as reinforcement learning, with the wet lab as verifier. Science becomes an "infinite token generator" — the lab isn't the product, the model is. The counterintuitive result: one general model trained on ~10 trillion experimentally-ver
10
Bunkerhill Health recently raised $55M to help build a true state-of-the-art AI platform for hospitals. In this episode of Founder Firesides, YC's Ankit Gupta sat down with their co-founder & CEO Nishith Khandwala to discuss how their tools dramatically speed up hospital operations, the cold email that landed them Cleveland Clinic as their first customer, and a future where even the most complicated surgeries are managed end-to-end by agents. https://www.bunkerhillhealth.com Chapters: 00:00 — $55M Series B Announcement 00:52 — What Bunker Hill Health Does 03:50 — Why It Takes Two Years to
07
Papers
01
解决问题:评估多模态大语言模型在科学可视化解读上的素养与局限性。方法:基于标准的SciVis素养测评(49题,18图表,8技术),对6款闭源与开源模型与485人类对照,采用闭域协议评测其任务与技术表现。值得关注:Gemini在多项子集超越人类平均,但开源模型仍落后;模型对科学插图、检索与空间理解突出,却在纹理、综合与定量估算上失分。此基准 போர்டர்கு能直观看出哪类模型适合嵌入科研/可视化推理系统,助Engineers 快速选型。
02
解决Transformer在多步推理中“记忆缺失”问题:自回归解码把中间状态压缩,难以跨步存活。T²MLR把上一token的中层向量缓存直接注入当前token早层,实现轻量级的中层递归,让抽象推理状态持久保留。由于只需对网络20%做递归,Inference成本低且可在已有1.7B模型上快速微调,理想的Agent/AI产品升级方式。
03
本文检验了将人类测验用的项目 recovery theory (IRT) 转移到 AI 评测的可靠性,探究AI benchmark 场景下因模型数少、项目多且能力分布偏斜导致的估计失真。作者对六大 LLM benchmark 进行响应矩阵仿真,评比四种 IRT 推估方法(MML、MCMC、VI、神经估计)在计算可行性、可扩展性和排名精度上的表现。研究揭示传统估计在大规模 benchmark 中不切实际,而可伸缩方法虽实捷却易偏误,为工程师挑选合适的 IRT 估计器及诊断手段提供实用准则。
04
Plover 解决实时 GUI 自动化中快速失控、 LOVE 失配等难题,核心做法是将任务计划持久化、可视化,再借助 planner‑executor 架构提供局部编辑、自然语言指导和截图支撑,让用户可随时检查、纠正或重写执行路径。对倾向于 Agent/AI 产品的工程师来说,它把自动化行为 // Stop, the output truncated incorrectly.Plover 解决实时 GUI 自动化中用户意图易被偏离的问题,核心做法是把任务计划持久化、可视化,并通过 planner–executor 架构支持局部编辑、自然语言指引和截图回溯,让用户能直接检查、纠正或重写执行路径。对热tax于 Agent/AI 产品的工程师来说,它把所有自动化行为变成透明、可维护的“可编辑计划”,显著降低调试成本并提升系统适应性。
05
大模型 RL 用来提升推理能力,而 Masked Diffusion 需解决 log‑likelihood 计算难题。论文把生成拆成“填词”与“再掩位”两步,构成二阶段 MDP,令策略梯度拆分为 token 与 masking 两项,一并优化后在数学推理和代码任务上达到优异成绩。对需要精准生成与动态控制的 Agent/AI 工程师来说,这是可直接落地的技术思路。
06
解决精神健康AI标注的质量瓶颈,缺少结构化证据及DSM‑5‑TR对齐。提出自演化、专家互动的LLM辅助框架,三阶段:候证据选取 → DSM‑5‑TR准则分析 → 病例综合,输出诊断与严重度标签。双内存(示例/反思)聚合专家反馈,能在不重新训练的情况下持续提升标注一致性和可解释性,为构建可审计、可解释的XAI系统提供实用工具。
07
当前缺乏能把 Issue maidir 的图片证据(截图、错误框、 UI 状态)融入仓库级定位的基准。我们构建 MM‑IssueLoc:652 条跨 23 种语言的 Issue‑PR 案例,标注 7 类图像、4 层关联度,并提供文本与图像双路评估及 VCE(视觉证据转文本)诊断工具,完整区分图像在定位中的助益、干扰或被忽视。实验表明现有 LLM/检索系统在多模态定位上仍远低于预期,提示仅依赖文本或下游补丁构建的评测无法揭示多模态优势——对张玉璟的 Agent/AI 产品研发而言,这是验证视觉信息真正价值的实战入口。
08
本文揭示,用来嵌入式控制的世界‑动作模型(WAM)在通过将动作与 înt预判耦合提升鲁棒性与安全性的设计上却存在漏洞。 omn 通过极小的视觉扰动制造“世界‑动作漂移攻击”,作者提出 BadWAM 框架,区分攻击强度与隐蔽性两条维度,分别提出 action‑only 与 imagination‑preserving 两种攻击方式。实验表明,即便是细微扰动也可将优秀模型从 96.5% 降至 43.1%,显示 WAM 系统对未检测攻击极为脆弱。对 Agent/AI 产品工程师来说,这提示必须在模型训练与部署时加入对漂移攻击的检测与防御,以保障控制系统 הס鲁棒性与可信度。
08
Issue Monitor
Ready now—Actionable issues
Needs review—Awaiting a fresh check
Data statusCheck statusLive status unavailable









