ZhangYvJing's

Daily Brief

← July 18, 2026 July 19, 2026 · Sunday July 20, 2026 →
00

Film / Book Chapter

Perfect Days
2023 / Wim Wenders

Perfect Days (2023) · Wim Wenders

今天适合看《Perfect Days》,因为它更像一次生活和判断方式的校准,能把注意力从持续输入里稍微抽出来,重新放回你真正想怎样生活和做事上。

The Beginning of Infinity
David Deutsch

The Beginning of Infinity · David Deutsch

Chapter 1: The Reach of Explanations

A broader chapter for sharpening what counts as a good explanation, useful when daily inputs are full of claims, demos, and partial narratives.

01

Insight

今天的议题从“低成本游戏化”延伸到“透明可复制的自动化”,整体重心正向可落地、成本透明的方向上升,而讨论的热度与技术深度并不总匹配。REO Trucks 推出 21,500 美元的汽油四驱皮卡,以“成熟发动机+可拆装维护”吸引美国经济型车主;同样以成本为先的思路也在 K3 这款 2.8T 开源大模型里显现,K3 的“开权重+1M 上下文екта”被多家评测超越 Fable5 但仍落后于 GPT‑5.6,价与性能的折冲再次陷入“性能不必定为先”这一行业争论;而 GoPro 的财务危机提醒我们,成本失衡会使一味追求新硬件的公司走向边缘。与此同时,AutoSynthesis、teLLMe、SearchOS 等egiatan 体系证明了从自然语言到可验证的决策链条正在被自动化:从元分析、因果推断到信息检索,各用例都在追求“可复制决策流程”与“成本意识”,这与传统的“性能指标”导向正 clicks 形成强烈对比。另一方面,安全评测论文强调在固定预算下衡量攻击防御代理的成本成功率,暗示机型的普及不再只是性能比拼,更是“预算内可持续”考量。可见,技术讨论从远离“噪声”中的模糊叙事到聚焦成本、可落地模型的共识正在形成;但仍有学术热潮与行业商业噪音交错,它们之间的关键词是“成本”,而非单一的性能高低。今天要看的不是单一技术细节,而是这种成本聚焦与可落地化的交叉点。Perfect Days (2023)。
03

Hacker News

01
REO Trucks I4 4WD Pickup Truck Starts at $21,500
REO Trucks 推出全Speak I4 四驱皮卡,起售价 21,500 美元。厂方选择成熟汽油发动机,因 90% 新车仍沿用燃油技术,旨在将更低成本与更高可靠性推向更广泛的美国市场。此举将使得车主获得更长的 500,000 英里发动机寿命、可自行拆装零件的维护方式,以及更低的维修成本,从而降低整体运营风险。
06
If You Build It, They Will Come
社会群体的快速加入方式已从被动参与转为主动组织。文中指出,需求远大于供应,只有投身组织的个体才能启动活动,普通成员普遍不愿付出。于是,想扩大人脉的人从仅仅参加转向策划并邀请,能显著降低等待成本,并直接提升其社交效率和合作成效。
07
Elixir-lang.org has a new design
Elixir 官方网站发布了全新的设计方案。该设计借助 Elixir 的不可变性、内存安全性和逐步类型系统,帮助开发者聚焦数据和业务领域,写出更清晰、更具目的性的代码。这样,团队能够构建能在失败后迅速恢复且易于维护的系统,从而降低维护成本并提升整体可靠性。
08
Google DeepMind 与 Isomorphic Labs 公布新生物韧性合作模式,重点阻止模型被滥用并支持政府与科研机构。该模式依托 AlphaFold、IsoDDE 和 AlphaGenome 等前沿 AI,能在预防、检测与应对阶段提升病毒监测、药物与疫苗等研发效率。结果使
10
Fable 5 在未公开的 NP‑hard 纤维网络设计任务上超越 GPT‑5.6 Sol,即使启用 /goal 也仅略增优势。实验表明 /goal 通过改变控制循环和搜索路径,有 procédure 有时把搜索聚焦至更优基底,却也可能让模型陷入劣势区间,从而使整体平均性能下降。此现象表明,持续化特性虽然能提升单次试验成功率,却会增加系统表现波动,对调优策略、成本评估以及风险管理带来不确定性。
11
not much happened today
Moonshot AI launched Kimi K3, a frontier-class open-weights model with 2.8T parameters, 1M-token context window, and native multimodal input. It features novel Kimi Delta Attention (KDA) enabling up to 6.3x faster decoding and Attention Residuals for ~25% higher training efficiency. K3 is live on mu
04

YouTube

01
A keynote exploring generative AI for code, deep-thinking algorithms, and the future of pre-training and transformer models for Gemini. Speaker: Benoit Schillings leads the Thinking, Reasoning, and Coding teams at Google DeepMind, directing foundational research toward AGI. His work focuses on advancing next-generation model reasoning and integrating software development best practices into AI code generation. Previously, as CTO at X, Benoit guided early-stage teams prototyping Alphabet's moonshot technologies across computing, biochemistry, and clean energy. LinkedIn: https://www.linkedin.
agent, ai_frontier, ai_product, engineering, market, security
02
Pablo Castro explores AI and knowledge systems for building better applications and agents. Speaker: Pablo Castro —Distinguished Engineer and CVP, Microsoft, leads the AI Knowledge team in Microsoft's CoreAI division, where he focuses on state-of-the-art information understanding and retrieval systems for AI applications and agents, including Foundry IQ, Azure AI Search, and Azure Content Understanding. LinkedIn: https://www.linkedin.com/in/pabloc Timestamps: 0:00 Introduction and speaker background 1:14 Defining the nature of knowledge: Intrinsic, Extrinsic, and Learned 1:27 Intrinsic kno
agent, ai_product, engineering
05
Code has become the fastest medium for producing technical content, but the most important piece isn't a secret agent skill or a magical framework: its the same thing that makes a great developer experience. Speakers: - Matt Palmer (Conductor): Matt Palmer is a DevRel & product leader focused on AI devtools, developer education, and making complex concepts accessible. Away from the keyboard, you can find him lifting, hiking, riding motorcycles, or caring for plants. X/Twitter: https://x.com/mattppal LinkedIn: https://www.linkedin.com/in/matt-palmer/
agent, ai_product, engineering
06
I pointed my lab at one problem, inference, after 200 users burned $1,000 in credits and the math just wouldn't close. So I built the thing, felt the cost, and went looking for why renting intelligence never pencils out. Turns out everyone in this market sells a gospel shaped like their own invoice. Jensen: build a token factory. Nadella: don't even think about the meter. Fireworks: own your model (on our infra). Three smart people, three different layers, three pitches that all end at "keep paying us." My rule: rent to learn, own to run. Rent the model while you're hunting PMF, own the infere
agent, ai_frontier, ai_product, engineering, market
08
We ran auto-improvement loops on a paper classification task against a ground-truth dataset. A real problem, narrow enough to measure precisely, and in fact one of the few clear cut target functions out there. We’ll share how to properly set up an agent for auto-improvement, what task specificity and target function quality is actually required for it to work, and why the most efficient path to a continuously improving agentic system is one where domain experts and automation know when to hand off to each other.​​​​​​​​​​​​​​​​ Speakers: - Annabell Schäfer (Langfuse): Annabell is a Growth En
agent, ai_product, engineering, market
09
As a developer, AI is fun, exciting, and full of potential – but users don't always feel the same way about it. From a UX perspective, AI comes with a whole new set of considerations around user trust, privacy, and security. From a UI perspective, AI brings new interaction patterns, new icons, new visual cues, and so much more! If we want people to get the most from what we build, we have to teach our users how to use AI. Let's look at ways to introduce new capabilities in our apps and guide our users through new patterns and processes – ideally without making them throw their phone out a wi
agent, ai_product, engineering, security
07

Papers

01
AutoSynthesis 把原来繁琐手工的定量证据合成搬进机器,让研究者只需一句自然语言问句,系统就能自动制定检索策略、抓取、筛选、抽取统计、计算效应量并做随机效应 Meta 分析,还能评估异质性与偏倚,并生成 PRISMA 报告。实验显示其合成结果与专家手工高度一致,证明自动化的可行与可扩展,对 Agent/AI 产品工程师关注的可复制决策流程尤为重要。
02
想知道雨天会不会导致拥堵?teLLMe把摄像头视频做成事件表,利用PC算法学习因果结构、DoWhy+线性回归估效应,并用LLM把自然语言提问映射为结构化查询,最终生成“Causal Card”总结效应、调整集、DAG支持与假设。对于Agent/AI产品工程师,它把海量观测数据转化为可检验的假设,直观呈现天气、时段ים 车流量的因果关系并显式不确定性,方便快速生成洞察与决策。
03
当信息寻址代理随交互历史膨胀,搜索进度难以追踪,往往陷入循环浪费预算,导致结果不完整。SearchOS 通过 SOCM 把搜索状态外化为 Frontier Task、Evidence Graph、Coverage Map 与 Failure Memory,并采用 pipeline‑parallel 调度与 Search Tool Middleware Harness,实时补齐未覆元素并记录失败模式,形成可复用的层级技能。对 Agent йолý 或 AI 产品工程师而言,它提供了可监控、可提升吞吐量的协作框架,显著提高搜索效率与完成度。
04
论文解决了如何从比特币的链上交易、历史价格与 Twitter 情绪中解码市场情绪。作者将这些数据融合成统一数据集,利用 XGBoost 训练情绪分类 😉 通过 SHAP 解释模型,量化链上特征对情绪预测的贡献,提升透明度。对 Agent/AI 产品工程者而言,能快速获取可解释的加密市场驱动信号,直接嵌入决策逻辑或风险评估模型,提升运营效率与风控能力。
05
本文指出安全代理评测往往只关注峰值成功率,却忽略成本。作者通过在固定计算与工具调用预算下,对 Cybench 红队与 Splunk BOTS 蓝队任务做费用‑成功比对,拆解推理与工具开销,揭示两侧的不同加速规律。对 Agent/AI 产品工程师而言,这套成本‑效能评估能精准衡量模型在 SOC 或 CTF 场景
06
SceneBind 解决多模态场景理解中“what‑and‑where”缺失的问题,采用全局语义向量+对象级语义‑空间槽的表示,并通过 SceneBind Matching 将全局相似度与对象对齐结合,支持跨模态检索与定位。其轻量化架构、预训练兼容性以及零样本迁移,使 Agent/AI 产品能够快速获取可定位的多模态推理能力,提升环境感知与决策效率。
07
科研论文里的图表往往在多轮修订中被不断重绘、重标,手工完成耗时又容易出错。SciDiagramEdit 试图用自然语言指令,直接在图的向量源上自动化这些编辑。研究者先在 arXiv 版次中抽取前后版本的图对,结合同步出现的编辑意图,构建一个基准集。随后用 “agentic learning” 与 “skill evolution”——即通过多轮执行轨迹不断微调代理的技能说明——让智能体逐步提升对各类编辑任务(如重排面板、改色、调整注释)的精准度。对张玉璟这类关注 AI 强化学习、可解释代理与前端交互的人来说,可能是一个把自然语言指令落到图形编辑这类工业场景的有价值范式。
08
MeanFlow generator 用平均速度快速采样,却缺乏可以与 RL 奖励兼容的优化方法。作者用 DiffusionNFT 的框架,引入一个“诱导瞬时速度预测器”,将平均速度与即时速度关联,再对该预测器使用 DiffusionNFT 的奖励目标,从而实现可定义的 RL 优化;采样仍按平均速度进行,保持几步即能完成。该方法在图像、视频生成上显著提升多项指标,且在仅几步采样的条件下,能击败许多需数十步的 RL‑tuned diffusion,适合追求高效、可调节的 Agent/AI 产品工程。