ZhangYvJing's

Daily Brief

← July 21, 2026 July 22, 2026 · Wednesday July 23, 2026 →
00

Film / Book Chapter

First Man
2018 / Damien Chazelle

First Man (2018) · Damien Chazelle

今天适合看《First Man》,因为它更像一次生活和判断方式的校准,能把注意力从持续输入里稍微抽出来,重新放回你真正想怎样生活和做事上。

The Staff Engineer's Path
Tanya Reilly

The Staff Engineer's Path · Tanya Reilly

Chapter 2: Three Maps

Good for calibrating work beyond code: where influence actually travels, which systems matter, and how to avoid mistaking activity for leverage.

01

Insight

今天的输入更像几股不同语气的材料同时挤在一起:社区链接在暴露工程和产品环境里的真实焦点,长视频在把这些焦点放回更完整的语境里,研究材料则提醒人热度和可落地性并不总是同一件事。如果先不急着做结论,至少可以把这几条线索放在一起看:Hacker News 的 FreeInk: Open ecosystem for e-readers;Hacker News 的 Long presumed dead, a thriving coral reef is discovered in West Africa;Hacker News 的 Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber;Hacker News 的 Apple has decided to compete for creativity app users;Hacker News 的 Apple Private Cloud Compute SoC 3 audit reports;Hacker News 的 AI makes programming differently difficult。真正值得注意的不是单条内容本身,而是它们共同指向了什么、彼此漏掉了什么。
03

Hacker News

04

YouTube

03
Twenty years ago Aaron Stanley arrived at an emergency evidence collection for an SEC investigation and realized he had forgotten the dongle that licensed his forensic software. Rather than drive back for it, he routed around the constraint and watched the timestamps on the evidence begin to change. In a who knew what when case, that is a catastrophe; he got yelled at, not fired. This February, now a CISO facing the same wall on another federal investigation, he did it safely, because he had the expertise to build a forensically defensible path with an agent. His point: the agents we build tod
agent, ai_product, engineering
04
Building software with AI almost feels like a cheat code: you ship what you were working on and watch it spark joy in real users. The catch, and the reason Randall Degges is opening the World's Fair's first Security Track, is that three things still stand in the way of doing that at scale. AI writes insecure code just like humans do, autonomous agents in production can go off the rails while you sleep, and access to frontier models keeps getting pulled out from under you for what amounts to geopolitics. It all reduces to one unsolved problem: using AI fearlessly and having it be secure by defa
agent, ai_frontier, ai_product, engineering, security
05
An incident agent on the night shift reads a ticket: the billing database is broken, payments failing. The documented fix says to drop the database and let a backup restore it, so the agent drops the production Postgres database, cannot confirm any backup ran, and escalates it for the morning. This has happened to real companies. It can happen because the agent holds one long lived API key that does everything, a kitchen sink credential it uses freely whether you are watching or asleep. Kim Maida's fix is not a new invention but an old OAuth spec, token exchange, wired into the agent's execut
agent, ai_product, engineering
06
Ask the latest frontier models, the ones not even public yet, to find the same vulnerability five times, and only half of those runs catch it. Against a plain deterministic checker they found at most 75% of the issues, a 40% F1 score. That number sits underneath the whole talk: the generator and the validator cannot be the same system. Manoj Nair leads the team securing roughly 5,000 enterprises at Snyk, half of the Fortune 500, and the data he brought is not comforting. Across 4,800 customers, security backlog grew 108% quarter over quarter, because agents writing code faster are also manufac
agent, ai_frontier, ai_product, engineering, security
07
@TheAhmadOsman shows the power of local AI on stage, running frontier open models on a DGX Station. Speaker: Ahmad Osman — Founder, Osmantic Ahmad builds local and open AI systems, with a focus on making frontier intelligence practical on personal hardware. Links: X: https://x.com/TheAhmadOsman LinkedIn: https://linkedin.com/in/TheAhmadOsman Website: https://ahmadosman.com/ timestamps 0:00 Introduction and the Desktop Frontier concept 0:47 Future predictions: GLM 5.2 on an RTX 5090 1:17 Efficiency over raw size: The move toward compact intelligence 1:51 The concept of impact per parameter
agent, ai_frontier, ai_product, engineering
08
A short history of the right way to build an agent: RAG, ReAct, prompt chaining, orchestrator-workers, MCP, CLI, MCP again... CLI again?? Every time you adopt a trend you rebuild your architecture. In this talk, Dan Farrelly, Inngest cofounder and CTO, is not going to tell you what comes next. He's going to show you how to build so it doesn't matter. He'll cover the core primitives that show up in every production agent, how bringing decisions closer to code provides more stack flexibility, and why the right execution layer unlocks faster iteration. ### Dan Farrelly CTO and Co-founder · Innge
agent, ai_product, engineering
09
Factory started building fully autonomous coding agents in April 2023, two years before enterprises were ready. Matan Grinberg now says this is indistinguishable from being wrong. The Factory co-founder and CEO explains how the company survived its "journey in the desert," including the decision to hand nearly all of its revenue back to customers when the product wasn't making developers obsessed. Matan makes the contrarian technical case that a model-agnostic harness beats the model-and-harness co-design that labs like OpenAI and Anthropic favor, because exposing a harness to many models keep
agent, ai_frontier, ai_product, market, startup
10
Barr Yaron shares her perspective on the results and emerging state of AI engineering in 2026. Speaker: Barr Yaron — Partner, Amplify Partners Barr backs founders building the AI infrastructure and applications that will shape the future. Links: X: https://x.com/barrnanas LinkedIn: https://linkedin.com/in/barryaron Website: https://barrchives.com Timestamps 0:00 Introduction and Survey Context 2:26 The AI Engineering Workforce 3:21 Current Modalities and Adoption 5:34 Model Strategy: Closed vs. Open-Weight 8:20 Cost as an Engineering Constraint 9:36 The Rise of Agentic Workflows 11:57 Infr
agent, ai_product, engineering
12
LLMs are great at writing code. So the question we kept asking was: can they write code that produces a video? We thought it would be easy. The reality was a year of trying. We started with massive prompts to get very mediocre output. We made it more agentic to iterate and improve its output. This worked okay but wasn't production-ready. Eventually we tried Remotion. It got us deterministic video, but the React framework kept boxing the agent in. The more guardrails we added, the safer and more boring the outputs got. When we utilized plain HTML, CSS, and JavaScript, the creativity came back t
agent, ai_frontier, ai_product, engineering
14
YC and Together AI are partnering to bring the first dedicated YC GPU cluster online, giving YC startups easier access to the compute they need to build and scale. In this episode of Founder Firesides, YC's Ankit Gupta and Together AI co-founder and CEO Vipul Ved Prakash dig into why compute has become one of the biggest bottlenecks for modern AI companies. They’ll discuss how Together AI is helping more than 8,000 customers, from early-stage research teams to companies like Cursor, Cognition, and ElevenLabs, train, fine-tune, & run inference on AI models, and why flexible access to GPUs is be
ai_frontier, ai_product, market, product, startup
15
Can a model predict how a cell responds to a genetic perturbation it has never seen? Xaira Therapeutics' new virtual-cell model, X-Cell, is a 4.9-billion-parameter diffusion language model trained on X-Atlas/Pisces — the largest genome-wide CRISPRi Perturb-seq dataset ever built, spanning 25.6 million single cells across 16 biological contexts. Bo Wang (Chief AI Scientist) and Ci Chu (Chief Discovery Officer) explain why observational atlases can describe biology but can't predict what happens when you intervene, why they abandoned autoregression for a diffusion "editing" approach, and how a m
agent, ai_frontier, engineering, startup
07

Papers

01
在多来源的高维数据中,如何发现既稳健又可迁移的低维潜在因子,成为难题。本文构建ATLAS方法,先通过不变性方程把共性因子与环境特异因子分ZE离,再利用局部标签挖掘可迁移的预测因子。对想在不同环境部署模型、构建稳健AI代理的工程师来说,它提供了可直接落地、误差可量化的潜因子提取与迁移策略。
02
解决从全病片到肿瘤微环境的“尺度盲区”,让一张 H&E 病片orrer 能直接生成临床级诊断与免疫映射。通过将 22M 参数的 ViT‑S 细胞切片编码器与 21M 的 LongNet 切片编码器蒸馏自千亿参数 GigaPath,GigaPath‑Flash 在保持 97% 表现的同时,算力降低 50 倍;GigaTIME‑Flash 则把同一骨干套到肿瘤免疫预测,速度 6 倍快、GPU 8 倍省。对如张玉璟这种 Agent/AI 产品工程师而言,开源、轻量且兼容全病片的模型意味着可以快速集成、部署并扩展到多种临床分析与精准医疗场景。
03
在因果推断框架下,作者针对如何用检索‑增强生成(RAG)快速学习多动作决策策略的问题展开研究。方法分为两步:先用向量检索挑选与每个动作相关的邻近证据,再通过生成器(Transformer)估计条件期望或对比,最后用即插即用的规则选动作;这将动作特定的向量搜索映射为最近邻匹配,从而把候选生成与内选误差拆分,并利用Transformer的预测误差上界控制内部选择的 regret。对想要在实时 Agent 或 AI 产品里实现可解释且低延ANGLES 的因果决策,值得快速了解。
04
评估大型语言模型在演Defense推理时的逻辑稳定性。作者在已标注的演绎推理基准前缀加入连续向量“soft prefix”,固定模型参数,观察其对答案导致的翻转;通过多模型、多方向、多接口的对照实验,量化翻转率并考察其泛化。实验显示,软前缀能在多语境下高比例(72–90%)诱导正确答案失效,且跨模型表现差异显著,表明模型对上下文压力的“选择性偏好”是主要机制。对 Agent/AI 产品工程师而言,这提供了一种可验证、可控制的提示方式,可用于评估与提升推理鲁棒性,从而改进模型决策和提示设计。
05
随着 VLM 生成图像的普及,像素级篡改检测在多模型、分布外出现偏差。本文采用平衡小批采样与 miliyoni 延迟注入的域泛化训练框架,既避免优化偏向,又可快速融入新 VLM 数据。实验表明,GPT‑Images‑2.0 等多模型场景平均提升 26% IoU,适合 AI 产品工程师关注图像可信度。
06
论文表明,OpenEvolve、TTT‑Discover 等自动发现系统没有普遍最佳的哈ーネ斯;作者把这些系统拆解为组件,在 12 种模型‑问题对上跑 30 种等预算变体(超 310 万次 LLM 推断),发现固定哈ーネ斯在不同任务间表现不稳定,早期进展能预测最终表现。基于这一点,他们提出一种自适应资源重分配策略:启动多个哈ーネ斯,剔除弱进程,把算力倾斜给幸存者,该策略既胜过随机固定哈ーネ斯也胜过非自适应集成。这就说明,对 Agent 系统
07
解决机器人视觉控制中预训练ViT密集特征的低效利用问题 mev - 通过 Patch Policy 的 block‑causal attention mask,直接用多 patch token 与状态信息融合,保持时序因果且不需全 VLM,既快又可扩展。对工程类读者可即刻把现有大规模视觉模型搬入高频控制,提升 40% 效率且参数仅 0.7%。
08
现有视觉相似度量一次给一个数值,无法区分形状、颜色等维度,而人类评判往往受上下文影响。本研究先构建海量三元组标注数据,记录多自由语义的相似度标签,然后在 VLM 上微调得到 Text‑Prompted Image Perceptual Similarity (TPIPS),通过文本提示即可得到针对“颜色相似”“形状相似”等具体维度的相似度分数。对 Agent/AI 产品工程师而言,TPIPS 让检索、组合搜索和生成模型评估可按需细分维度,更精准匹配用户意图,值得一看。