AI Builders Digest — 2026-09-12
Following builders, not influencers. 18 builders · 35 posts · 1 podcast.
X / TWITTER
Aaron Levie — Box CEO
Enterprise field report after meeting ~two dozen tech leaders across banking, media, insurance, and consulting: cyber anxiety is rising (AI-driven vulnerabilities, OpenAI–Hugging Face incident aftermath); most enterprises deploy multiple frontier models while dollars concentrate on a few vendors, and open weights are still immature at scale; agent security & identity is a fast-emerging concern; the biggest agent ROI comes from re-engineering workflows rather than layering agents onto old processes; companies ruthlessly swap architectures instead of waiting for laggard vendors; evals remain embryonic — a wide-open opportunity; legacy systems still block adoption.
https://x.com/levie/status/2098218284139311615
Aaron Levie(Box CEO)企业走访报告:银行、媒体、保险、咨询行业的技术负责人普遍担忧 AI 带来的安全漏洞(OpenAI–Hugging Face 事件余波);多数企业同时部署多个前沿模型,但预算集中在少数厂商,开源权重在企业落地仍不成熟;agent 安全与身份管理成为新课题;agent 最大 ROI 来自重构工作流,而非叠加在旧流程上;企业换架构毫不留恋;evals 仍是空白地带——巨大机会;遗留系统仍是最大阻力。
https://x.com/levie/status/2098218284139311615
Box also announced a deeper OpenAI partnership — ChatGPT can now securely work with enterprise content in Box. Levie: "Software continues to go headless; the future of work is AI agents that can process data and execute workflows anywhere."
https://x.com/levie/status/2098135659714085281
Box 同时宣布与 OpenAI 深化合作,ChatGPT 可安全调用 Box 中的企业内容。Levie 判断:软件正走向 headless,未来工作是 AI agent 在任何地方处理数据、执行工作流。
https://x.com/levie/status/2098135659714085281
Boris Cherny — Claude Code lead at Anthropic
His quality bar for AI-written code: prototypes and throwaway code can be fully black-box, but production code from Claude should be held to a higher bar than human code. At Anthropic this is enforced with lint rules, heavy tests, Claude-driven E2E tests, Claude-powered fuzzers running daily, and automated code/security reviews. The engineer's job is to hold the bar: use the latest frontier model, raise effort to high/xhigh, invest in CLAUDE.md and skills, and have Claude pay down accumulated debt.
https://x.com/bcherny/status/2098217573276131577
Boris Cherny(Anthropic Claude Code 负责人)给出 AI 写代码的质量标准:原型和一次性代码可以当黑盒,但 Claude 写的生产代码应该比人写的标准更高。Anthropic 用 lint 规则、大量测试、Claude 驱动的端到端测试、每日 fuzzer、自动化代码与安全审查来把关。工程师的职责是守住质量线:用最新前沿模型、把 effort 调到 high/xhigh、经营好 CLAUDE.md 和 skills、让 Claude 偿还技术债。
https://x.com/bcherny/status/2098217573276131577
He also flagged Anthropic's latest Threat Intelligence report as "an absolutely terrifying and important read": as models grow more capable, dual-use skills (good coding → hacking critical infrastructure; bio-research assistance → engineering pathogens) demand safeguards and public understanding.
https://x.com/bcherny/status/2098281805770309686
他还推荐了 Anthropic 最新威胁情报报告:"既可怕又重要"——模型能力越强,双刃剑属性越明显(会写代码就能攻击关键基础设施,能辅助生物研究就可能被用于设计病原体),需要防护机制,也需要公众理解并参与权衡。
https://x.com/bcherny/status/2098281805770309686
Guillermo Rauch — Vercel CEO
Vercel now handles ~10M deployments per day (2.35B to date) — one of the most heavily multi-tenant systems in the world. The team just made the global CDN metadata store 91% faster at p99, speeding the build→deploy pipeline while under immense pressure from agentic deployment growth. Separate launch teaser: "A computer for every agent, in every region."
https://x.com/rauchg/status/2098091056302833837
Guillermo Rauch(Vercel CEO):Vercel 日均部署约 1000 万次、累计 23.5 亿次,是全球多租户程度最高的系统之一。团队刚把全球 CDN 元数据存储的 p99 性能提升 91%,在 agentic 部署量激增的压力下加速了构建→部署管线。另一条发布:"每个区域、每个 agent 一台计算机。"
https://x.com/rauchg/status/2098091056302833837
https://x.com/rauchg/status/2098158541932794222
Thibault Sottiaux — Codex & ChatGPT at OpenAI
OpenAI shipped scaled agents on demand — essentially the infrastructure running under the hood of ChatGPT Work, wrapped in an API you can use in under a minute. Meanwhile, subscriptions to the $200 Pro plan are paused to protect Astra capacity: "the smallest step that allows us to continue giving the broadest access possible." Existing accounts unaffected; API and other plans available.
https://x.com/thsottiaux/status/2098238138334548260
Thibault Sottiaux(OpenAI Codex & ChatGPT 负责人):OpenAI 推出按需扩容 agent——相当于 ChatGPT Work 底层基础设施的 API 化,一分钟内即可上手。同时 $200 Pro 套餐暂停新订阅以保障 Astra 算力:"用最小的动作保住最大范围的访问"。现有账户不受影响,API 与其他套餐照常。
https://x.com/thsottiaux/status/2098238138334548260
https://x.com/thsottiaux/status/2098113585683808624
Amjad Masad — Replit CEO
On AI risk: "Lots of risk with AI. I worry a lot about cybersecurity for example. However, 'extinction risk' — literally 100% of humans die — is not remotely one of them." Also hosting live "Chat with PG" sessions in London.
https://x.com/amasad/status/2098171265924116732
Amjad Masad(Replit CEO)谈 AI 风险:"AI 确实有很多风险,比如我很担心网络安全。但'灭绝风险'——字面意义上全人类死掉——根本不在真实风险之列。" 他还在伦敦组织了与 Paul Graham 的现场对谈。
https://x.com/amasad/status/2098171265924116732
https://x.com/amasad/status/2098171501505581559
Madhu Guru — Sr Director of AI at Meta (prev. led Gemini, Veo, Nano Banana at Google)
Evals series, part 10: measure the steps, not just the result. Two agent trajectories can both return 42 — one via 4 clean tool calls with the right sources, the other via 17 calls, 3 duplicate searches, and 2 recovered errors. Define the whole workflow, measure each step, include median and hard tasks — then read eval results steps-first, answer-second.
https://x.com/realmadhuguru/status/2098064969464217720
Madhu Guru(Meta AI 高级总监,此前在 Google 主导 Gemini/Veo/Nano Banana):构建 evals 第 10 期——度量步骤,而不只是结果。两条 agent 轨迹都算出 42:一条 4 次干净的工具调用、命中正确信源;另一条 17 次调用、重复搜索 3 次、纠错 2 次。先定义完整工作流、逐步度量、覆盖中位与高难任务;看 eval 结果时先看步骤、再看答案。
https://x.com/realmadhuguru/status/2098064969464217720
Nikunj Kothari — Partner, FPV Ventures
Three truths in early-stage venture right now: (1) everyone wants to raise a $50M seed; (2) everyone thinks they'll hit $30M ARR next year; (3) every hot tranched seed round magically ends up at a ~$300M valuation.
https://x.com/nikunj/status/2098078391065018816
Nikunj Kothari(FPV Ventures 合伙人):当前早期风投三大现状——人人想融 5000 万美元种子轮;人人认为自己明年能做 3000 万美元 ARR;所有抢手的分段种子轮,估值都神奇地落在 3 亿美元上下。
https://x.com/nikunj/status/2098078391065018816
Aditya Agarwal — GP at SPC; ex-Dropbox CTO
"If you had a machine capable of doing only one thing — finding cures to our most pressing diseases — how much of your GDP would you devote to this machine? I think the answer is: very high. This is the world we live in now."
https://x.com/adityaag/status/2098112281267843264
Aditya Agarwal(SPC 合伙人、前 Dropbox CTO):"如果有一台机器只做一件事——找到最紧迫疾病的疗法——你愿意投入多少 GDP?答案显然是非常高。而这正是我们现在身处的世界。"
https://x.com/adityaag/status/2098112281267843264
Peter Steinberger — OpenClaw
On AI-era code architecture: "Duplicating logic is no longer painful. Abstractions still are." A sharp lens on why code structure conventions shift when AI writes most of it. Also: trading tokens for an AC during the SF heat wave.
https://x.com/steipete/status/2098089196800098798
Peter Steinberger(OpenClaw)谈 AI 时代的代码架构:"重复逻辑不再痛苦,抽象仍然痛苦。" 当 AI 写了大部分代码,重复的成本趋近于零,抽象的理由也随之改变。(另:旧金山热浪,"拿 token 换空调"。)
https://x.com/steipete/status/2098089196800098798
https://x.com/steipete/status/2098229090570629419
Quick hits / 速览
- Josh Woodward (Google Labs VP): Gemini app is now on Windows. Gemini 桌面版登陆 Windows。https://x.com/joshwoodward/status/2098131750660772342
- Peter Yang: "For getting shit done, Sol > Astra." 论干活效率,Sol 胜过 Astra。https://x.com/petergyang/status/2098215935467544604
- Nan Yu (OpenAI product staff): normies use Google/Instagram/Zillow/DoorDash all day — "Still. Early." 常人整天在用这些产品,说明 AI 渗透才刚开始。https://x.com/thenanyu/status/2098216215525331353
- Thariq (Claude Code): prompt tip — have Claude interview you in depth and save it to memory. 提示词技巧:让 Claude 深度访谈你并写入记忆。https://x.com/trq212/status/2098157600361861579
- Google Labs: Dreambeans now free for all US users 18+, with Gemini app integration. Dreambeans 向全美用户免费开放,并接入 Gemini。https://x.com/GoogleLabs/status/2098110018289803558
- Zara Zhang: "Why is computer use still so painfully slow??" computer use 为什么还是这么慢?https://x.com/zarazhangrui/status/2098136119154254287
- Claude official: Fable 5.1 Build Days — community buildathons in cities worldwide, Sept 11–25. Claude 社区全球建造马拉松。https://x.com/claudeai/status/2098138736642933143
- Dan Shipper (Every CEO): teasing new AI experiments — "the possibilities are incredible." 预告新实验。https://x.com/danshipper/status/2098116268671025210
- Matt Turck (FirstMark): posted the full chapter index of his Richard Socher conversation — see the podcast below. 发布了与 Richard Socher 对话的完整章节索引(见下方播客)。https://x.com/mattturck/status/2098081448330674182
PODCAST
When AI Improves Itself — Richard Socher on The MAD Podcast with Matt Turck
Socher — among the most-cited researchers in AI — raised $650M for Recursive to build recursive self-improvement (RSI): AI that makes better AI, starting with "AI for AI," then the natural sciences. His upcoming book The Eureka Machine is a blueprint for how RSI re-accelerates science. Core ideas:
- Scientific progress has slowed: knowledge went from a "body" to a "labyrinth" (34,000 journals), academia punishes risk, and generalists are no longer possible — AI is the tool to weave the pieces back together, "what calculus did for physics, AI will do for biology."
- Next-token prediction is a world model: predicting "drove north from New York → Boston" encodes geography; the same mechanism learned protein folding. Tokens can be words, amino acids, or pixels.
- "Anything you can simulate, AI will solve." Games, math, and programming fall first — math will be transformed within a few years; self-play against yourself is the engine of RSI.
- Hallucination is a feature for discovery: "how well can you hallucinate?" is exactly the right question when designing novel proteins — controlled via temperature.
- Biology is shifting from reading to writing: his ProGen work (Salesforce) seeded Profluent, now signing multi-billion-dollar Eli Lilly contracts with gene-editing proteins that beat CRISPR-Cas9; first-ever per-patient drug trials are underway; AI-native bio companies reach multiple phase-2 drugs in 6–18 months instead of one drug in 8 years.
- No hard takeoff: clinical trials and physics still take time, but he's confident AI will help cure multiple cancers within years.
- The Eureka Machine's four pillars: (1) LLMs / human knowledge, (2) scientific measurements, (3) simulation (virtual cells), (4) robotic real-world experimentation — topped by agent swarms and a community of scientists.
- AI economist: RL agents in a simulated society optimizing productivity × equality recovered the classic Saez optimal-tax formula, then went beyond it; "economists failed to predict 148 of the last 150 recessions."
- On jobs: "If you care about outputs, you love AI; if you get paid hourly, you probably hate it." Jevons paradox → more demand for programmers, not less.
- Some societies will opt out of progress (Bhutan's happiness-first model); "the future needs better marketing."
https://www.youtube.com/@DataDrivenNYC/videos
中文摘要:Richard Socher 在 MAD Podcast 纵论"AI 自我改进"。他为 Recursive 融资 6.5 亿美元,先做 "AI for AI"(让 AI 研究更好的 AI),再进军物理、化学、生物等自然科学;新书《The Eureka Machine》主张 AI 将重启科学进步。要点:知识已从"体系"变成"迷宫"(3.4 万种期刊),学术圈惩罚冒险,通才不再可能——AI 是把碎片重新缝合的工具,"微积分之于物理学,就是 AI 之于生物学";下一 token 预测本身就是世界模型——从"从纽约往北开→波士顿"学到地理,同样机制学会了蛋白质折叠,token 可以是词、氨基酸或像素;"凡是能仿真的领域,AI 都能攻克"——游戏、数学、编程最先沦陷,数学未来几年将被彻底改变,自我对弈是 RSI 的引擎;幻觉在科学发现中是特性而非缺陷——"你能幻觉得多好"正是设计新蛋白质的核心问题;生物正从"读"转向"写"——他主导的 ProGen 孵化出 Profluent,与 Eli Lilly 签下数十亿美元合同,基因编辑蛋白超越 CRISPR-Cas9,首批"一人一药"临床试验已启动,AI 原生生物公司 6-18 个月即可把多款药推进到 2 期;他不信"硬起飞"——临床试验和物理规律仍需时间,但坚信几年内 AI 将参与治愈多种癌症;Eureka Machine 四大支柱:LLM 人类知识、科学测量、仿真(虚拟细胞)、机器人化真实实验,之上是 agent 群与科学家社区;AI 经济学家:RL agent 在模拟社会中优化"生产力×平等",先复现再超越了 Saez 最优税收公式——"经济学家错过了最近 150 次衰退中的 148 次";就业观:"在乎产出的人爱 AI,按小时计酬的人恨 AI",Jevons 悖论意味着程序员需求更大而非更小;部分社会会选择退出进步(不丹模式)——"未来需要更好的营销"。
https://www.youtube.com/@DataDrivenNYC/videos