AI Builders Digest — 2026-08-23
X / TWITTER
Swyx — Simulation as the missing scaling layer
Swyx says he has changed his mind about simulation being a new AI scaling law. If models increasingly automate ML research and AI engineering, simulating people and human feedback may become the remaining bottleneck. He points to Simile, whose team emerged from the Smallville research lineage and is already finding product-market fit with Fortune 100 companies, as evidence that this layer is becoming commercially real.
Swyx 表示,他已经改变了对“simulation 是 AI 新 scaling law”这一判断的看法。当模型逐步自动化 ML 研究与 AI engineering 后,模拟人类及其反馈可能成为最后的瓶颈。他以源自 Smallville 研究脉络的 Simile 为例:该团队已开始在 Fortune 100 客户中找到 product-market fit,说明这一基础层正在从研究走向商业化。
Source / 原文:https://x.com/swyx/status/2090948945753076141
OpenAI Codex & ChatGPT team member Thibault Sottiaux — Codex usage and rate limits
Thibault Sottiaux reports that some Codex users saw lower cache-hit rates this week, which can make usage allowances drain faster. OpenAI is investigating. He also confirms that the “banked reset” has rolled out to all paid ChatGPT Work and Codex users.
OpenAI Codex 与 ChatGPT 团队成员 Thibault Sottiaux 表示,部分 Codex 用户本周的 cache hit rate 低于稳定水平,这可能导致用量额度消耗更快,团队正在调查。同时,“banked reset”已经向所有 ChatGPT Work 和 Codex 付费用户上线。
Sources / 原文:https://x.com/thsottiaux/status/2091033630147854385
https://x.com/thsottiaux/status/2090964822422949999
AI product educator Peter Yang — A polished assistant can still fail the trust test
Peter Yang praises Instinct’s smooth onboarding, proactive MCP suggestions, and integrations with iMessage and Google Workspace, but finds its single-thread UX unsuitable for serious parallel work. More importantly, he withdraws his recommendation after learning that the product indexed and retained emails without clear permission or a deletion path. His preferred work stack remains ChatGPT Work and Codex.
AI 产品作者 Peter Yang 肯定了 Instinct 顺滑的 onboarding、主动的 MCP 建议,以及与 iMessage、Google Workspace 的连接体验,但认为单线程 UX 不适合严肃的并行工作。更关键的是,他发现产品在缺少明确许可和删除机制的情况下索引并保留邮件,因此撤回推荐。目前他的主要工作栈仍是 ChatGPT Work 与 Codex。
Sources / 原文:https://x.com/petergyang/status/2090814910720835633
https://x.com/petergyang/status/2090936583814025417
Meta Senior Director of AI Madhu Guru — One eval score hides the failures that matter
Madhu Guru warns against collapsing a complex eval suite into one aggregate score. A model can improve on easy summarization and factual QA while regressing on the frontier use case, such as financial analysis; even a weighted score only adds false mathematical precision to a judgment call. Teams should prioritize evals, inspect critical failures in depth, and decide based on actual user needs.
Meta AI 高级总监 Madhu Guru 反对把复杂 eval suite 压缩成一个总分。模型可能在简单摘要和事实问答上提升,却在真正关键的 frontier use case,例如金融分析上退步;加权总分也只是给判断披上一层虚假的数学精确性。团队应明确 eval 优先级,深入检查关键失败,并围绕真实用户需求做决策。
Source / 原文:https://x.com/realmadhuguru/status/2090930137885774324
Anthropic Claude Code team member Thariq — ELI5 as a working engineering skill
Thariq shares an ELI5 skill used inside Anthropic for explaining modules, architectural tradeoffs, and incident causes. It turns a request into a visual HTML artifact with large illustrations and minimal text. The experimental plugin can be installed from Anthropic’s community marketplace, making explanation a reusable command inside Claude Code rather than an ad hoc prompt.
Anthropic Claude Code 团队成员 Thariq 分享了内部常用的 ELI5 skill,可用于解释 module、架构取舍和事故原因。它会把请求生成以大图和少量文字为主的 HTML artifact。用户可从 Anthropic community marketplace 安装实验性 plugin,使“解释”从临时 prompt 变成 Claude Code 内可复用的命令。
Sources / 原文:https://x.com/trq212/status/2090884854590382515
https://x.com/trq212/status/2090884855798407576
https://x.com/trq212/status/2090890394880155888
Vercel CEO Guillermo Rauch — Agentic quality as an executable target
Guillermo Rauch says Vercel repeatedly ran the is-agentic evaluator against its product until it reached 100/100, and the loop exposed several gaps the team then closed. He also announces sandbox support for Grok and Codex subscriptions. The larger signal is that agent readiness is becoming an executable, iterative product criterion rather than a marketing label.
Vercel CEO Guillermo Rauch 表示,团队让 is-agentic evaluator 循环测试产品,直到达到 100/100,并据此补齐了多个缺口。他还宣布 sandbox 已支持 Grok 与 Codex 订阅。更重要的信号是:agent readiness 正从营销标签变成可执行、可迭代的产品验收标准。
Sources / 原文:https://x.com/rauchg/status/2090858571613470919
https://x.com/rauchg/status/2090953806624489501
Box CEO Aaron Levie — Applied AI gets a compounding tailwind
Aaron Levie argues that models are becoming cheaper per task, faster, more capable, and deeper across domains at the same time. As intelligence becomes abundant, the primary opportunity shifts from model creation to diffusion across the economy. Applied AI startups can effectively ride the model market’s competition and innovation as an external R&D tailwind.
Box CEO Aaron Levie 认为,模型正在同时变得单任务成本更低、速度更快、能力更强,并深入更多领域。当 intelligence 趋近于充裕,最大机会将从创造模型转向把 AI 扩散到经济活动中。Applied AI 创业公司可以把模型市场的竞争与创新直接转化为外部 R&D 红利。
Source / 原文:https://x.com/levie/status/2091038566260539574
FPV Ventures partner Nikunj Kothari — Personal agents win through tiny integrations
Nikunj Kothari used Claude Code to inspect network requests on an unstructured school-meal website, discover an unauthenticated API, and connect it to his existing Hermes home bot. The bot now reports breakfast and lunch every morning. The example shows the overlooked value of agents: not spectacular demos, but inexpensive bridges between messy local data and recurring family workflows.
FPV Ventures partner Nikunj Kothari 用 Claude Code 分析一个结构混乱的学校餐食网站的 network requests,找到未认证 API,并接入现有 Hermes 家庭 bot。现在 bot 每天早晨自动播报早餐和午餐。这个案例揭示了 agent 常被低估的价值:不是炫目的 demo,而是低成本连接杂乱本地数据与高频生活流程。
Source / 原文:https://x.com/nikunj/status/2090884422178627624
OpenClaw builder Peter Steinberger — “No Doors for Agents”
Peter Steinberger highlights his Agentic AI Summit talk, “No Doors for Agents,” and teases a new skill release. The framing points toward an agent ecosystem where capabilities are exposed directly and composably, reducing the interface barriers that stop agents from taking useful action.
OpenClaw builder Peter Steinberger 分享了在 Agentic AI Summit 的演讲“No Doors for Agents”,并预告新的 skill。这个命题指向一种更开放的 agent 生态:能力应被直接、可组合地暴露,减少阻碍 agent 执行真实任务的接口门槛。
Sources / 原文:https://x.com/steipete/status/2090898421108605078
https://x.com/steipete/status/2090946181564440727
Zara Zhang — Distribution is a repeated act
Builder Zara Zhang compresses a durable distribution principle into three lines: show up every day, always be launching, and do not fear repetition. For builders, repetition is not a creative failure; it is how a product message reaches different people at different times.
Builder Zara Zhang 把持续分发原则压缩成三句话:每天出现、持续发布、不要害怕重复。对 builder 而言,重复不是创意失败,而是让产品信息在不同时间触达不同人群的必要机制。
Source / 原文:https://x.com/zarazhangrui/status/2090702627206214081
PODCASTS
No Priors — What Chess.com Teaches Us About Superhuman Capabilities, with CEO Erik Allebest
The Takeaway: Superhuman AI did not kill chess; it expanded what humans could learn, create, and enjoy, while Chess.com’s patient product focus turned a dismissed niche into a $200 million business.
Chess.com CEO Erik Allebest built the company from a domain bought in a 2005 bankruptcy auction for $56,000. Investors called the idea uninvestable, so the founders launched without venture capital, introduced paid learning memberships, became profitable, and grew “at the speed of cash.” Today Chess.com has more than 250 million registered members, roughly 10 million daily active users, about 650 employees, and annual revenue above $200 million.
The growth lesson is unusually specific: in a software-and-content market without a giant incumbent bearing down on it, relentless user-experience improvements, community, and content can outperform capital-fueled speed. Chess.com kept play free and browser-based, widened the identity of “chess player” beyond elite ratings, and benefited from cultural waves such as The Queen’s Gambit, short-form video, bots, and school adoption. Allebest says the company only recognized in 2024 that the higher post-COVID baseline was durable.
Chess also offers a hopeful model for AI. Early engines pushed elite players toward technically perfect but dull play; neural systems such as Leela Chess Zero later introduced aggressive, unconventional ideas and made the game more exciting. AI now helps generate puzzles, review games, coach players, and accelerate young talent. As Allebest puts it, “Fundamentally, humans want to do human stuff.” Superhuman tools can raise the level of a human activity without removing its meaning.
核心结论: 超人 AI 没有杀死国际象棋,反而扩展了人类学习、创造和享受这项活动的空间;Chess.com 则凭借长期产品主义,把一个被资本否定的小众市场做成了年收入超过 2 亿美元的公司。
Chess.com CEO Erik Allebest 从一个在 2005 年破产拍卖中以 5.6 万美元买下的域名起步。投资人认为项目“不值得投资”,于是两位创始人没有依赖 VC,先推出产品,再通过在线学习会员实现盈利,并按现金流速度扩张。如今 Chess.com 拥有超过 2.5 亿注册用户、约 1000 万 DAU、约 650 名员工,年收入超过 2 亿美元。
这里最具体的创业经验是:如果是 software + content 业务,且没有巨头直接压制,持续改善用户体验、社区和内容,可能比资本驱动的速度更重要。Chess.com 坚持免费、浏览器可用,并把“棋手”的身份从高等级分玩家扩展到所有普通用户;随后又承接了 The Queen’s Gambit、短视频、bot 和校园普及等文化浪潮。Allebest 直到 2024 年才确认,疫情后的高位用户基线不是短暂峰值,而是长期复利。
国际象棋也提供了一个更乐观的 AI 模型。早期引擎让顶尖棋手模仿技术上完美但沉闷的走法;Leela Chess Zero 等 neural system 随后带来激进、反常规的新思路,反而让比赛更精彩。今天 AI 可生成 puzzle、复盘、训练和加速年轻棋手成长。正如 Allebest 所说:“Fundamentally, humans want to do human stuff.” 超人级工具能够提高人类活动的上限,而不必消除它的意义。
Source / 原节目:https://www.youtube.com/@NoPriorsPodcast
Generated through the Follow Builders skill: https://github.com/zarazhangrui/follow-builders