SightDiff – before/after visual proof of what your AI agent changed
Show HN: SightDiff – before/after visual proof of what your AI agent changed
我们发现了什么
Show HN: SightDiff – before/after visual proof of what your AI agent changed.
- 来源:Hacker News(发现于 2026-08-14)
- 证据等级:C · 存在定价或订阅线索;有收费设计不等于已有收入。
- 商业模式:待核验
- 主题:AI Agent
- 初筛评分:14.4/100 · 收录 1 次
证据,比故事更重要。
规则清洗与初筛,未经人工商业核验。原文语境、实际客户和付费情况仍需自行验证。
引用与数字披露
来源类型(原作者自述/第三方测算/媒体转引)需采集端标注,本版尚未落字段。
- 作者
- 未标注
- 抓取日期
- 来源类型
- 未标注
- 币种
- 未标注
- 口径
- 未标注
- 披露主体
- 未标注
- 披露日期
- 未标注
中文辅助译文(全文)
● 本地 · pre-commit · 任意智能体
你的 AI 智能体说完成了。看看真正发生了什么变化。
在 git commit 之前,本地渲染出其所触及的每一个页面和状态的前后对比。被改动的界面会被标记出来;未改动的界面会被验证为完全一致。
获取早期访问 观看 60 秒演示
66% 的开发者表示,他们最大的挫败感是 AI 编写的代码几乎正确,但又不完全正确。— Stack Overflow 开发者调查 2025
智能体被要求改动一个页面——证明页却发现它改动了两个。sightdiff snap → 智能体运行 → sightdiff check
工作原理
两条命令,包裹住你的智能体所做的任何事情。
无需集成,无需云端、无需 CI 流水线。一个正在运行的开发服务器,和一份列出你所关心的页面与状态的精简配置——SightDiff 甚至可以通过爬取你的应用来替你写出这份配置。
1
捕获基线
在你发出提示之前,一条命令会以像素级精度截取每一个已配置页面和状态的截图——包含元素状态、需鉴权访问的视图、经过遮罩的动态内容。
$ sightdiff snap
2
让智能体工作
Claude Code、Cursor、Copilot——或者匆忙中的人。SightDiff 并不挂接到智能体内部,这恰恰是它能与所有这些工具配合的原因。
> 实现新的筛选器……
3
在提交前检查
一切都会被重新渲染并按像素进行 diff:一个汇总页打开,被改动的界面排在前面并高亮出差异区域,未改动的界面被验证为完全一致。当存在标记项时退出码非零——可作为 pre-commit 关卡使用。
$ sightdiff check
关键之处
你没检查的那个页面,就是会出问题的那一个。
在上面演示中,智能体被要求为一个页面添加筛选器——它完成了。但它对一个共享 CSS 类的修改也影响到了仪表板——一个没人要求改动的页面。这正是你即将提交的那个回归。
已改动 特性开关 按你要求的改动
已改动 仪表板 你没要求的那一处
完全一致 审计日志 经核验,而非假设
完全一致 设置 经核验,而非假设
演示中那份真实的证明页——前后对比并高亮改动区域,未改动的页面折叠为一行已核验记录。
为什么要单独做一个工具
智能体在给自己的作业打分。
现代智能体可以打开浏览器并"验证"自己的工作。有时它们确实做到了。有时它们会跳过、误读,或者悄悄篡改证据——而你在三天后从用户的截图中才发现。
SightDiff 在智能体触及不到的地方生成证明:由一个独立的本地进程从你实际运行的应用中捕获,与 git 状态绑定。这份证明页是给你看的,而不是给模型的。
"我需要的工具,要让智能体能够清楚地向我展示它们的工作,同时尽量减少它们在所完成事项上作弊的机会。" — Simon Willison,在抓到智能体编辑演示输出而不是实际运行之后 · simonwillison.net
早期访问
两种参与方式。其中一种让我保持诚实。
候补名单
$0 你的邮箱,我的构建进展
beta 发布时第一时间获得访问
偶尔的构建日志更新——不发送垃圾邮件,可随时退订
督促我把它做出来 创始用户
$10 /月 · 随时可退款
仅从你真正能够运行的第一个版本开始计费
直接联系我——你的工作流将决定 v1 的形态
创始价格终身有效
发送一封邮件即可取消或退款,无需说明理由
成为创始用户 → 这是一次对一个尚未完全成形的工具的押注。我是一名在职工程师,为自己日常的工作流打造它;付费是你让我加快进度的方式。
常见问题
你应该问的问题。
我的代码或界面是否曾离开过我的机器? 没有。捕获、渲染和 diff 都在本地运行。不会上传任何内容——这也是当你的 CI、VPN 或法务部门说不时,它依然能用的理由。
它支持哪些智能体? 全部都支持——Claude Code、Cursor、Copilot,或者匆忙中的人。SightDiff 观察的是你的应用,而不是智能体本身,因此无需集成,智能体也无法作弊。
运行它需要什么? 一个本地开发服务器,和一份列出需要捕获的页面与状态的小型配置——sightdiff discover 可以通过爬取你的应用来为你生成。如果浏览器能渲染它,SightDiff 就能截图它;beta 阶段首先打磨 React 配合 Vite 或 Next.js,你在候补回复中说明的技术栈将决定后续的优先顺序。
它与 Chromatic 或 Percy 有什么不同? 那些都是出色的 CI 时工具:它们在你提交并推送之后,在云端检查 pull request。SightDiff 工作在此之前的那个时刻——本地、工作区脏、智能体刚刚停下来、由你决定是否信任此次改动。
它目前还不能做什么? 今天它会对你已配置的每个界面进行截图,并依靠像素 diff 来标记——它还不能将你的代码 diff 映射到仅受影响的界面,基线目前通过 snap 获取而非在后台持续采集。这两项都在路线图上,由创始用户决定先后顺序。
像你亲眼看过那样提交——因为你确实看过。
获取早期访问
译文由上游机器翻译生成,可能有误;判断请以英文原文为准。
英文原文(来源本站未改写)
● Local · pre-commit · any agent
Your AI agent says done. See what actually changed.
Before/after proof of every page and state your agent touched — rendered locally, before you git commit . Changed surfaces get flagged; untouched ones are verified identical.
Get early access Watch the 60-second demo
66% of developers say their top frustration is AI code that's almost right, but not quite. — Stack Overflow Developer Survey 2025
The agent was asked to change one page — the proof sheet caught it changing two. sightdiff snap → agent works → sightdiff check
How it works
Two commands, wrapped around anything your agent does.
No integration, no cloud, no CI pipeline. A running dev server and a tiny config listing the pages and states you care about — SightDiff can even write that config for you by crawling your app.
1
Snap a baseline
Before you prompt, one command captures pixel-perfect screenshots of every configured page and state — element states, auth-gated views, masked dynamic content included.
$ sightdiff snap
2
Let the agent work
Claude Code, Cursor, Copilot — or a human in a hurry. SightDiff doesn't hook into the agent, which is exactly why it works with all of them.
> implement the new filter…
3
Check before you commit
Everything is re-rendered and pixel-diffed. One sheet opens: changed surfaces first with highlighted regions, untouched ones verified identical. Exits non-zero when flagged — usable as a pre-commit gate .
$ sightdiff check
The part that matters
The page you didn't check is the one that breaks.
In the demo above, the agent was asked to add a filter to one page — and delivered. But its edit to a shared CSS class also shifted the dashboard, a page nobody asked about. That's the regression you'd have committed.
CHANGED Feature flags the change you asked for
CHANGED Dashboard the one you didn't
IDENTICAL Audit log verified, not assumed
IDENTICAL Settings verified, not assumed
The actual proof sheet from the demo — before/after with changed regions highlighted, unchanged pages collapsed to a single verified line.
Why a separate tool
Agents grade their own homework.
Modern agents can open a browser and “verify” their work. Sometimes they do. Sometimes they skip it, misread it, or quietly edit the evidence — and you find out from a user screenshot three days later.
SightDiff renders proof outside the agent's reach : captured by a separate local process from your actual running app, keyed to git state. This sheet is for you , not for the model.
“I need tools that allow agents to clearly demonstrate their work to me, while minimizing the opportunities for them to cheat about what they've done.” — Simon Willison, after catching agents editing demo output instead of running it · simonwillison.net
Early access
Two ways in. One of them keeps me honest.
Waitlist
$0 your email, my build updates
First access when the beta ships
Occasional build-log updates — no spam, unsubscribe anytime
Forces me to build it Founding user
$10 /mo · refundable anytime
Charged only from the first release you can actually run
Direct line to me — your workflow shapes v1
Founding price locked for life
Cancel or refund with one email, no questions asked
Become a founding user → This is a bet on a tool that doesn't fully exist yet. I'm a working engineer building it for my own daily loop; paying is how you tell me to hurry.
FAQ
Questions you should ask.
Does my code or UI ever leave my machine? No. Capture, rendering, and diffing all run locally. Nothing is uploaded — which is also why it works when your CI, your VPN, or your legal team says no.
Which agents does it work with? Any of them — Claude Code, Cursor, Copilot, or a human in a hurry. SightDiff watches your app, not the agent, so there's nothing to integrate and nothing for the agent to game.
What do I need to run it? A local dev server and a small config listing the pages and states to capture — sightdiff discover writes it for you by crawling your app. If a browser can render it, SightDiff can shoot it; the beta's polish targets React with Vite or Next.js first, and your waitlist reply telling me your stack decides what comes next.
How is this different from Chromatic or Percy? Those are excellent CI-time tools: they check pull requests in the cloud, after you commit and push. SightDiff lives at the moment before — local, dirty working tree, agent just stopped, you deciding whether to trust the change.
What doesn't it do yet? Today it shoots every surface you've configured and lets the pixel diff do the flagging — it doesn't yet map your code diff to only the affected surfaces, and baselines are taken with snap rather than continuously in the background. Both are on the roadmap, and founding users decide the order.
Commit like you've seen it — because you have.
Get early access
这条还缺什么证据?
下面每条都由本条已有字段推出(等级、理由、商业模式、来源次数、是否演示), 本站不生成推测性结论;通用验证方法放在方法论页。
- 可核验的收入或付费证据查官网定价页与付费口径;第三方数据源(如 GetLatka)只作旁证,需标注来源与时点。
- 商业模式未定确认按席位/按用量/授权还是开源托管版收费;开源项目另查 LICENSE 与是否存在付费版。
- 只有单一来源找一手站点或其他渠道是否重复出现同一产品;社区热帖数量不等于商业进展。
通用验证清单(谁有这个问题/谁愿意付费/一个人能交付哪一小步)见我们的筛选方法。