01 / THE SIGNAL

我们发现了什么

Vespper 是一个 MCP,让 AI 智能体能够高效地编辑 Word 文档,由我们微调的模型驱动。在此处查看产品工作原理概览:https://youtu.be/odKxsgPjzzw 我们在为制药公司构建 AI 文档编辑器一年后,开始着手解决这个问题。

  • 来源:Hacker News(发现于 2026-09-29)
  • 证据等级:D · 发现产品或需求信号,暂未获得可核验的商业证据。
  • 商业模式:待核验
  • 主题:独立产品
  • 初筛评分:20.5/100 · 收录 1 次
#独立开发#待验证#产品发现
02 / SOURCE & EVIDENCE

证据,比故事更重要。

发现产品或需求信号,暂未获得可核验的商业证据。

规则清洗与初筛,未经人工商业核验。原文语境、实际客户和付费情况仍需自行验证。

引用与数字披露

来源类型(原作者自述/第三方测算/媒体转引)需采集端标注,本版尚未落字段。

短句引用
作者
未标注
抓取日期
来源类型
未标注
数字口径
币种
未标注
口径
未标注
披露主体
未标注
披露日期
未标注

本条正文译文未完成(采集端 translation.body_ok=false),此处只展示英文原文。

英文原文(来源本站未改写)

Hey HN!We're Dudu and Topaz from Vespper ( https://vespper.com ).Vespper is an MCP that lets AI agents efficiently edit Word documents, powered by our fine-tuned model.It's currently 3× faster, 2× cheaper and more accurate than the closest alternative.Check out an overview of how the product works here: https://youtu.be/odKxsgPjzzw We came to work on this problem after spending a year building an AI document editor for pharma companies.Before that, Topaz(myself) was a senior SWE at Snyk, working on distributed systems, and Dudu was a deep learning engineer at Viz.ai, building computer vision models for stroke detection.Our editor helped pharma companies generate regulatory documents (e.g.

CSRs) to speed up their submissions.Initially, the output was Markdown, displayed in a WYSIWYG editor.However, users preferred working with their own Word templates.That's when the problems began.AI agents aren't great at editing Word documents.A Word document is a zip file of verbose XML files following the OOXML spec.Even "small" changes require backflips, for example: adding a numbered list requires creating an entry in numbering.xml with a fresh ID and linking it back in document.xml, bolding a sentence requires splitting it into 3+ run elements.The list goes on.

This makes editing the zip directly (unzip + grep + sed) a bad idea for agents because they burn a lot of time + tokens on these mechanics.In practice, today's tooling falls into roughly three categories.You can let the agent write code against low-level libraries like python-docx or the Open XML SDK, you can give it an MCP with opinionated editing tools (SuperDoc, Office CLI, Adeu, etc), or you can round-trip the file through Markdown/HTML with something like pandoc/mammoth.js.None of them really work.

The first two categories still burn the agent's context on Word mechanics instead of the task at hand (MCPs also introduce a new DSL to learn), and the third is very lossy (pandoc/mammoth.js/etc don't preserve enough fidelity).From firsthand experience, these problems hurt performance in downstream tasks.When we tried having our agent fill large documents, things broke quickly.The context window was already packed with customer data (files, user context, global rules, etc.), and the agent burned tokens + time on exploring the document and debugging failed edits.

Filling a single CSR (clinical study report) took ~50 minutes, and the result was bad (missed fields/sections, broken styling, etc.).Harvey.ai's team reached similar conclusions: https://www.harvey.ai/blog/building-an-agent-for-complex-doc...That's when we shifted our focus.We designed an MCP that lets agents edit Word docs as if they were editing HTML.The agent receives HTML, makes find-and-replace edits, and we reconcile those edits back into the original .docx file.We picked HTML over Markdown because it's structurally much closer to OOXML and because CSS associates styles with elements roughly the way OOXML does.W

出处https://vespper.com/blog/launching-vespper-docx-mcp抓取日期 · 采集源 Hacker News

03 / EVIDENCE GAPS

这条还缺什么证据?

下面每条都由本条已有字段推出(等级、理由、商业模式、来源次数、是否演示), 本站不生成推测性结论;通用验证方法放在方法论页。

  • 可核验的收入或付费证据查官网定价页与付费口径;第三方数据源(如 GetLatka)只作旁证,需标注来源与时点。
  • 商业模式未定确认按席位/按用量/授权还是开源托管版收费;开源项目另查 LICENSE 与是否存在付费版。
  • 只有单一来源找一手站点或其他渠道是否重复出现同一产品;社区热帖数量不等于商业进展。

通用验证清单(谁有这个问题/谁愿意付费/一个人能交付哪一小步)见我们的筛选方法。

04 / SIGNAL HISTORY

发现时间线