我们发现了什么
Show HN:ClientCoded – AI代理的QA平台。我们为AI代理构建了一个QA平台。
- 来源:Hacker News(发现于 2026-09-14)
- 证据等级:D · 发现产品或需求信号,暂未获得可核验的商业证据。
- 商业模式:待核验
- 主题:AI Agent
- 初筛评分:16.7/100 · 收录 1 次
证据,比故事更重要。
规则清洗与初筛,未经人工商业核验。原文语境、实际客户和付费情况仍需自行验证。
引用与数字披露
来源类型(原作者自述/第三方测算/媒体转引)需采集端标注,本版尚未落字段。
- 作者
- 未标注
- 抓取日期
- 来源类型
- 未标注
- 币种
- 未标注
- 口径
- 未标注
- 披露主体
- 未标注
- 披露日期
- 未标注
中文辅助译文(全文)
嗨,HN。我们构建了一个面向 AI 智能体的 QA 平台。我的联合创始人在 Veeva Systems 花了 6 年时间为受监管的软件搭建 QA 框架。我则在销售领域深耕 5 年,帮助 AI 原生初创公司实现增长。我们都亲眼目睹了企业在部署智能体时,由于缺乏结构化的测试和监控手段而面临的种种问题。在我们搭建平台前采访过的工程团队表示,他们目前只进行人工抽查,或仅用 LLM 作为裁判来检查差异。我们为 35 个平台(Salesforce、Jira、Stripe、Zendesk、Datadog 以及其他 29 个)生成合成测试环境。每个环境包含约 200 个对抗性查询,覆盖 7 个类别并附带计算得出的真实标签:干净查询、模糊问题、多步操作、范围边界测试、矛盾输入、无效假设以及依赖上下文的问题。真实标签最初通过对合成数据集运行 SQL 计算得出,并非由 LLM 凭空猜测。我们还为预生产对话型智能体提供对抗性测试。系统会生成不同的人物角色,将智能体推离预设脚本,并从 10 个维度对其通过/失败进行评分,让你能在真实客户场景中看到会发生什么。我们的生产监控也只需一个 Webhook,即可对每一段对话进行实时评估。我们是两位自筹资金的创始人,非常期待听到你对我们的方法的反馈!https://clientcoded.com/
译文由上游机器翻译生成,可能有误;判断请以英文原文为准。
英文原文(来源本站未改写)
Hi HN.We built a QA platform for AI agents.My cofounder spent 6 years at Veeva Systems building QA frameworks for regulated software.I spent 5 years in sales helping AI-native startups with their growth.We've both seen the issues companies face when deploying agents without a structured way of testing and monitoring them.The engineering teams we interviewed prior to building said they are doing manual spot-checks or just using LLM as a judge to check discrepancies.We generate synthetic test environments for 35 platforms (Salesforce, Jira, Stripe, Zendesk, Datadog, and 29 more).
Each environment has ~200 adversarial queries with computed ground truth across 7 categories: clean lookups, ambiguous questions, multi-step operations, scope boundary tests, contradictory inputs, invalid assumptions, and context-dependent questions.The ground truth is initially computed by running SQL against the synthetic dataset so it's not just guessed by an LLM.We also do adversarial testing for pre-production conversational agents.The system generates different personas that push agents off-script and scores pass/fail across 10 dimensions so you can see what happens in realistic customer scenarios.
Our production monitoring is just one webhook with every conversation evaluated in real time as well.We are two founders bootstrapping and we would love feedback on our approach!https://clientcoded.com/
这条还缺什么证据?
下面每条都由本条已有字段推出(等级、理由、商业模式、来源次数、是否演示), 本站不生成推测性结论;通用验证方法放在方法论页。
- 可核验的收入或付费证据查官网定价页与付费口径;第三方数据源(如 GetLatka)只作旁证,需标注来源与时点。
- 商业模式未定确认按席位/按用量/授权还是开源托管版收费;开源项目另查 LICENSE 与是否存在付费版。
- 只有单一来源找一手站点或其他渠道是否重复出现同一产品;社区热帖数量不等于商业进展。
通用验证清单(谁有这个问题/谁愿意付费/一个人能交付哪一小步)见我们的筛选方法。