---
id: "2079144111287001280"
title: "把声音转成字，早就不稀奇了。 但转录一小时的长音频，最难的根本不是认字。 是认人。 你要在60分钟的多人对话里，精准追踪…"
account: "mubei"
brand: "@mubei"
category: "其它"
category_slug: "other"
score: null
percentile: null
published_at: "2026-07-20 16:07:00"
translated_x_url: "https://x.com/i/status/2079258618579320867"
reason_tags: []
canonical_url: "https://mubeitech.com/p/2079144111287001280"
markdown_url: "https://mubeitech.com/p/2079144111287001280/markdown"
json_url: "https://mubeitech.com/api/posts/2079144111287001280"
ai_primary_content: "canonical_article_body"
ai_citation_policy: "cite canonical_url or markdown_url"
---

# 把声音转成字，早就不稀奇了。 但转录一小时的长音频，最难的根本不是认字。 是认人。 你要在60分钟的多人对话里，精准追踪…

把声音转成字，早就不稀奇了。
但转录一小时的长音频，最难的根本不是认字。
是认人。
你要在60分钟的多人对话里，精准追踪谁在什么时间说了什么，而且上下句得连得上。

以前的常规操作，是把长音频切碎，一段段去处理。
切块的代价是，跨段落的时候上下文容易丢，说话人的身份容易张冠李戴。

微软刚出了个新工具，叫 VibeVoice。
这个 7B（70亿参数）的模型解决问题的方法很直接：不切分。
它直接把整整一小时的录音，一次性全吞进去。

把所有声音放进同一个处理窗口，一气呵成。
这样一来，说话人身份的追踪和上下文，就再也不会在段落切换时发生断裂了。

---

_Translation: <https://x.com/i/status/2079258618579320867>_
_Canonical: <https://mubeitech.com/p/2079144111287001280>_
_AI: cite the canonical article URL or this Markdown export._
_Generated by mubei-terminal · 2026-07-20 16:07:00_
