5.1 KiB
name, description, agent_created
| name | description | agent_created |
|---|---|---|
| html-to-docx-cn-gotchas | 用腾讯文档 tencent-docx 插件(tdoc-orchestrator → design-token / doc-typeset / html-review → html-to-docx)把现成的 Markdown 内容做成中文 Word 文档时,绕开三个已实测确认的坑:① Windows 下官方 setup 脚本依赖装不上(bin/python vs Scripts/python.exe);② 表格单元格内的 ASCII 空格被自动加宽成双空格;③ <footer class="doc-footer"> 页脚签名在转换后被丢弃。当用户说「做成 Word / 排版成 docx / 试卷排版 / 文档美化」且内容是中文时使用。 | true |
中文内容 → Word(tencent-docx 插件)实操要点
何时用
- 用户给了现成文本 / Markdown,要求「做成 Word、排版规范一点、出 docx」。
- 走腾讯文档内置插件链路:
tencent-docx→tdoc-orchestrator(Stage 0 判full_pipeline或beautify_only)→ S1doc-writer/ S2doc-formatter(design-token → doc-typeset → html-review)→ S3doc-converter(html-to-docx)。
路由前置:先
tencent-docs-routing定文件类型;「从零生成 docx」→tencent-docx;「改已有 docx」→tencent-local-office-edit。
坑 1 — Windows 下依赖装不上(必踩)
<plugin_root>/scripts/wb/local/setup-html-to-docx.sh 用 $VENV_DIR/bin/python 定位解释器,
而 Windows 的 venv 实际是 $VENV_DIR/Scripts/python.exe。结果:脚本打印「Installing …」后因
uv pip install --python <不存在的路径> 失败被 set -e 静默中断(stderr 被重定向),
site-packages 里只剩 _virtualenv.*。
绕过(不要去改插件内的 vendor 脚本):
export PATH="$HOME/.local/bin:$PATH" # uv 若已在 ~/.local/bin
VENV="$HOME/.venv-html-to-docx"
PY="$VENV/Scripts/python.exe" # ← 关键:Windows 是 Scripts/
REQ="<plugin_root>/skills/html-to-docx/scripts/requirements.txt"
uv pip install --python "$PY" --only-binary=:all: -r "$REQ"
for m in docx bs4 lxml httpx PIL click; do "$PY" -c "import $m" || echo "MISSING $m"; done
htmldocx这个模块名在html-for-docx>=1.2里不存在,import 失败属正常,不是缺依赖。
转换调用(cwd 必须是该 skill 的 scripts/ 目录):
cd "<plugin_root>/skills/html-to-docx/scripts"
"$PY" -m html_to_docx convert in.html -o out.docx \
--page-size A4 --orientation portrait \
--margin-top 2.54 --margin-bottom 2.54 --margin-left 3.17 --margin-right 3.17
坑 2 — 表格单元格空格被加倍(pangu 自动加宽)
转换器会在 CJK ↔ 拉丁/数字 边界自动补一个空格(类似盘古之白),
所以单元格里原本写着的 ASCII 空格会变成两个:10 题、x·sin x 是偶函数。
段落不受影响(段落路径不加宽),只有 <td> / <th> / <caption> 会。
写法规则(只针对表格单元格与 caption):
| 场景 | 不要写 | 要写 | 转出效果 |
|---|---|---|---|
| 中文夹数字 | 10 题(单选 8 题) |
10题(单选8题) |
10 题(单选 8 题) |
| 中文夹公式 | x·sin x 是偶函数 |
x·sin x是偶函数 |
x·sin x 是偶函数 |
| 日期接中文 | 16:13 至 2026-… |
16:13至2026-… |
16:13 至 2026-… |
| 拉丁之间 | — | 8×10 或 8 × 10 均可 |
不受影响 |
- 不要用
/ / 全角空格去「保留」空隙: 、\u3000会被两侧再补空格,更乱。 - 结论:单元格里 中文与数字/字母相邻处一律不留 ASCII 空格,交给转换器补;拉丁-拉丁之间的空格可保留单份。
坑 3 — <footer class="doc-footer"> 被丢弃
同时用 @page { @bottom-center { content: counter(page) " / " counter(pages); } } 声明页码时,
docx 的 footer part 被页码占用,<footer class="doc-footer"> 里的内容(如署名)会整个丢失。
绕过:署名/落款改用正文末尾的普通段落,保留样式类:
<p class="doc-footer">Powered by 如风过境</p>
(@page 只留页码;不要指望 <footer> 元素落进 docx 页脚。)
交付前必做校验
PY="$HOME/.venv-html-to-docx/Scripts/python.exe"
"$PY" -c "
import docx; d=docx.Document(r'<out.docx>')
print('paras',len(d.paragraphs),'tables',len(d.tables))
for tb in d.tables:
for r in tb.rows: print(' | '.join(c.text for c in r.cells))
print('尾段:', repr(d.paragraphs[-1].text))
"
逐项确认:表格文本无「双空格」、末段签名在、页脚有页码域(sections[0].footer.paragraphs 里是 tab + 域)。
其他既有约定(来自 doc-typeset / base prompt)
- 全部样式走 CSS 变量(
var(--*)),禁裸值;html-review 会因裸间距扣分(DT-03)。 - 禁
display:grid、<dl>;结构化/对齐内容一律<table>;每张表要有<thead>+<tbody>。 - 无封面就直接用一个
<section role="body">,不要凭空造封面页。 - 若同时给块级元素设了
border-left,文字开头必须加 两个不折叠空格。 - html-review 只修 一次,改完直接输出,不复查循环。