Files
su-planning-office/skills/html-to-docx-cn-gotchas/SKILL.md
T

5.1 KiB
Raw Blame History

name, description, agent_created
name description agent_created
html-to-docx-cn-gotchas 用腾讯文档 tencent-docx 插件(tdoc-orchestrator → design-token / doc-typeset / html-review → html-to-docx)把现成的 Markdown 内容做成中文 Word 文档时,绕开三个已实测确认的坑:① Windows 下官方 setup 脚本依赖装不上(bin/python vs Scripts/python.exe);② 表格单元格内的 ASCII 空格被自动加宽成双空格;③ <footer class="doc-footer"> 页脚签名在转换后被丢弃。当用户说「做成 Word / 排版成 docx / 试卷排版 / 文档美化」且内容是中文时使用。 true

中文内容 → Word(tencent-docx 插件)实操要点

何时用

  • 用户给了现成文本 / Markdown,要求「做成 Word、排版规范一点、出 docx」。
  • 走腾讯文档内置插件链路:tencent-docx → tdoc-orchestrator(Stage 0 判 full_pipeline 或 beautify_only)→ S1 doc-writer / S2 doc-formatter(design-token → doc-typeset → html-review)→ S3 doc-converter(html-to-docx)。

路由前置:先 tencent-docs-routing 定文件类型;「从零生成 docx」→ tencent-docx;「改已有 docx」→ tencent-local-office-edit。

坑 1 — Windows 下依赖装不上(必踩)

<plugin_root>/scripts/wb/local/setup-html-to-docx.sh 用 $VENV_DIR/bin/python 定位解释器, 而 Windows 的 venv 实际是 $VENV_DIR/Scripts/python.exe。结果:脚本打印「Installing …」后因 uv pip install --python <不存在的路径> 失败被 set -e 静默中断(stderr 被重定向), site-packages 里只剩 _virtualenv.*。

绕过(不要去改插件内的 vendor 脚本):

export PATH="$HOME/.local/bin:$PATH"          # uv 若已在 ~/.local/bin
VENV="$HOME/.venv-html-to-docx"
PY="$VENV/Scripts/python.exe"                  # ← 关键:Windows 是 Scripts/
REQ="<plugin_root>/skills/html-to-docx/scripts/requirements.txt"
uv pip install --python "$PY" --only-binary=:all: -r "$REQ"
for m in docx bs4 lxml httpx PIL click; do "$PY" -c "import $m" || echo "MISSING $m"; done

htmldocx 这个模块名在 html-for-docx>=1.2 里不存在,import 失败属正常,不是缺依赖。

转换调用(cwd 必须是该 skill 的 scripts/ 目录):

cd "<plugin_root>/skills/html-to-docx/scripts"
"$PY" -m html_to_docx convert in.html -o out.docx \
  --page-size A4 --orientation portrait \
  --margin-top 2.54 --margin-bottom 2.54 --margin-left 3.17 --margin-right 3.17

坑 2 — 表格单元格空格被加倍(pangu 自动加宽)

转换器会在 CJK ↔ 拉丁/数字 边界自动补一个空格(类似盘古之白), 所以单元格里原本写着的 ASCII 空格会变成两个:10 题、x·sin x 是偶函数。 段落不受影响(段落路径不加宽),只有 <td> / <th> / <caption> 会。

写法规则(只针对表格单元格与 caption):

场景 不要写 要写 转出效果
中文夹数字 10 题(单选 8 题) 10题(单选8题) 10 题(单选 8 题)
中文夹公式 x·sin x 是偶函数 x·sin x是偶函数 x·sin x 是偶函数
日期接中文 16:13 至 2026-… 16:13至2026-… 16:13 至 2026-…
拉丁之间 — 8×10 或 8 × 10 均可 不受影响
  • 不要用 &nbsp; / &#8201; / 全角空格去「保留」空隙:&#8201;、\u3000 会被两侧再补空格,更乱。
  • 结论:单元格里 中文与数字/字母相邻处一律不留 ASCII 空格,交给转换器补;拉丁-拉丁之间的空格可保留单份。

同时用 @page { @bottom-center { content: counter(page) " / " counter(pages); } } 声明页码时, docx 的 footer part 被页码占用,<footer class="doc-footer"> 里的内容(如署名)会整个丢失。

绕过:署名/落款改用正文末尾的普通段落,保留样式类:

<p class="doc-footer">Powered by 如风过境</p>

(@page 只留页码;不要指望 <footer> 元素落进 docx 页脚。)

交付前必做校验

PY="$HOME/.venv-html-to-docx/Scripts/python.exe"
"$PY" -c "
import docx; d=docx.Document(r'<out.docx>')
print('paras',len(d.paragraphs),'tables',len(d.tables))
for tb in d.tables:
    for r in tb.rows: print(' | '.join(c.text for c in r.cells))
print('尾段:', repr(d.paragraphs[-1].text))
"

逐项确认:表格文本无「双空格」、末段签名在、页脚有页码域(sections[0].footer.paragraphs 里是 tab + 域)。

其他既有约定(来自 doc-typeset / base prompt)

  • 全部样式走 CSS 变量(var(--*)),禁裸值;html-review 会因裸间距扣分(DT-03)。
  • 禁 display:grid、<dl>;结构化/对齐内容一律 <table>;每张表要有 <thead> + <tbody>。
  • 无封面就直接用一个 <section role="body">,不要凭空造封面页。
  • 若同时给块级元素设了 border-left,文字开头必须加 &nbsp;&nbsp; 两个不折叠空格。
  • html-review 只修 一次,改完直接输出,不复查循环。