Files
su-planning-office/skills/html-to-docx-cn-gotchas/SKILL.md
T

96 lines
5.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: html-to-docx-cn-gotchas
description: 用腾讯文档 tencent-docx 插件(tdoc-orchestrator → design-token / doc-typeset / html-review → html-to-docx)把现成的 Markdown 内容做成中文 Word 文档时,绕开三个已实测确认的坑:① Windows 下官方 setup 脚本依赖装不上(bin/python vs Scripts/python.exe);② 表格单元格内的 ASCII 空格被自动加宽成双空格;③ <footer class="doc-footer"> 页脚签名在转换后被丢弃。当用户说「做成 Word / 排版成 docx / 试卷排版 / 文档美化」且内容是中文时使用。
agent_created: true
---
# 中文内容 → Word(tencent-docx 插件)实操要点
## 何时用
- 用户给了现成文本 / Markdown,要求「做成 Word、排版规范一点、出 docx」。
- 走腾讯文档内置插件链路:`tencent-docx` → `tdoc-orchestrator`(Stage 0 判 `full_pipeline` 或 `beautify_only`)→ S1 `doc-writer` / S2 `doc-formatter`(design-token → doc-typeset → html-review)→ S3 `doc-converter`(html-to-docx)。
> 路由前置:先 `tencent-docs-routing` 定文件类型;「从零生成 docx」→ `tencent-docx`;「改已有 docx」→ `tencent-local-office-edit`。
## 坑 1 — Windows 下依赖装不上(必踩)
`<plugin_root>/scripts/wb/local/setup-html-to-docx.sh` 用 `$VENV_DIR/bin/python` 定位解释器,
而 Windows 的 venv 实际是 `$VENV_DIR/Scripts/python.exe`。结果:脚本打印「Installing …」后因
`uv pip install --python <不存在的路径>` 失败被 `set -e` 静默中断(stderr 被重定向),
`site-packages` 里只剩 `_virtualenv.*`。
**绕过**(不要去改插件内的 vendor 脚本):
```bash
export PATH="$HOME/.local/bin:$PATH" # uv 若已在 ~/.local/bin
VENV="$HOME/.venv-html-to-docx"
PY="$VENV/Scripts/python.exe" # ← 关键:Windows 是 Scripts/
REQ="<plugin_root>/skills/html-to-docx/scripts/requirements.txt"
uv pip install --python "$PY" --only-binary=:all: -r "$REQ"
for m in docx bs4 lxml httpx PIL click; do "$PY" -c "import $m" || echo "MISSING $m"; done
```
> `htmldocx` 这个模块名在 `html-for-docx>=1.2` 里不存在,import 失败属正常,不是缺依赖。
转换调用(`cwd` 必须是该 skill 的 `scripts/` 目录):
```bash
cd "<plugin_root>/skills/html-to-docx/scripts"
"$PY" -m html_to_docx convert in.html -o out.docx \
--page-size A4 --orientation portrait \
--margin-top 2.54 --margin-bottom 2.54 --margin-left 3.17 --margin-right 3.17
```
## 坑 2 — 表格单元格空格被加倍(pangu 自动加宽)
转换器会在 **CJK ↔ 拉丁/数字** 边界自动补一个空格(类似盘古之白),
所以单元格里原本写着的 ASCII 空格会变成两个:`10 题`、`x·sin x 是偶函数`。
**段落不受影响**(段落路径不加宽),只有 `<td>` / `<th>` / `<caption>` 会。
**写法规则**(只针对表格单元格与 caption):
| 场景 | 不要写 | 要写 | 转出效果 |
|---|---|---|---|
| 中文夹数字 | `10 题(单选 8 题)` | `10题(单选8题)` | `10 题(单选 8 题)` |
| 中文夹公式 | `x·sin x 是偶函数` | `x·sin x是偶函数` | `x·sin x 是偶函数` |
| 日期接中文 | `16:13 至 2026-…` | `16:13至2026-…` | `16:13 至 2026-…` |
| 拉丁之间 | — | `8×10` 或 `8 × 10` 均可 | 不受影响 |
- **不要**用 `&nbsp;` / `&#8201;` / 全角空格去「保留」空隙:`&#8201;`、`\u3000` 会被两侧再补空格,更乱。
- 结论:单元格里 **中文与数字/字母相邻处一律不留 ASCII 空格,交给转换器补**;拉丁-拉丁之间的空格可保留单份。
## 坑 3 — `<footer class="doc-footer">` 被丢弃
同时用 `@page { @bottom-center { content: counter(page) " / " counter(pages); } }` 声明页码时,
docx 的 footer part 被页码占用,`<footer class="doc-footer">` 里的内容(如署名)会**整个丢失**。
**绕过**:署名/落款改用正文末尾的普通段落,保留样式类:
```html
<p class="doc-footer">Powered by 如风过境</p>
```
(`@page` 只留页码;不要指望 `<footer>` 元素落进 docx 页脚。)
## 交付前必做校验
```bash
PY="$HOME/.venv-html-to-docx/Scripts/python.exe"
"$PY" -c "
import docx; d=docx.Document(r'<out.docx>')
print('paras',len(d.paragraphs),'tables',len(d.tables))
for tb in d.tables:
for r in tb.rows: print(' | '.join(c.text for c in r.cells))
print('尾段:', repr(d.paragraphs[-1].text))
"
```
逐项确认:表格文本无「双空格」、末段签名在、页脚有页码域(`sections[0].footer.paragraphs` 里是 tab + 域)。
## 其他既有约定(来自 doc-typeset / base prompt)
- 全部样式走 CSS 变量(`var(--*)`),禁裸值;html-review 会因裸间距扣分(DT-03)。
- 禁 `display:grid`、`<dl>`;结构化/对齐内容一律 `<table>`;每张表要有 `<thead>` + `<tbody>`。
- 无封面就直接用一个 `<section role="body">`,不要凭空造封面页。
- 若同时给块级元素设了 `border-left`,文字开头必须加 `&nbsp;&nbsp;` 两个不折叠空格。
- html-review 只修 **一次**,改完直接输出,不复查循环。