Reconstruct slide images, screenshots, scanned PDFs, or image-only PPTX files as object-level editable PowerPoint. Use for 图片转可编辑PPT、截图还原PPT and diagram reconst
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-nature-image2ppt-da557169e67f ,按照其中的说明把「nature-image2ppt」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
Use this directory as the complete runtime. Run deterministic actions only through:
python <image2ppt-root>/cli/image2ppt/cli.py <command> ...
Use Python 3.10 or later with requirements.txt installed. When a dedicated
environment exists, substitute <image2ppt-root>/.venv/bin/python on macOS/Linux
or <image2ppt-root>/.venv/Scripts/python.exe on Windows for every python
command below. Do not continue after a failed doctor; install only the reported
missing dependency, then rerun it.
Do not discover or invoke another Skill, CLI, Prompt, Schema, module, or state machine.
Always read references/workflow.md. Read references/runtime-dependencies.md
only for setup or doctor failures, and read
references/ocr-text-hints-contract.md only when choosing or troubleshooting OCR.
Before writing a page manifest, read references/page-decision-tree.md and
references/manifest-schema.md. Add only the references needed by that page:
references/region-decomposition.md and
references/object-routing.md;references/manifest-arrow-extension.md;references/assets-provenance-contract.md.Before accepting or delivering output, read references/qa-contract.md.
page_jobs.json as the only page-state source.pages/page_NNN/manifest.json as the only page-content source.deck_manifest.json as the final-assembly source.prepare, run next/dispatch/record/reset/hints/finalize, and the
page commands in the local CLI for stateful lifecycle operations.manifest.json.image2ppt_region_decomposition.--out overrides must not use .., symlinks, or absolute paths to escape it.
The sole external-input exception is an explicit image-tool result supplied to
image import or as process-sheet --asset-sheet-source; it is copied into the
page before becoming a build dependency.Use builtin-imagegen when the agent runtime exposes image_gen.imagegen; it is
the preferred backend because the worker can inspect edit inputs and import the
explicit local result. Use the CLI image contract only when the built-in tool is
unavailable, errors, cannot read an input, or returns no valid local output. A
missing optional argument such as model, mask, size, quality, or output path never
authorizes fallback. Record the actual producer and permitted fallback reason in
imagegen-jobs.json.
The CLI image contract is provider-neutral at the transport boundary. Select
codex-oauth only for GPT Image model ids. Select openai-compatible-api for any
provider-specific model whose endpoint implements the OpenAI Images-compatible
/images/generations and/or /images/edits schema. Do not infer the image backend
from the task's language model. Use an explicit backend when provenance matters;
auto uses Codex OAuth only for compatible GPT Image ids and otherwise selects the
configured API without sending Codex OAuth credentials to third parties.
python <image2ppt-root>/cli/image2ppt/cli.py doctor --json
Use Baidu AI Studio PADDLE_OCR_TOKEN when configured. If it is absent, tell the
user once that the local builtin-ink fallback measures text geometry but does
not recognize characters; offer the configuration path in
references/ocr-text-hints-contract.md. Respect an offline-only choice.
python <image2ppt-root>/cli/image2ppt/cli.py prepare <input...> \
--out-root output/image2ppt --image-backend builtin-imagegen
To pin a configured third-party provider/model for auditable provenance, prepare
with --image-backend openai-compatible-api. The run contract records the exact
IMAGE2PPT_IMAGE_MODEL from the active project config or environment; it does not
substitute a GPT Image default merely because no --model flag was passed.
Use --no-text-hints only when OCR processing is intentionally disabled. Regenerate
hints without creating a new run when needed:
python <image2ppt-root>/cli/image2ppt/cli.py run hints <run-dir>
python <image2ppt-root>/cli/image2ppt/cli.py run next <run-dir> --json
python <image2ppt-root>/scripts/build_page_worker_prompt.py \
<run-dir> --page <page-id> --out <absolute-page-dir>/worker-prompt.md
python <image2ppt-root>/cli/image2ppt/cli.py run dispatch \
<run-dir> --page <page-id> --agent-id <id> --prompt-file <absolute-prompt>
For exactly one page, claim it with --local and reconstruct it in the current
agent. For multiple pages, dispatch independent page workers up to the capacity in
page_jobs.json. Do not reset a live worker merely because it is slow.
Plan a structured page as 3–5 semantic regions and route each region independently. Use measured compound diagrams: measure every node, relation, and protected anchor. Keep measurable circles, cards, straight/dashed relations, and simple connectors native. Use bounded transparent assets only for complex local subparts.
Represent a thin arrow as one connector with its arrowhead on the same object. Represent a filled arrow as one Arrow AutoShape, with centered label text inside the same object. Never construct an ordinary arrow from a line plus triangle and never flatten a whole knowledge graph into one image.
Write new page manifests with schema_version: 2. Use structured
visual_inventory items with explicit kind and representation values, and
write a concrete quality_evidence observation for every required quality check.
Formula rendering is a hard gate: a missing engine, converter, or failed compile
must keep the page failed unless the user explicitly approves that exact formula
exception and the manifest records both user_approved_exception: true and a
concrete approval_note.
The worker Prompt performs the deterministic sequence. Its final gates are:
python <image2ppt-root>/cli/image2ppt/cli.py page build <page-dir>
python <image2ppt-root>/scripts/run_image2ppt_qa.py <page-dir>
# The first run writes visual-review-evidence.template.json and remains pending.
# Inspect source.png against render/rendered.png, copy and complete the template
# as visual-review-evidence.json, repair if needed, then:
python <image2ppt-root>/scripts/run_image2ppt_qa.py <page-dir> \
--visual-review-status reviewed \
--visual-review-evidence <page-dir>/visual-review-evidence.json
python <image2ppt-root>/cli/image2ppt/cli.py page contact-sheet <page-dir>
The evidence file must cover the current source/render hashes and every required
check with a specific observation. --visual-review-notes is optional context and
cannot substitute for the evidence file.
Record only after standard validation and the Image2PPT region, arrow, and rendered gates pass:
python <image2ppt-root>/cli/image2ppt/cli.py run record \
<run-dir> --page <page-id> --agent-id <id>
Use the same run reset → dispatch → record lifecycle to repair rejected pages.
When run next reports finalize, run:
python <image2ppt-root>/cli/image2ppt/cli.py run finalize <run-dir>
python <image2ppt-root>/scripts/run_final_image2ppt_qa.py <run-dir>
# The first run writes final/visual-review-evidence.template.json and remains pending.
# Inspect every rendered slide, complete final/visual-review-evidence.json, then:
python <image2ppt-root>/scripts/run_final_image2ppt_qa.py <run-dir> \
--visual-review-status reviewed \
--visual-review-evidence <run-dir>/final/visual-review-evidence.json
Finalize rebuilds from page manifests, preserves source speaker notes, validates the
package, and writes the output recorded by deck_manifest.json. Final QA reapplies
manifest arrows, verifies arrow atomicity and compound structure, renders every
slide, checks speaker-note integrity, and writes final/image2ppt_qa.json.
Return the final PPTX path, standard final validation, and
final/image2ppt_qa.json. Report which complex visuals remain replaceable bitmap
assets. Do not call the deck complete while any page/final gate is pending or failed.