---
name: markdownto-guided-narration
description: Author spoken tours of Markdown documents or webpages with deterministic passage and image targets.
---

# Markdown To guided narration

Read `mdto spec guided-narration` or the adjacent `SPEC.md` before authoring.
Outside a checkout, read the [published spec](https://markdownto.ai/specs/guided-narration.md).
The spec and fixtures define conformance; the editorial advice here adapts to the user's brief.

## Write speech against the source

Set `markdownto: guided-narration@0.1` and `source` to a relative Markdown path or
HTTPS URL. Read or capture the source through an available tool before choosing targets.
The source reference itself never instructs the renderer or validator to fetch a website.
For a third-party webpage, supply a capture or imported source in the reader; ownership
or permission to modify the original website is not required for this separate manuscript.

Each paragraph is one spoken beat of 1–1200 characters, with exactly one target field.
Use `##` for chapter navigation; headings are not spoken. Write transitions in the speech.
Do not put lists, tables, HTML, blockquotes, or stage directions in the manuscript.
Do not invent timestamps, CSS selectors, block IDs, or automatic summarization instructions.

Prefer `[target-quote:: exact passage]` copied from the source. Matching normalizes
whitespace but remains case-sensitive. If the quote appears in multiple blocks, choose a
more distinctive quote or add `[target-prefix:: previous block ending]` and/or
`[target-suffix:: next block beginning]`. These refer to neighboring blocks, not context
inside the matched block. Use `[target:: #existing-id]` only for a known captured block ID.
Keep the field in the spoken paragraph; a blank line starts another beat.

To explain an image, target its captured caption or alt text, or its existing captured
block ID. Say what to look at and explain what it means. Describe only details the source
actually shows. The embedded reader supports bundled demo illustrations and embedded raster
images; a remote image URL is not automatically fetched. Inspect the rendered capture to
confirm that the intended picture is present, not merely its caption.

## Make the tour worth hearing

Choose an audience and a purpose from the brief. Explain relationships and consequences
rather than reading every visible sentence aloud. Keep each beat tied to one useful visual
stop. Give the listener enough context to understand a diagram without pretending a conceptual
illustration is historical evidence. Distinguish sourced facts from analogies and interpretation.
Use shorter paragraphs when a long explanation moves attention across several elements.

## Resolve before delivery

Validate the manuscript with `mdto validate tour.md`. Read `mdto guided-narration --help`
for the current inspect and resolve commands. Resolution takes an explicit JSON capture:
`{ "source": "./article.md", "blocks": [{ "id": "intro", "text": "Exact source passage" }] }`.
The capture's source string must match the manuscript exactly; IDs must be unique and nonempty.
Resolve every target against that capture and fix missing or ambiguous matches before playback.
The reader's Download capture for agents action exports its actual blocks for this purpose.

## Record it, once the targets resolve

`mdto guided-narration estimate` counts the tour offline: characters, estimated duration, and
estimated cost per beat, plus what producing right now would actually buy against the cache.
Nothing is spent and no provider is called, so run it freely and read the *incremental* figure —
it is the one that answers "what does this edit cost".

`mdto guided-narration produce` records one track per beat. It prints the estimate every time and
stops until `--yes` acknowledges it, so there is no accidental spend; `--dry-run` shows the plan
and stops regardless. Choose the voice with `--voice` and the speed with `--pace`: the manuscript
carries neither, on purpose. Use `--provider mock --voice mock-narrator` to exercise the whole
path offline before spending anything, and `--beat N` or `--beats 2-5` to re-record a passage you
rewrote. Identical beats are bought once, and an interrupted run resumes from its cache.

Output lands in `guided-narration/` beside the manuscript, in the version-addressed shape the
agentsfs Hub reads: a joined MP3, a receipt, an audio index naming each beat's recording, the
per-beat files, and a pointer written last. The paths inside the pointer and receipt are
repository-relative. If the manuscript lives inside an embedded agentsfs, that directory's prefix
is stripped when it publishes, so pass `--source-path <path-relative-to-the-agentsfs-root>` —
otherwise the Hub refuses the pointer silently and simply shows no recording.

A host that renders the reader can hand it that recording instead of a speech service. Put a
`recording` inside `guided-restore`'s `saved` object — `{version: 1, voice, audio: [{text,
audioBase64 | url, mimeType, durationMs}]}`, one entry per beat, keyed by the beat's
whitespace-collapsed narration — and the reader plays it and asks no provider for anything.
`<basename>.audio.json` (`guided-narration-audio@0.1`) is the natural source: its `text` is
already collapsed and each beat's `file` becomes an entry's `url`. Prefer `url` over
`audioBase64` for anything long, and let the reader page's own `connect-src` reach the host so
it can fetch those files. A beat the recording misses is read by the computer voice, so a
partial recording still helps; an invalid one is refused with `guided-recording-refused` and
the reader carries on without it.

Render and inspect the tour, including any image stops. Confirm the actual speech, target
order, and chapter navigation. The runtime speaks authored paragraphs without rewriting them.
The hosted playground uses Hub-authenticated Gemini when available and computer voice otherwise;
validation and rendering do not generate audio. Do not claim cloud playback was verified from
conformance checks alone. For a local checkout, use `node site/tools/preview.mjs 4382` so
authenticated audio can reach the existing service through the fixed local relay.
