Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,13 +70,13 @@ Commands registered on the **`docgen`** CLI include:
- **`freeze`** — ``docgen freeze`` builds the **`docgen-gui`** onedir (`pip install 'docgen[packaging]'`). Optional ``--smoke`` runs the binary headless. Do not run a full freeze in routine pytest; set ``DOCGEN_FREEZE_SMOKE=1`` for the optional test.
- **`tts`** — text-to-speech for segment files (OpenAI or xAI `/v1/tts`).
- **`timestamps`** — word/segment timing (`timing.json`). Default engine **`local`** aligns the known narration text against the mp3 offline (ffmpeg silencedetect, no API); **`--engine whisper`** uses OpenAI whisper-1 or xAI `/v1/stt` when `ai.provider` is grok. Both emit the same Whisper-shaped blocks. Failed ffmpeg silencedetect raises `AlignmentError` (empty stderr is not treated as full-span speech). OpenAI whisper-1 word/segment `start`/`end` must be finite JSON numbers (bool/NaN raise `AIError`). Grok `/v1/stt` rejects empty word tokens and inverted `end < start` intervals (`AIError`). Empty ``segments.all`` raises ``TimestampError`` (same as TTS) and does not leave a stale ``timing.json`` as success.
- **`image-generate`** — render scene-spec **image elements** (`image:` + `prompt:` boxes) via OpenAI Images or xAI Imagine into the bundle (also runs for missing assets inside `generate-all`).
- **`image-generate`** — render scene-spec **image elements** (`image:` + `prompt:` boxes) via OpenAI Images or xAI Imagine into the bundle (also runs for missing assets inside `generate-all`). Default **`align_with_docs`** wraps each prompt with narration + source snippets and fails when the authored prompt shares no documented terms. After the PNG is written, **OCR** plus a **vision review** (`align_review`, default true) check the pixels against the same corpus and retry once on FAIL.
- **`manim`** — render Manim scenes declared in config.
- **`compose`** — mux narration audio with visual sources via ffmpeg. With no segment ids, uses ``segments.all`` (same as ``generate-all``), not ``segments.default``. A ``type: mixed`` row raises ``ComposeError`` if any listed source is missing (no silent subset mux).
- **`validate`** / **`validate --pre-push`** — drift, narration lint, Manim hints, **`timing_sync`**, **`story_end`** (last paced reveal vs audio end; hard fail), **`scene_assets`** (pre-render: stuck-board cadence, frame-budget overlaps, `MANIM_FONT` consistency, stale helpers / stale compiled class including hand-edited generated-region labels and `run_time` — hard fail; also a `generate-all` gate before Manim), **`av_sync`** (hard fail on `--pre-push` / `generate-all`; prefers scene-spec labels as OCR anchors), **`subject_beat_coverage`** (declarative specs vs narration topic beats; hard fail when enabled), and related visual-sync checks (`ocr_scan`, `layout`, `freeze_ratio` — hard fail on `--pre-push` / `generate-all`). Missing tesseract fails `ocr_scan` / `av_sync` / `layout` (not skip-PASS). Missing audio or an LFS pointer fails `timing_sync` and recording media gates (`stream_presence`, `av_drift`, `ocr_scan`, `av_sync`) (not skip-PASS). Missing `*.scene.yaml` fails `story_end` / `subject_beat_coverage` for `type: manim` (not skip-PASS).
- **`validate`** / **`validate --pre-push`** — drift, narration lint, Manim hints, **`timing_sync`**, **`story_end`** (last paced reveal vs audio end; hard fail), **`scene_assets`** (pre-render: stuck-board cadence, frame-budget overlaps, `MANIM_FONT` consistency, stale helpers / stale compiled class including hand-edited generated-region labels and `run_time` — hard fail; also a `generate-all` gate before Manim), **`av_sync`** (hard fail on `--pre-push` / `generate-all`; prefers scene-spec labels as OCR anchors), **`subject_beat_coverage`** (declarative specs vs narration topic beats; hard fail when enabled), **`image_prompt_alignment`** (image-element prompts must share documented terms with narration/source; hard fail when enabled), **`image_asset_alignment`** (OCR of generated PNGs vs docs; optional vision when `validation.image_asset_alignment.review`; hard fail when enabled), and related visual-sync checks (`ocr_scan`, `layout`, `freeze_ratio` — hard fail on `--pre-push` / `generate-all`). Missing tesseract fails `ocr_scan` / `av_sync` / `layout` / `image_asset_alignment` (not skip-PASS). Missing audio or an LFS pointer fails `timing_sync` and recording media gates (`stream_presence`, `av_drift`, `ocr_scan`, `av_sync`) (not skip-PASS). Missing `*.scene.yaml` fails `story_end` / `subject_beat_coverage` / `image_prompt_alignment` / `image_asset_alignment` for `type: manim` (not skip-PASS).
- **`lint`** — narration lint helper.
- **`narration-generate`** — LLM-assisted narration from hints and repo context; optional **`--revise --revision-notes`** for in-place edits (same contract as the wizard Revise button).
- **`scene-spec-generate`** — LLM emits declarative **`*.scene.yaml`**; enforces frame budget + **subject-beat coverage** (dwell OK; cover topic shifts; reject invented labels).
- **`scene-spec-generate`** — LLM emits declarative **`*.scene.yaml`**; enforces frame budget + **subject-beat coverage** (dwell OK; cover topic shifts; reject invented labels) and **image-prompt alignment** (image `prompt:` must use documented terms from narration/source).
- **`scene-compile`** — compile specs into **`scenes.py`** (generated regions only).
- **`yaml-generate`** — merge defaults and hint wiring into **`docgen.yaml`**.
- **`clean-bundle`** — remove regenerable outputs per policy.
Expand All @@ -91,6 +91,8 @@ Commands registered on the **`docgen`** CLI include:
- **Manim / `scenes.py` (marker blocks):** Fix generators under `src/docgen/**` (`manim_scene_support.py`, `scene_spec.py`, `scene_spec_generate.py`, `validate`, `yaml_generate`, tests). **Do not** patch generated classes inside a consumer's **`animations/scenes.py`** between **`BEGIN/END GENERATED SCENE`** markers; re-run **`scene-spec-generate`** / **`scene-compile --retime`** and **`manim`** instead. Preferred consumer order: narration → TTS → timestamps → scene-spec/compile → Manim → compose.
- **Beat sync (fail-closed):** when `timing.json` has words, every story box label must match a spoken phrase (`wait_word`); unmatched labels and leftover LLM indices are rejected. Opt out with ``pace: none``. Legacy row-level ``wait_segment`` is upgraded to ``wait_word`` and written back on ``scene-compile`` (direct ``compile_scene_class`` still rejects leftover ``wait_segment``). Fuzzy containment matching is not used. **`scene-compile` clamps FadeIn / page-fade `run_time` against the next word start** so `_TimedScene._clock` cannot race past waits (issue #66 — do not emit cascading first-board dumps). After a reveal, a **dwell** slot may play `Indicate` / `Circumscribe` when the gap to the next `wait_word` is long enough (also clamped). Long holds emit additional mid-hold pulses (`timed_wait` + emphasis) so the board does not freeze after the first Indicate. Optional box fields: `shape` (rounded/pill/diamond), `reveal` (fade/grow/slide), `emphasis` (none/pulse/ring). Page transitions FadeOut revealed boxes, not the parent `VGroup`. `scene-compile` refreshes stale `_box` / `_arrow` / `_TimedScene` helpers in `scenes.py`.
- **Subject-beat coverage:** implemented in `scene_spec.layout_density_violations` / `cluster_subject_beats`; enforced by **`scene-spec-generate`** and **`validate`** (`validation.subject_beat_coverage.enabled`, default true). Not a blind label count.
- **Image-prompt alignment:** implemented in `scene_spec.image_prompt_alignment_violations`; **`scene-spec-generate`** and **`image-generate`** reject prompts that share no documented terms; **`validate`** (`validation.image_prompt_alignment.enabled`, default true) is the same gate. `image-generate` also wraps the Images API prompt with narration/source (`image_generation.align_with_docs`, default true).
- **Image-asset alignment:** OCR + vision review of the generated PNG (`docgen.image_align`). **`image-generate`** fails closed and retries once (`image_generation.align_review`, default true). **`validate`** (`validation.image_asset_alignment`, OCR on by default, vision opt-in) is the same pixel gate.
- **`docgen benchmark` is required** after clock / compile / `_TimedScene` / dwell changes. Pytest string assertions are not a substitute. Do not remove the CI `benchmark` job. See **Required gate** below.
- Prefer **stable CLI / library contracts** and **documented exit codes** so CI can depend on them.
- **`narration_from_source`:** hints in config + **`docgen narration-generate`** — owner-supplied context paths, not opaque bulk edits to outputs.
Expand Down
39 changes: 31 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,21 +55,30 @@ If you still need the legacy behaviour, pin a pre-removal commit
- **Image assets in Manim scenes** — a scene-spec box may be an **image
element** (`image: images/<name>.png` + `prompt:`); `docgen image-generate`
renders the prompt via OpenAI Images (default `gpt-image-1`) or xAI Imagine
(`grok-imagine-image-2.0` when `ai.provider` is `grok`). The compiled scene
shows it with the `_image` helper. `generate-all` fills in missing assets
automatically.
(`grok-imagine-image-2.0` when `ai.provider` is `grok`). By default the
authored prompt is **grounded** in the segment narration plus
`manim_scene_generation` source snippets, and prompts that share no
documented terms fail closed (`image_prompt_alignment`). After the PNG is
written, OCR rejects invented on-image labels and a vision model reviews
the pixels against the same docs (`image_asset_alignment` /
`image_generation.align_review`). The compiled scene shows the asset with
the `_image` helper. `generate-all` fills in missing assets automatically.
- **ffmpeg composition** — combine narration audio and Manim video into final
segments, with a freeze-tail guard.
- **Validation** — A/V drift, freeze ratio, OCR error scan, layout, narration lint,
Manim scene lint, **timing_sync** (stale `timing.json` vs regenerated mp3 —
hard fail), **story_end** (paced visual story finishes long before narration —
hard fail), and **av_sync** (OCR check that scene-spec label anchors appear on
screen near their spoken time — hard fail on `--pre-push` / `generate-all`).
hard fail), **image_prompt_alignment** (image-element `prompt:` must share
documented terms with narration/source — hard fail), **image_asset_alignment**
(OCR of generated PNGs vs docs; optional vision review — hard fail), and **av_sync** (OCR
check that scene-spec label anchors appear on screen near their spoken time
— hard fail on `--pre-push` / `generate-all`).
Missing tesseract fails `ocr_scan` / `av_sync` / `layout` instead of skip-PASS.
Missing audio or an LFS pointer fails `timing_sync` and recording media gates
(`stream_presence`, `av_drift`, `ocr_scan`, `av_sync`) instead of skip-PASS.
Missing `*.scene.yaml` fails `story_end` / `subject_beat_coverage` for
`type: manim` instead of skip-PASS. A hand-edited generated-region label
Missing `*.scene.yaml` fails `story_end` / `subject_beat_coverage` /
`image_prompt_alignment` / `image_asset_alignment` for `type: manim`
instead of skip-PASS. A hand-edited generated-region label
or `run_time` in `scenes.py` fails `scene_assets` compile_sync (clock /
benchmark still execute a fresh `compile_scene_class`, not the on-disk file).
- **GitHub Pages** — auto-generate `index.html`, deploy workflow, LFS rules,
Expand Down Expand Up @@ -128,7 +137,11 @@ using the Cursor key:
image_generation:
model: gpt-image-1 # or dall-e-3, gpt-image-1-mini, …
size: 1536x1024
align_with_docs: true # wrap prompts with narration/source; fail invented terms
align_review: true # vision-review the PNG; retry once on FAIL
# review_model: gpt-4o
# quality: high
# style: "…" # optional override of the educational-diagram prefix
```

```bash
Expand Down Expand Up @@ -241,7 +254,7 @@ docgen --repo /path/to/your-project generate-all
| `docgen freeze [--dist DIR] [--smoke]` | PyInstaller onedir for **`docgen-gui` only** (`pip install 'docgen[packaging]'`). Not the full Manim CLI |
| `docgen tts [--segment 01] [--dry-run]` | Generate TTS audio |
| `docgen timestamps [--engine local\|whisper]` | Extract word/segment timestamps from TTS audio → `timing.json` (default `local`: offline narration-text alignment; `whisper`: OpenAI transcription). Empty `segments.all` is `TimestampError` (stale `timing.json` is not success) |
| `docgen image-generate [--segment 01 \| --all \| --spec PATH] [--force] [--dry-run] [--model …] [--size …]` | Generate scene-spec image assets (`image:` + `prompt:` boxes) via the OpenAI Images API into the bundle |
| `docgen image-generate [--segment 01 \| --all \| --spec PATH] [--force] [--dry-run] [--model …] [--size …]` | Generate scene-spec image assets (`image:` + `prompt:` boxes) via the OpenAI Images API; grounds prompts in narration/source, then OCR + vision-reviews the PNG unless `align_with_docs` / `align_review` is false |
| `docgen manim [--scene StackDAGScene]` | Render Manim animations |
| `docgen compose [01 02 03] [--ffmpeg-timeout 900]` | Compose segments (audio + video). Omit ids to walk `segments.all` (same as `generate-all`); a mapped segment with missing audio/visuals is a hard fail |
| `docgen validate [--max-drift 2.75] [--pre-push]` | Run all validation checks |
Expand Down Expand Up @@ -344,7 +357,11 @@ timestamps:
image_generation: # scene-spec image elements (docgen image-generate)
model: gpt-image-1 # Cursor/OpenAI Images; Grok remaps gpt-image-* to Imagine
size: 1536x1024
align_with_docs: true # ground prompts in narration/source (default)
align_review: true # vision-review generated PNGs (default)
# review_model: gpt-4o
# quality: high # optional, model-specific
# style: "…" # optional Images-API prefix override

manim:
quality: 1080p30 # supports 480p15, 720p30, 1080p30, 1080p60, 1440p30, 1440p60, 2160p60
Expand All @@ -355,6 +372,12 @@ manim:
validation:
subject_beat_coverage:
enabled: true # scene-spec-generate + validate: cover narration topic beats
image_prompt_alignment:
enabled: true # image-element prompts must use documented terms
image_asset_alignment:
enabled: true # OCR generated PNGs vs narration/source
ocr: true
review: false # set true to vision-review existing assets in validate

compose:
ffmpeg_timeout_sec: 300 # can also be overridden with: docgen compose --ffmpeg-timeout N
Expand Down
104 changes: 102 additions & 2 deletions src/docgen/ai_client.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,8 @@

Chat + images for OpenAI/Grok go through the ``openai`` SDK. xAI is
``base_url=https://api.x.ai/v1`` plus model aliases. Anthropic chat uses
``POST https://api.anthropic.com/v1/messages`` (no TTS/images there).
``POST https://api.anthropic.com/v1/messages`` (no image *generation* there;
Claude can still *review* a PNG via ``chat_completion_with_image``).

TTS/STT: OpenAI ``/v1/audio/speech`` + ``whisper-1``, or xAI ``/v1/tts``
and ``/v1/stt``.
Expand Down Expand Up @@ -393,6 +394,89 @@ def _create() -> Any:
return text


def chat_completion_with_image(
*,
system_prompt: str,
user_message: str,
image_bytes: bytes,
media_type: str = "image/png",
model: str,
temperature: float,
cfg: "Config | None" = None,
) -> str:
"""Vision chat: attach one image plus text (OpenAI, Grok, or Anthropic)."""
import base64

settings = resolve_ai_settings(cfg)
resolved = resolve_chat_model(model, settings)
if not image_bytes:
raise AIError("vision chat needs non-empty image bytes")
mime = (media_type or "image/png").strip() or "image/png"
b64 = base64.b64encode(image_bytes).decode("ascii")
if settings.is_anthropic:
return _anthropic_chat(
system_prompt=system_prompt,
user_message=user_message,
model=resolved,
temperature=temperature,
settings=settings,
image_b64=b64,
image_media_type=mime,
)

import openai

from docgen.openai_retry import call_with_rate_limit_retries

client = openai_client(cfg)

def _create() -> Any:
return client.chat.completions.create(
model=resolved,
messages=[
{"role": "system", "content": system_prompt},
{
"role": "user",
"content": [
{"type": "text", "text": user_message},
{
"type": "image_url",
"image_url": {"url": f"data:{mime};base64,{b64}"},
},
],
},
],
temperature=float(temperature),
)

try:
response = call_with_rate_limit_retries(_create)
except openai.AuthenticationError as exc:
raise AIError(
f"{_vendor(settings)} rejected {settings.api_key_env} (authentication failed): {exc}. "
f"{settings.auth_help()}"
) from exc
except openai.PermissionDeniedError as exc:
raise AIError(
f"{_vendor(settings)} permission denied for model {resolved!r}: {exc}."
) from exc
except openai.RateLimitError as exc:
raise AIError(
f"{_vendor(settings)} rate-limited for model {resolved!r}: {exc}."
) from exc
except openai.APIConnectionError as exc:
raise AIError(
f"{_vendor(settings)} connection error: {exc} — re-run when connectivity is restored."
) from exc
try:
text = (response.choices[0].message.content or "").strip()
except (IndexError, AttributeError):
text = ""
if not text:
raise AIError(f"{_vendor(settings)} vision chat returned no text content.")
return text


def synthesize_speech(
*,
text: str,
Expand Down Expand Up @@ -556,15 +640,31 @@ def _anthropic_chat(
model: str,
temperature: float,
settings: AISettings,
image_b64: str | None = None,
image_media_type: str = "image/png",
) -> str:
if not settings.api_key:
raise AIError(f"Anthropic chat needs ANTHROPIC_API_KEY. {settings.auth_help()}")
if image_b64:
user_content: Any = [
{
"type": "image",
"source": {
"type": "base64",
"media_type": image_media_type or "image/png",
"data": image_b64,
},
},
{"type": "text", "text": user_message},
]
else:
user_content = user_message
payload = {
"model": model,
"max_tokens": ANTHROPIC_MAX_TOKENS,
"temperature": float(temperature),
"system": system_prompt,
"messages": [{"role": "user", "content": user_message}],
"messages": [{"role": "user", "content": user_content}],
}
raw = _http_json(
_anthropic_messages_url(settings),
Expand Down
3 changes: 3 additions & 0 deletions src/docgen/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -1142,6 +1142,9 @@ def image_generate_cmd(
if res.status == "dry-run":
click.echo(f"[image-generate] {target.name}: would generate {res.relpath}")
click.echo(f" prompt: {res.prompt}")
if res.effective_prompt and res.effective_prompt != res.prompt:
click.echo(" aligned prompt:")
click.echo(res.effective_prompt)
elif res.status == "exists":
click.echo(f"[image-generate] {target.name}: {res.relpath} exists (skip; use --force)")
else:
Expand Down
Loading
Loading