diff --git a/AGENTS.md b/AGENTS.md index 4fa24aa..5b141a8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -70,13 +70,13 @@ Commands registered on the **`docgen`** CLI include: - **`freeze`** — ``docgen freeze`` builds the **`docgen-gui`** onedir (`pip install 'docgen[packaging]'`). Optional ``--smoke`` runs the binary headless. Do not run a full freeze in routine pytest; set ``DOCGEN_FREEZE_SMOKE=1`` for the optional test. - **`tts`** — text-to-speech for segment files (OpenAI or xAI `/v1/tts`). - **`timestamps`** — word/segment timing (`timing.json`). Default engine **`local`** aligns the known narration text against the mp3 offline (ffmpeg silencedetect, no API); **`--engine whisper`** uses OpenAI whisper-1 or xAI `/v1/stt` when `ai.provider` is grok. Both emit the same Whisper-shaped blocks. Failed ffmpeg silencedetect raises `AlignmentError` (empty stderr is not treated as full-span speech). OpenAI whisper-1 word/segment `start`/`end` must be finite JSON numbers (bool/NaN raise `AIError`). Grok `/v1/stt` rejects empty word tokens and inverted `end < start` intervals (`AIError`). Empty ``segments.all`` raises ``TimestampError`` (same as TTS) and does not leave a stale ``timing.json`` as success. -- **`image-generate`** — render scene-spec **image elements** (`image:` + `prompt:` boxes) via OpenAI Images or xAI Imagine into the bundle (also runs for missing assets inside `generate-all`). +- **`image-generate`** — render scene-spec **image elements** (`image:` + `prompt:` boxes) via OpenAI Images or xAI Imagine into the bundle (also runs for missing assets inside `generate-all`). Default **`align_with_docs`** wraps each prompt with narration + source snippets and fails when the authored prompt shares no documented terms. After the PNG is written, **OCR** plus a **vision review** (`align_review`, default true) check the pixels against the same corpus and retry once on FAIL. - **`manim`** — render Manim scenes declared in config. - **`compose`** — mux narration audio with visual sources via ffmpeg. With no segment ids, uses ``segments.all`` (same as ``generate-all``), not ``segments.default``. A ``type: mixed`` row raises ``ComposeError`` if any listed source is missing (no silent subset mux). -- **`validate`** / **`validate --pre-push`** — drift, narration lint, Manim hints, **`timing_sync`**, **`story_end`** (last paced reveal vs audio end; hard fail), **`scene_assets`** (pre-render: stuck-board cadence, frame-budget overlaps, `MANIM_FONT` consistency, stale helpers / stale compiled class including hand-edited generated-region labels and `run_time` — hard fail; also a `generate-all` gate before Manim), **`av_sync`** (hard fail on `--pre-push` / `generate-all`; prefers scene-spec labels as OCR anchors), **`subject_beat_coverage`** (declarative specs vs narration topic beats; hard fail when enabled), and related visual-sync checks (`ocr_scan`, `layout`, `freeze_ratio` — hard fail on `--pre-push` / `generate-all`). Missing tesseract fails `ocr_scan` / `av_sync` / `layout` (not skip-PASS). Missing audio or an LFS pointer fails `timing_sync` and recording media gates (`stream_presence`, `av_drift`, `ocr_scan`, `av_sync`) (not skip-PASS). Missing `*.scene.yaml` fails `story_end` / `subject_beat_coverage` for `type: manim` (not skip-PASS). +- **`validate`** / **`validate --pre-push`** — drift, narration lint, Manim hints, **`timing_sync`**, **`story_end`** (last paced reveal vs audio end; hard fail), **`scene_assets`** (pre-render: stuck-board cadence, frame-budget overlaps, `MANIM_FONT` consistency, stale helpers / stale compiled class including hand-edited generated-region labels and `run_time` — hard fail; also a `generate-all` gate before Manim), **`av_sync`** (hard fail on `--pre-push` / `generate-all`; prefers scene-spec labels as OCR anchors), **`subject_beat_coverage`** (declarative specs vs narration topic beats; hard fail when enabled), **`image_prompt_alignment`** (image-element prompts must share documented terms with narration/source; hard fail when enabled), **`image_asset_alignment`** (OCR of generated PNGs vs docs; optional vision when `validation.image_asset_alignment.review`; hard fail when enabled), and related visual-sync checks (`ocr_scan`, `layout`, `freeze_ratio` — hard fail on `--pre-push` / `generate-all`). Missing tesseract fails `ocr_scan` / `av_sync` / `layout` / `image_asset_alignment` (not skip-PASS). Missing audio or an LFS pointer fails `timing_sync` and recording media gates (`stream_presence`, `av_drift`, `ocr_scan`, `av_sync`) (not skip-PASS). Missing `*.scene.yaml` fails `story_end` / `subject_beat_coverage` / `image_prompt_alignment` / `image_asset_alignment` for `type: manim` (not skip-PASS). - **`lint`** — narration lint helper. - **`narration-generate`** — LLM-assisted narration from hints and repo context; optional **`--revise --revision-notes`** for in-place edits (same contract as the wizard Revise button). -- **`scene-spec-generate`** — LLM emits declarative **`*.scene.yaml`**; enforces frame budget + **subject-beat coverage** (dwell OK; cover topic shifts; reject invented labels). +- **`scene-spec-generate`** — LLM emits declarative **`*.scene.yaml`**; enforces frame budget + **subject-beat coverage** (dwell OK; cover topic shifts; reject invented labels) and **image-prompt alignment** (image `prompt:` must use documented terms from narration/source). - **`scene-compile`** — compile specs into **`scenes.py`** (generated regions only). - **`yaml-generate`** — merge defaults and hint wiring into **`docgen.yaml`**. - **`clean-bundle`** — remove regenerable outputs per policy. @@ -91,6 +91,8 @@ Commands registered on the **`docgen`** CLI include: - **Manim / `scenes.py` (marker blocks):** Fix generators under `src/docgen/**` (`manim_scene_support.py`, `scene_spec.py`, `scene_spec_generate.py`, `validate`, `yaml_generate`, tests). **Do not** patch generated classes inside a consumer's **`animations/scenes.py`** between **`BEGIN/END GENERATED SCENE`** markers; re-run **`scene-spec-generate`** / **`scene-compile --retime`** and **`manim`** instead. Preferred consumer order: narration → TTS → timestamps → scene-spec/compile → Manim → compose. - **Beat sync (fail-closed):** when `timing.json` has words, every story box label must match a spoken phrase (`wait_word`); unmatched labels and leftover LLM indices are rejected. Opt out with ``pace: none``. Legacy row-level ``wait_segment`` is upgraded to ``wait_word`` and written back on ``scene-compile`` (direct ``compile_scene_class`` still rejects leftover ``wait_segment``). Fuzzy containment matching is not used. **`scene-compile` clamps FadeIn / page-fade `run_time` against the next word start** so `_TimedScene._clock` cannot race past waits (issue #66 — do not emit cascading first-board dumps). After a reveal, a **dwell** slot may play `Indicate` / `Circumscribe` when the gap to the next `wait_word` is long enough (also clamped). Long holds emit additional mid-hold pulses (`timed_wait` + emphasis) so the board does not freeze after the first Indicate. Optional box fields: `shape` (rounded/pill/diamond), `reveal` (fade/grow/slide), `emphasis` (none/pulse/ring). Page transitions FadeOut revealed boxes, not the parent `VGroup`. `scene-compile` refreshes stale `_box` / `_arrow` / `_TimedScene` helpers in `scenes.py`. - **Subject-beat coverage:** implemented in `scene_spec.layout_density_violations` / `cluster_subject_beats`; enforced by **`scene-spec-generate`** and **`validate`** (`validation.subject_beat_coverage.enabled`, default true). Not a blind label count. +- **Image-prompt alignment:** implemented in `scene_spec.image_prompt_alignment_violations`; **`scene-spec-generate`** and **`image-generate`** reject prompts that share no documented terms; **`validate`** (`validation.image_prompt_alignment.enabled`, default true) is the same gate. `image-generate` also wraps the Images API prompt with narration/source (`image_generation.align_with_docs`, default true). +- **Image-asset alignment:** OCR + vision review of the generated PNG (`docgen.image_align`). **`image-generate`** fails closed and retries once (`image_generation.align_review`, default true). **`validate`** (`validation.image_asset_alignment`, OCR on by default, vision opt-in) is the same pixel gate. - **`docgen benchmark` is required** after clock / compile / `_TimedScene` / dwell changes. Pytest string assertions are not a substitute. Do not remove the CI `benchmark` job. See **Required gate** below. - Prefer **stable CLI / library contracts** and **documented exit codes** so CI can depend on them. - **`narration_from_source`:** hints in config + **`docgen narration-generate`** — owner-supplied context paths, not opaque bulk edits to outputs. diff --git a/README.md b/README.md index aec8a80..0e6e018 100644 --- a/README.md +++ b/README.md @@ -55,21 +55,30 @@ If you still need the legacy behaviour, pin a pre-removal commit - **Image assets in Manim scenes** — a scene-spec box may be an **image element** (`image: images/.png` + `prompt:`); `docgen image-generate` renders the prompt via OpenAI Images (default `gpt-image-1`) or xAI Imagine - (`grok-imagine-image-2.0` when `ai.provider` is `grok`). The compiled scene - shows it with the `_image` helper. `generate-all` fills in missing assets - automatically. + (`grok-imagine-image-2.0` when `ai.provider` is `grok`). By default the + authored prompt is **grounded** in the segment narration plus + `manim_scene_generation` source snippets, and prompts that share no + documented terms fail closed (`image_prompt_alignment`). After the PNG is + written, OCR rejects invented on-image labels and a vision model reviews + the pixels against the same docs (`image_asset_alignment` / + `image_generation.align_review`). The compiled scene shows the asset with + the `_image` helper. `generate-all` fills in missing assets automatically. - **ffmpeg composition** — combine narration audio and Manim video into final segments, with a freeze-tail guard. - **Validation** — A/V drift, freeze ratio, OCR error scan, layout, narration lint, Manim scene lint, **timing_sync** (stale `timing.json` vs regenerated mp3 — hard fail), **story_end** (paced visual story finishes long before narration — - hard fail), and **av_sync** (OCR check that scene-spec label anchors appear on - screen near their spoken time — hard fail on `--pre-push` / `generate-all`). + hard fail), **image_prompt_alignment** (image-element `prompt:` must share + documented terms with narration/source — hard fail), **image_asset_alignment** + (OCR of generated PNGs vs docs; optional vision review — hard fail), and **av_sync** (OCR + check that scene-spec label anchors appear on screen near their spoken time + — hard fail on `--pre-push` / `generate-all`). Missing tesseract fails `ocr_scan` / `av_sync` / `layout` instead of skip-PASS. Missing audio or an LFS pointer fails `timing_sync` and recording media gates (`stream_presence`, `av_drift`, `ocr_scan`, `av_sync`) instead of skip-PASS. - Missing `*.scene.yaml` fails `story_end` / `subject_beat_coverage` for - `type: manim` instead of skip-PASS. A hand-edited generated-region label + Missing `*.scene.yaml` fails `story_end` / `subject_beat_coverage` / + `image_prompt_alignment` / `image_asset_alignment` for `type: manim` + instead of skip-PASS. A hand-edited generated-region label or `run_time` in `scenes.py` fails `scene_assets` compile_sync (clock / benchmark still execute a fresh `compile_scene_class`, not the on-disk file). - **GitHub Pages** — auto-generate `index.html`, deploy workflow, LFS rules, @@ -128,7 +137,11 @@ using the Cursor key: image_generation: model: gpt-image-1 # or dall-e-3, gpt-image-1-mini, … size: 1536x1024 + align_with_docs: true # wrap prompts with narration/source; fail invented terms + align_review: true # vision-review the PNG; retry once on FAIL + # review_model: gpt-4o # quality: high + # style: "…" # optional override of the educational-diagram prefix ``` ```bash @@ -241,7 +254,7 @@ docgen --repo /path/to/your-project generate-all | `docgen freeze [--dist DIR] [--smoke]` | PyInstaller onedir for **`docgen-gui` only** (`pip install 'docgen[packaging]'`). Not the full Manim CLI | | `docgen tts [--segment 01] [--dry-run]` | Generate TTS audio | | `docgen timestamps [--engine local\|whisper]` | Extract word/segment timestamps from TTS audio → `timing.json` (default `local`: offline narration-text alignment; `whisper`: OpenAI transcription). Empty `segments.all` is `TimestampError` (stale `timing.json` is not success) | -| `docgen image-generate [--segment 01 \| --all \| --spec PATH] [--force] [--dry-run] [--model …] [--size …]` | Generate scene-spec image assets (`image:` + `prompt:` boxes) via the OpenAI Images API into the bundle | +| `docgen image-generate [--segment 01 \| --all \| --spec PATH] [--force] [--dry-run] [--model …] [--size …]` | Generate scene-spec image assets (`image:` + `prompt:` boxes) via the OpenAI Images API; grounds prompts in narration/source, then OCR + vision-reviews the PNG unless `align_with_docs` / `align_review` is false | | `docgen manim [--scene StackDAGScene]` | Render Manim animations | | `docgen compose [01 02 03] [--ffmpeg-timeout 900]` | Compose segments (audio + video). Omit ids to walk `segments.all` (same as `generate-all`); a mapped segment with missing audio/visuals is a hard fail | | `docgen validate [--max-drift 2.75] [--pre-push]` | Run all validation checks | @@ -344,7 +357,11 @@ timestamps: image_generation: # scene-spec image elements (docgen image-generate) model: gpt-image-1 # Cursor/OpenAI Images; Grok remaps gpt-image-* to Imagine size: 1536x1024 + align_with_docs: true # ground prompts in narration/source (default) + align_review: true # vision-review generated PNGs (default) + # review_model: gpt-4o # quality: high # optional, model-specific + # style: "…" # optional Images-API prefix override manim: quality: 1080p30 # supports 480p15, 720p30, 1080p30, 1080p60, 1440p30, 1440p60, 2160p60 @@ -355,6 +372,12 @@ manim: validation: subject_beat_coverage: enabled: true # scene-spec-generate + validate: cover narration topic beats + image_prompt_alignment: + enabled: true # image-element prompts must use documented terms + image_asset_alignment: + enabled: true # OCR generated PNGs vs narration/source + ocr: true + review: false # set true to vision-review existing assets in validate compose: ffmpeg_timeout_sec: 300 # can also be overridden with: docgen compose --ffmpeg-timeout N diff --git a/src/docgen/ai_client.py b/src/docgen/ai_client.py index 16a3bb2..536ae97 100644 --- a/src/docgen/ai_client.py +++ b/src/docgen/ai_client.py @@ -9,7 +9,8 @@ Chat + images for OpenAI/Grok go through the ``openai`` SDK. xAI is ``base_url=https://api.x.ai/v1`` plus model aliases. Anthropic chat uses -``POST https://api.anthropic.com/v1/messages`` (no TTS/images there). +``POST https://api.anthropic.com/v1/messages`` (no image *generation* there; +Claude can still *review* a PNG via ``chat_completion_with_image``). TTS/STT: OpenAI ``/v1/audio/speech`` + ``whisper-1``, or xAI ``/v1/tts`` and ``/v1/stt``. @@ -393,6 +394,89 @@ def _create() -> Any: return text +def chat_completion_with_image( + *, + system_prompt: str, + user_message: str, + image_bytes: bytes, + media_type: str = "image/png", + model: str, + temperature: float, + cfg: "Config | None" = None, +) -> str: + """Vision chat: attach one image plus text (OpenAI, Grok, or Anthropic).""" + import base64 + + settings = resolve_ai_settings(cfg) + resolved = resolve_chat_model(model, settings) + if not image_bytes: + raise AIError("vision chat needs non-empty image bytes") + mime = (media_type or "image/png").strip() or "image/png" + b64 = base64.b64encode(image_bytes).decode("ascii") + if settings.is_anthropic: + return _anthropic_chat( + system_prompt=system_prompt, + user_message=user_message, + model=resolved, + temperature=temperature, + settings=settings, + image_b64=b64, + image_media_type=mime, + ) + + import openai + + from docgen.openai_retry import call_with_rate_limit_retries + + client = openai_client(cfg) + + def _create() -> Any: + return client.chat.completions.create( + model=resolved, + messages=[ + {"role": "system", "content": system_prompt}, + { + "role": "user", + "content": [ + {"type": "text", "text": user_message}, + { + "type": "image_url", + "image_url": {"url": f"data:{mime};base64,{b64}"}, + }, + ], + }, + ], + temperature=float(temperature), + ) + + try: + response = call_with_rate_limit_retries(_create) + except openai.AuthenticationError as exc: + raise AIError( + f"{_vendor(settings)} rejected {settings.api_key_env} (authentication failed): {exc}. " + f"{settings.auth_help()}" + ) from exc + except openai.PermissionDeniedError as exc: + raise AIError( + f"{_vendor(settings)} permission denied for model {resolved!r}: {exc}." + ) from exc + except openai.RateLimitError as exc: + raise AIError( + f"{_vendor(settings)} rate-limited for model {resolved!r}: {exc}." + ) from exc + except openai.APIConnectionError as exc: + raise AIError( + f"{_vendor(settings)} connection error: {exc} — re-run when connectivity is restored." + ) from exc + try: + text = (response.choices[0].message.content or "").strip() + except (IndexError, AttributeError): + text = "" + if not text: + raise AIError(f"{_vendor(settings)} vision chat returned no text content.") + return text + + def synthesize_speech( *, text: str, @@ -556,15 +640,31 @@ def _anthropic_chat( model: str, temperature: float, settings: AISettings, + image_b64: str | None = None, + image_media_type: str = "image/png", ) -> str: if not settings.api_key: raise AIError(f"Anthropic chat needs ANTHROPIC_API_KEY. {settings.auth_help()}") + if image_b64: + user_content: Any = [ + { + "type": "image", + "source": { + "type": "base64", + "media_type": image_media_type or "image/png", + "data": image_b64, + }, + }, + {"type": "text", "text": user_message}, + ] + else: + user_content = user_message payload = { "model": model, "max_tokens": ANTHROPIC_MAX_TOKENS, "temperature": float(temperature), "system": system_prompt, - "messages": [{"role": "user", "content": user_message}], + "messages": [{"role": "user", "content": user_content}], } raw = _http_json( _anthropic_messages_url(settings), diff --git a/src/docgen/cli.py b/src/docgen/cli.py index 9501b4c..49c075b 100644 --- a/src/docgen/cli.py +++ b/src/docgen/cli.py @@ -1142,6 +1142,9 @@ def image_generate_cmd( if res.status == "dry-run": click.echo(f"[image-generate] {target.name}: would generate {res.relpath}") click.echo(f" prompt: {res.prompt}") + if res.effective_prompt and res.effective_prompt != res.prompt: + click.echo(" aligned prompt:") + click.echo(res.effective_prompt) elif res.status == "exists": click.echo(f"[image-generate] {target.name}: {res.relpath} exists (skip; use --force)") else: diff --git a/src/docgen/config.py b/src/docgen/config.py index 9f37b47..0aba5d2 100644 --- a/src/docgen/config.py +++ b/src/docgen/config.py @@ -331,6 +331,8 @@ def __post_init__(self) -> None: "story_end", "narration_lint", "subject_beat_coverage", + "image_prompt_alignment", + "image_asset_alignment", ): self._sub_block(validation, nested, label=f"validation.{nested}") # List-valued keys: a string must not be iterated as characters. @@ -485,6 +487,30 @@ def __post_init__(self) -> None: require_yaml_string(ig["size"], label="image_generation.size", source=src) if ig.get("quality") is not None: require_yaml_string(ig["quality"], label="image_generation.quality", source=src) + if ig.get("align_with_docs") is not None: + require_yaml_bool( + ig["align_with_docs"], + label="image_generation.align_with_docs", + source=src, + ) + if ig.get("style") is not None: + require_yaml_string(ig["style"], label="image_generation.style", source=src) + if ig.get("align_review") is not None: + require_yaml_bool( + ig["align_review"], + label="image_generation.align_review", + source=src, + ) + if ig.get("review_model") is not None: + require_yaml_string( + ig["review_model"], label="image_generation.review_model", source=src + ) + if ig.get("align_review_retries") is not None: + require_yaml_number( + ig["align_review_retries"], + label="image_generation.align_review_retries", + source=src, + ) ai = self._block("ai") if ai.get("provider") is not None: require_yaml_string(ai["provider"], label="ai.provider", source=src) @@ -766,6 +792,40 @@ def __post_init__(self) -> None: label="validation.subject_beat_coverage.enabled", source=src, ) + ipa = self._sub_block( + validation, + "image_prompt_alignment", + label="validation.image_prompt_alignment", + ) + if ipa.get("enabled") is not None: + require_yaml_bool( + ipa["enabled"], + label="validation.image_prompt_alignment.enabled", + source=src, + ) + iaa = self._sub_block( + validation, + "image_asset_alignment", + label="validation.image_asset_alignment", + ) + if iaa.get("enabled") is not None: + require_yaml_bool( + iaa["enabled"], + label="validation.image_asset_alignment.enabled", + source=src, + ) + if iaa.get("ocr") is not None: + require_yaml_bool( + iaa["ocr"], + label="validation.image_asset_alignment.ocr", + source=src, + ) + if iaa.get("review") is not None: + require_yaml_bool( + iaa["review"], + label="validation.image_asset_alignment.review", + source=src, + ) def _source_label(self) -> str: return self.yaml_path.name if self.yaml_path else "docgen.yaml" @@ -950,10 +1010,49 @@ def image_generation_config(self) -> dict[str, Any]: defaults: dict[str, Any] = { "model": "gpt-image-1", "size": "1536x1024", + "align_with_docs": True, + "align_review": True, + "align_review_retries": 1, } defaults.update(self._block("image_generation")) return defaults + @property + def image_align_with_docs(self) -> bool: + """When true (default), ground image prompts in narration/source and fail unaligned ones.""" + block = self._block("image_generation") + if "align_with_docs" in block: + return bool(block.get("align_with_docs")) + return True + + @property + def image_prompt_alignment_enabled(self) -> bool: + """When true (default), validate + scene-spec-generate enforce image-prompt grounding. + + Config: ``validation.image_prompt_alignment.enabled`` (bool). + """ + block = self._sub_block( + self._block("validation"), + "image_prompt_alignment", + label="validation.image_prompt_alignment", + ) + if "enabled" in block: + return bool(block.get("enabled")) + return True + + @property + def image_asset_alignment_config(self) -> dict[str, Any]: + """Pixel checks on generated scene images (OCR + optional vision).""" + defaults: dict[str, Any] = {"enabled": True, "ocr": True, "review": False} + defaults.update( + self._sub_block( + self._block("validation"), + "image_asset_alignment", + label="validation.image_asset_alignment", + ) + ) + return defaults + # -- Manim ----------------------------------------------------------------- @property diff --git a/src/docgen/image_align.py b/src/docgen/image_align.py new file mode 100644 index 0000000..b31dc4d --- /dev/null +++ b/src/docgen/image_align.py @@ -0,0 +1,150 @@ +"""Pixel-level alignment of generated scene images against documentation. + +Prompt grounding (``image_prompt_alignment``) only checks the caption. This +module looks at the **PNG**: + +* **OCR** — readable tokens that are not in narration/source fail (invented + on-image labels). Empty OCR is fine (many diagrams have no text). +* **Vision review** — a chat model that accepts images (OpenAI, Grok, or + Claude) answers PASS/FAIL against the same documentation corpus. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from pathlib import Path +from typing import TYPE_CHECKING, Callable + +if TYPE_CHECKING: + from docgen.config import Config + +DEFAULT_REVIEW_MODEL = "gpt-4o" + +_REVIEW_SYSTEM = """You review one illustration for a documentation video. + +Decide whether the IMAGE depicts the DOCUMENTED SUBJECT (narration + source). +PASS when the picture is a reasonable, simplified depiction of those concepts. +FAIL when it shows a different subject, invented product/API names, decorative +text that is not copied from the docs, fake UI chrome, or generic clip-art +unrelated to the documentation. + +Reply with EXACTLY two lines and nothing else: +PASS + +or +FAIL + +""" + + +@dataclass(frozen=True) +class ImageReviewResult: + passed: bool + reason: str + ocr_text: str = "" + + +def ocr_image_text(path: Path) -> str | None: + """OCR a still image. ``None`` if unreadable or tesseract is unavailable. + + Empty string means the file decoded but no text was found. + """ + try: + import cv2 + import pytesseract + + pytesseract.get_tesseract_version() + except Exception: + return None + img = cv2.imread(str(path)) + if img is None: + return None + gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) + _, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) + try: + return str(pytesseract.image_to_string(thresh) or "") + except Exception: + return None + + +def parse_review_verdict(text: str) -> ImageReviewResult: + """Parse a PASS/FAIL vision reply. Unparseable text fails closed.""" + lines = [ln.strip() for ln in (text or "").splitlines() if ln.strip()] + if not lines: + return ImageReviewResult(False, "vision review returned empty text") + verdict = lines[0].split()[0].upper().strip(".:") + reason = lines[1] if len(lines) > 1 else " ".join(lines[1:]) or lines[0] + if verdict == "PASS": + return ImageReviewResult(True, reason) + if verdict == "FAIL": + return ImageReviewResult(False, reason) + preview = (text or "").strip() + if len(preview) > 80: + preview = preview[:77] + "..." + return ImageReviewResult(False, f"vision review reply was not PASS/FAIL: {preview!r}") + + +def build_review_user_message( + *, + corpus_text: str, + authored_prompt: str = "", + label: str = "", +) -> str: + parts = [ + "DOCUMENTED SUBJECT:", + (corpus_text or "").strip() or "(none)", + ] + if (authored_prompt or "").strip(): + parts.extend(["", "AUTHORED IMAGE PROMPT:", authored_prompt.strip()]) + if (label or "").strip(): + parts.extend(["", "ON-SCREEN LABEL:", label.strip()]) + parts.extend(["", "Review the attached image."]) + return "\n".join(parts) + + +def _media_type_for(path: Path) -> str: + ext = path.suffix.lower() + return { + ".jpg": "image/jpeg", + ".jpeg": "image/jpeg", + ".webp": "image/webp", + ".gif": "image/gif", + }.get(ext, "image/png") + + +def review_image_against_docs( + path: Path, + *, + corpus_text: str, + authored_prompt: str = "", + label: str = "", + cfg: "Config | None" = None, + model: str = "", + chat_fn: Callable[..., str] | None = None, +) -> ImageReviewResult: + """Send ``path`` + documentation to a vision-capable chat model.""" + from docgen.ai_client import AIError, chat_completion_with_image + + try: + raw = path.read_bytes() + except OSError as exc: + return ImageReviewResult(False, f"could not read image {path}: {exc}") + if not raw: + return ImageReviewResult(False, f"image {path} is empty") + user = build_review_user_message( + corpus_text=corpus_text, authored_prompt=authored_prompt, label=label + ) + invoke = chat_fn or chat_completion_with_image + try: + text = invoke( + system_prompt=_REVIEW_SYSTEM, + user_message=user, + image_bytes=raw, + media_type=_media_type_for(path), + model=(model or "").strip() or DEFAULT_REVIEW_MODEL, + temperature=0.0, + cfg=cfg, + ) + except (AIError, RuntimeError, OSError, TypeError, ValueError) as exc: + return ImageReviewResult(False, f"vision review failed: {exc}") + return parse_review_verdict(text) diff --git a/src/docgen/image_generate.py b/src/docgen/image_generate.py index b41722b..980cbb8 100644 --- a/src/docgen/image_generate.py +++ b/src/docgen/image_generate.py @@ -11,8 +11,15 @@ ``docgen image-generate`` scans specs, calls the Images API (OpenAI or xAI Imagine) for elements whose asset is missing (or ``--force``), and writes PNG bytes to -``/``. ``docgen manim`` then loads the asset via the -``_image`` helper in ``scenes.py``. +``/``. By default the authored ``prompt`` is **grounded** in +the segment narration plus ``manim_scene_generation`` source snippets so the +image model sees the documentation, not only a short caption. Prompts that +share no documented terms fail closed (same check as ``validate`` / +``image_prompt_alignment``). After the PNG is written, **OCR** rejects +invented on-image labels and a **vision review** (OpenAI / Grok / Claude) +checks the pixels against the same corpus (``image_generation.align_review``). +A failed review retries once with the critique, then deletes the asset. +``docgen manim`` then loads the asset via the ``_image`` helper in ``scenes.py``. """ from __future__ import annotations @@ -20,16 +27,34 @@ import base64 from dataclasses import dataclass from pathlib import Path -from typing import TYPE_CHECKING, Callable +from typing import TYPE_CHECKING, Any, Callable +from docgen.image_align import ( + ImageReviewResult, + ocr_image_text, + review_image_against_docs, +) from docgen.openai_retry import call_with_rate_limit_retries -from docgen.scene_spec import iter_image_elements, load_scene_spec +from docgen.scene_spec import ( + image_ocr_alignment_violations, + image_prompt_alignment_violations, + iter_image_elements, + load_scene_spec, +) if TYPE_CHECKING: from docgen.config import Config DEFAULT_IMAGE_MODEL = "gpt-image-1" DEFAULT_IMAGE_SIZE = "1536x1024" +DEFAULT_IMAGE_STYLE = ( + "Clean educational diagram for a documentation video. Flat vector " + "illustration, high contrast, no watermark, no signature, no decorative " + "fake UI or invented product logos. Any readable text must be short ASCII " + "copied from the documented subject. Do not add components that are not " + "named in the documentation." +) +_ALIGN_CORPUS_EXCERPT = 2200 class ImageGenerationError(RuntimeError): @@ -42,6 +67,7 @@ class ImageAssetResult: path: Path status: str # "generated" | "exists" | "dry-run" prompt: str + effective_prompt: str = "" def generate_image_bytes( @@ -128,6 +154,69 @@ def generate_image_bytes( ) +def _excerpt(text: str, limit: int = _ALIGN_CORPUS_EXCERPT) -> str: + collapsed = " ".join((text or "").split()) + if len(collapsed) <= limit: + return collapsed + cut = collapsed[: limit - 1] + if " " in cut: + cut = cut.rsplit(" ", 1)[0] + return cut + "…" + + +def build_aligned_image_prompt( + authored: str, + *, + corpus_text: str = "", + label: str = "", + style: str = DEFAULT_IMAGE_STYLE, +) -> str: + """Wrap an authored scene-spec prompt with documentation + style constraints.""" + parts: list[str] = [(style or "").strip() or DEFAULT_IMAGE_STYLE, ""] + corpus = (corpus_text or "").strip() + if corpus: + parts.append( + "Documented subject (use these terms; do not invent names, logos, " + "or extra components):" + ) + parts.append(_excerpt(corpus)) + parts.append("") + lab = (label or "").strip() + if lab: + parts.append(f"On-screen timing label (must remain accurate): {lab}") + parts.append("") + parts.append("Illustration request:") + parts.append((authored or "").strip()) + return "\n".join(parts).strip() + "\n" + + +def collect_alignment_corpus(cfg: "Config", spec: dict[str, Any]) -> str: + """Narration + scene-generation hints + source snippets for one spec.""" + parts: list[str] = [] + seg_id = str(spec.get("segment_id") or "").strip() + if seg_id: + found = cfg.find_segment_asset(cfg.narration_dir, seg_id, ".md") + if found is not None and found.is_file(): + try: + parts.append(found.read_text(encoding="utf-8")) + except OSError: + pass + from docgen.manim_scene_support import ( + collect_source_snippets, + merged_scene_generation_settings, + ) + + settings = merged_scene_generation_settings(cfg, seg_id) + for h in settings.hints: + if str(h).strip(): + parts.append(str(h).strip()) + for label, text in collect_source_snippets(cfg, settings, extra_paths=[]): + body = str(text or "").strip() + if body: + parts.append(f"{label}\n{body}") + return "\n\n".join(p for p in parts if str(p).strip()) + + def _resolve_asset_path(cfg: "Config", relpath: str) -> Path: p = Path(relpath) if p.is_absolute() or ".." in p.parts: @@ -147,6 +236,8 @@ def generate_images_for_spec( model_override: str | None = None, size_override: str | None = None, image_fn: Callable[[str], bytes] | None = None, + review_fn: Callable[..., ImageReviewResult] | None = None, + ocr_fn: Callable[[Path], str | None] | None = None, ) -> list[ImageAssetResult]: """Generate missing image assets referenced by one ``*.scene.yaml``. @@ -154,7 +245,9 @@ def generate_images_for_spec( missing **and** has no ``prompt`` fails loud — either commit the file or give the toolchain a prompt to generate it from. - ``image_fn`` is an injection point for tests (prompt → PNG bytes). + ``image_fn`` / ``review_fn`` / ``ocr_fn`` are injection points for tests. + Live vision review runs when ``align_review`` is on and ``image_fn`` is + not injected (or ``review_fn`` is provided). """ spec = load_scene_spec(spec_path) elements = iter_image_elements(spec) @@ -163,15 +256,44 @@ def generate_images_for_spec( size = (size_override or "").strip() or str(icfg.get("size") or DEFAULT_IMAGE_SIZE) quality = icfg.get("quality") quality = str(quality).strip() if quality else None + align = bool(icfg.get("align_with_docs", True)) + align_review = bool(icfg.get("align_review", True)) + try: + pixel_retries = int(icfg.get("align_review_retries", 1) or 0) + except (TypeError, ValueError): + pixel_retries = 1 + pixel_retries = max(0, pixel_retries) + review_model = str(icfg.get("review_model") or "").strip() + style = str(icfg.get("style") or "").strip() or DEFAULT_IMAGE_STYLE + corpus = collect_alignment_corpus(cfg, spec) if align else "" + live_review = align and align_review and (review_fn is not None or image_fn is None) + + if align: + issues = image_prompt_alignment_violations(spec, corpus_text=corpus) + if issues: + joined = "\n ".join(issues) + raise ImageGenerationError( + f"{spec_path}: image prompt alignment failed — rewrite each " + f"`prompt` so it uses documented terms from narration/source " + f"(or set image_generation.align_with_docs: false):\n {joined}" + ) results: list[ImageAssetResult] = [] for el in elements: rel = str(el["image"]).strip() prompt = str(el.get("prompt") or "").strip() out = _resolve_asset_path(cfg, rel) + label = str(el.get("label") or "").strip() + effective = ( + build_aligned_image_prompt( + prompt, corpus_text=corpus, label=label, style=style + ) + if align and prompt + else prompt + ) if out.is_file() and not force: - results.append(ImageAssetResult(rel, out, "exists", prompt)) + results.append(ImageAssetResult(rel, out, "exists", prompt, effective)) continue if not prompt: raise ImageGenerationError( @@ -179,7 +301,7 @@ def generate_images_for_spec( f"({out}); add the file to the bundle or set a prompt in the spec." ) if dry_run: - results.append(ImageAssetResult(rel, out, "dry-run", prompt)) + results.append(ImageAssetResult(rel, out, "dry-run", prompt, effective)) continue fn = image_fn or ( @@ -187,17 +309,95 @@ def generate_images_for_spec( prompt=p, model=model, size=size, quality=quality, cfg=cfg ) ) - data = fn(prompt) - if not data: + critique = "" + kept = False + last_issues: list[str] = [] + for attempt in range(1 + pixel_retries): + to_send = effective + if critique: + to_send = ( + f"{effective}\n\n--- PIXEL REVIEW FAILED ---\n{critique}\n" + "Redraw so the image matches the documented subject. " + "Do not invent labels or extra components." + ) + data = fn(to_send) + if not data: + raise ImageGenerationError( + f"{spec_path}: image element {rel!r} — provider returned empty bytes" + ) + out.parent.mkdir(parents=True, exist_ok=True) + out.write_bytes(data) + last_issues = _pixel_alignment_issues( + out, + relpath=rel, + corpus=corpus, + authored_prompt=prompt, + label=label, + cfg=cfg, + align=align, + live_review=live_review, + review_model=review_model, + review_fn=review_fn, + ocr_fn=ocr_fn, + ) + if not last_issues: + kept = True + break + critique = "\n".join(last_issues) + if not kept: + out.unlink(missing_ok=True) + joined = "\n ".join(last_issues) raise ImageGenerationError( - f"{spec_path}: image element {rel!r} — provider returned empty bytes" + f"{spec_path}: image {rel!r} failed pixel alignment " + f"(OCR / vision review):\n {joined}" ) - out.parent.mkdir(parents=True, exist_ok=True) - out.write_bytes(data) - results.append(ImageAssetResult(rel, out, "generated", prompt)) + results.append(ImageAssetResult(rel, out, "generated", prompt, effective)) return results +def _pixel_alignment_issues( + path: Path, + *, + relpath: str, + corpus: str, + authored_prompt: str, + label: str, + cfg: "Config", + align: bool, + live_review: bool, + review_model: str, + review_fn: Callable[..., ImageReviewResult] | None, + ocr_fn: Callable[[Path], str | None] | None, +) -> list[str]: + if not align or not corpus.strip(): + return [] + issues: list[str] = [] + scanned = (ocr_fn or ocr_image_text)(path) + if scanned: + issues.extend( + image_ocr_alignment_violations( + scanned, corpus_text=corpus, relpath=relpath + ) + ) + if live_review: + if review_fn is not None: + verdict = review_fn( + path, corpus_text=corpus, authored_prompt=authored_prompt, label=label + ) + else: + verdict = review_image_against_docs( + path, + corpus_text=corpus, + authored_prompt=authored_prompt, + label=label, + cfg=cfg, + model=review_model, + ) + if not verdict.passed: + issues.append(f"{relpath}: vision review FAIL — {verdict.reason}") + return issues + + def spec_files_for_bundle(cfg: "Config") -> list[Path]: """All committed ``animations/specs/*.scene.yaml`` files, sorted.""" specs_dir = cfg.animations_dir / "specs" @@ -210,6 +410,8 @@ def generate_missing_images_for_bundle( cfg: "Config", *, image_fn: Callable[[str], bytes] | None = None, + review_fn: Callable[..., ImageReviewResult] | None = None, + ocr_fn: Callable[[Path], str | None] | None = None, ) -> list[str]: """Generate only **missing** image assets across all bundle specs. @@ -218,7 +420,9 @@ def generate_missing_images_for_bundle( """ msgs: list[str] = [] for spec_path in spec_files_for_bundle(cfg): - for res in generate_images_for_spec(cfg, spec_path, image_fn=image_fn): + for res in generate_images_for_spec( + cfg, spec_path, image_fn=image_fn, review_fn=review_fn, ocr_fn=ocr_fn + ): if res.status == "generated": msgs.append(f"{spec_path.name}: generated {res.relpath}") return msgs diff --git a/src/docgen/init.py b/src/docgen/init.py index 8f16208..98ed0d6 100644 --- a/src/docgen/init.py +++ b/src/docgen/init.py @@ -338,6 +338,8 @@ def _write_config(plan: InitPlan) -> str: "image_generation": { "model": "gpt-image-1", # Cursor/OpenAI Images; override with --model "size": "1536x1024", + "align_with_docs": True, # ground prompts in narration/source; fail invented terms + "align_review": True, # vision-review the PNG against the same docs }, "compose": { "ffmpeg_timeout_sec": 300, diff --git a/src/docgen/scene_spec.py b/src/docgen/scene_spec.py index ed7d94c..d22fbd2 100644 --- a/src/docgen/scene_spec.py +++ b/src/docgen/scene_spec.py @@ -1339,6 +1339,92 @@ def content_tokens(text: str) -> set[str]: return {w for w in words if len(w) > 2 and w not in _CONTENT_STOPWORDS} +# Visual-style words an image prompt may use without being documented terms. +_IMAGE_STYLE_TOKENS = frozenset( + """ + diagram diagrams illustration illustrations icon icons flat clean vector + isometric cartoon watercolor sketch photo photographic realistic abstract + artwork image images picture pictures visual background foreground style + colored colour color palette high contrast simple minimal educational + documentary video watermark signature logo logos chrome banner poster + scene board rounded pill diamond arrow arrows flow flowchart infographic + thumbnail render rendering drawing draw depict showing shows show + depicting depicts labeled labelled label labels text white black blue + green orange red gray grey dark light bright soft hard wide tall small + large tiny big thin thick line lines box boxes node nodes panel panels + layout grid row rows column columns + """.split() +) + + +def image_prompt_alignment_violations( + spec: dict[str, Any], + *, + corpus_text: str, +) -> list[str]: + """Reject image prompts that share no documented terms with narration/source. + + Style adjectives (``flat``, ``diagram``, ``illustration``, …) are ignored. + A missing corpus (no narration and no source) is a no-op — there is nothing + to align against. Specs with no image elements, or image elements that omit + ``prompt`` (committed assets), also pass. + """ + elements = iter_image_elements(spec) + if not elements: + return [] + corpus = content_tokens(corpus_text) + if not corpus: + return [] + issues: list[str] = [] + for el in elements: + rel = str(el.get("image") or "").strip() or "(unnamed image)" + prompt = str(el.get("prompt") or "").strip() + if not prompt: + continue + toks = content_tokens(prompt) + substance = toks - _IMAGE_STYLE_TOKENS + if not substance: + issues.append( + f"{rel}: image prompt is only visual style (no documented subject terms)" + ) + continue + if not (substance & corpus): + preview = prompt if len(prompt) <= 80 else prompt[:77] + "..." + issues.append( + f"{rel}: image prompt shares no documented terms with " + f"narration/source: {preview!r}" + ) + return issues + + +def image_ocr_alignment_violations( + ocr_text: str, + *, + corpus_text: str, + relpath: str = "image", +) -> list[str]: + """Reject OCR tokens that look like invented documented terms. + + Style words and a missing corpus are ignored. A single short OCR ghost + (``<6`` letters) is tolerated; two confident invented tokens, or one + token of length ≥ 6, fail. + """ + corpus = content_tokens(corpus_text) + if not corpus: + return [] + substance = content_tokens(ocr_text) - _IMAGE_STYLE_TOKENS + if not substance: + return [] + invented = {t for t in substance if t not in corpus} + confident = {t for t in invented if len(t) >= 4} + if len(confident) >= 2 or any(len(t) >= 6 for t in confident): + sample = ", ".join(repr(x) for x in sorted(confident)[:6]) + return [ + f"{relpath}: on-image OCR has terms not in narration/source: {sample}" + ] + return [] + + def cluster_subject_beats( sentences: list[str], *, diff --git a/src/docgen/scene_spec_generate.py b/src/docgen/scene_spec_generate.py index 693d07e..73b9410 100644 --- a/src/docgen/scene_spec_generate.py +++ b/src/docgen/scene_spec_generate.py @@ -36,6 +36,7 @@ sanitize_pacing_conflicts, compile_scene_class, cluster_subject_beats, + image_prompt_alignment_violations, layout_budget_violations, layout_density_violations, layout_stack_budget, @@ -90,10 +91,14 @@ may instead be an image element with: - image: bundle-relative asset path, e.g. ``images/.png`` (no absolute paths, no "..") - width / height: positive numbers (frame budget rules above apply; images count like boxes) - - prompt: string — a clear visual description; ``docgen image-generate`` renders it via the OpenAI Images API - - label: optional single word from the narration used as the timing anchor for the reveal + - prompt: string — a clear visual description grounded in the narration **and** SOURCE + DOCUMENTATION; ``docgen image-generate`` renders it and rejects prompts that invent + undocumented product names or share no documented terms + - label: optional spoken phrase from the narration used as the timing anchor for the reveal Image elements must NOT carry ``color`` or ``font_size``. Prefer labeled boxes for diagrams; use images -only for illustrative artwork the hints explicitly request. +only for illustrative artwork the hints explicitly request. Image ``prompt`` text must name +the same concepts as the narration/source (not generic "a diagram" / invented architecture). +Any words you expect to appear *inside* the artwork must be short ASCII copied from the docs. Optional per-box (**Whisper ``words`` only**); omit if unsure — compile fills from each box ``label`` → first transcript match: - wait_word: non-negative int — index into ``timing.json`` → ``words``; that box waits until that token's **start**, then fades in (**one box at a time** within each row). @@ -251,6 +256,22 @@ def build_scene_spec_user_message( if str(h).strip(): parts.append(f"- {str(h).strip()}") + if source_snippets: + parts.append("") + parts.append("--- SOURCE DOCUMENTATION ---") + parts.append( + "Use these excerpts when authoring box labels **and** image prompts. " + "Do not invent product names, APIs, or architecture that is not in this " + "source or the narration." + ) + for label, text in source_snippets: + body = str(text or "").strip() + if not body: + continue + parts.append("") + parts.append(f"### {label}") + parts.append(body) + if reference_scenes: parts.append("") parts.append( @@ -449,6 +470,7 @@ def _parse_and_harden_llm_spec( raw: str, enforce_density: bool, density_slack: int = 0, + corpus_text: str = "", ) -> dict[str, Any]: """Parse YAML, auto-layout, validate schema/budget/(optional) density, compile-lint.""" body = strip_yaml_fences(raw) @@ -499,6 +521,17 @@ def _parse_and_harden_llm_spec( f"segment {seg_id}: scene spec failed subject-beat coverage:\n {joined}\nDraft: {draft}" ) + if getattr(cfg, "image_prompt_alignment_enabled", True): + align_issues = image_prompt_alignment_violations( + merged_spec, corpus_text=corpus_text or narration_text + ) + if align_issues: + draft = _save_draft(cfg, seg_id, body) + joined = "\n ".join(align_issues) + raise SceneGenerationError( + f"segment {seg_id}: image prompt alignment failed:\n {joined}\nDraft: {draft}" + ) + try: _, _ = linted_class_block_from_spec(cfg, merged_spec, timing_key=seg_name) except SceneGenerationError as exc: @@ -536,6 +569,15 @@ def generate_scene_spec( existing = scenes_path.read_text(encoding="utf-8") if scenes_path.exists() else "" reference_scenes = extract_reference_classes(existing) snippets = collect_source_snippets(cfg, settings, extra_paths=extra_paths) + corpus_parts = [narration_text] + for h in list(settings.hints) + list(extra_hints): + if str(h).strip(): + corpus_parts.append(str(h).strip()) + for label, text in snippets: + body = str(text or "").strip() + if body: + corpus_parts.append(f"{label}\n{body}") + corpus_text = "\n\n".join(p for p in corpus_parts if str(p).strip()) system_prompt = scene_spec_system_prompt(cfg, seg_id) user_message = build_scene_spec_user_message( @@ -578,13 +620,23 @@ def generate_scene_spec( for attempt in range(3): msg = user_message if attempt > 0 and last_sparse is not None: - msg = ( - f"{user_message}\n\n--- RETRY: SUBJECT-BEAT COVERAGE FAILED ---\n" - f"{last_sparse}\n" - f"Cover each of the {n_beats} subject beats with a spoken-phrase label. " - "Hold the board across sentences in the same beat; add a new label only " - "when the topic shifts. Do not invent unspoken diagram terms." - ) + err = str(last_sparse) + if "image prompt alignment" in err: + msg = ( + f"{user_message}\n\n--- RETRY: IMAGE PROMPT ALIGNMENT FAILED ---\n" + f"{last_sparse}\n" + "Rewrite each image prompt so it uses documented terms from the " + "narration and SOURCE DOCUMENTATION. Do not invent product names " + "or generic 'a diagram' artwork with no subject." + ) + else: + msg = ( + f"{user_message}\n\n--- RETRY: SUBJECT-BEAT COVERAGE FAILED ---\n" + f"{last_sparse}\n" + f"Cover each of the {n_beats} subject beats with a spoken-phrase label. " + "Hold the board across sentences in the same beat; add a new label only " + "when the topic shifts. Do not invent unspoken diagram terms." + ) try: raw = invoke( system_prompt=system_prompt, @@ -612,12 +664,22 @@ def generate_scene_spec( raw=raw, enforce_density=True, density_slack=0, + corpus_text=corpus_text, ) last_sparse = None break except SceneGenerationError as exc: - if "subject-beat coverage" not in str(exc): + err = str(exc) + retryable = ( + "subject-beat coverage" in err or "image prompt alignment" in err + ) + if not retryable: raise + if "image prompt alignment" in err: + if attempt >= 2: + raise + last_sparse = exc + continue # Near-miss: accept without another LLM call when close enough. try: merged_spec = _parse_and_harden_llm_spec( @@ -630,6 +692,7 @@ def generate_scene_spec( raw=raw, enforce_density=True, density_slack=near_miss_slack, + corpus_text=corpus_text, ) last_sparse = None break diff --git a/src/docgen/validate.py b/src/docgen/validate.py index e79031b..0c8b672 100644 --- a/src/docgen/validate.py +++ b/src/docgen/validate.py @@ -362,6 +362,8 @@ def validate_segment( if self.config.visual_map.get(seg_id, {}).get("type") == "manim": report.checks.append(self._check_manim_scene_lint()) report.checks.append(self._check_subject_beat_coverage(seg_id)) + report.checks.append(self._check_image_prompt_alignment(seg_id)) + report.checks.append(self._check_image_asset_alignment(seg_id)) report.checks.append(self._check_scene_assets(seg_id)) return report.to_dict() @@ -372,9 +374,10 @@ def run_pre_push(self) -> None: Missing recordings are reported as warnings, not failures — a project that hasn't generated videos yet should still be pushable. Quality checks on *existing* recordings — including visual-sync (``av_sync``, - ``subject_beat_coverage``, ``ocr_scan``, ``layout``, ``freeze_ratio``) - — and narration lint are hard failures so ``generate-all`` cannot - print ``Pipeline complete`` after a desynced mux. + ``subject_beat_coverage``, ``image_prompt_alignment``, + ``image_asset_alignment``, ``ocr_scan``, ``layout``, ``freeze_ratio``) + — and narration lint are hard failures so + ``generate-all`` cannot print ``Pipeline complete`` after a desynced mux. """ reports = self.run_all() hard_fail = False @@ -646,6 +649,161 @@ def _check_subject_beat_coverage(self, seg_id: str) -> CheckResult: ["Subject beats covered; no invented unspoken labels"], ) + def _check_image_prompt_alignment(self, seg_id: str) -> CheckResult: + """Ensure scene-spec image prompts share documented terms (not generic art).""" + if not self.config.image_prompt_alignment_enabled: + return CheckResult( + "image_prompt_alignment", + True, + ["validation.image_prompt_alignment disabled in config (skipped)"], + ) + + seg_name = self.config.resolve_segment_name(seg_id) + spec_path = self.config.animations_dir / "specs" / f"{seg_name}.scene.yaml" + if not spec_path.is_file(): + return _missing_manim_spec_result("image_prompt_alignment", spec_path.name) + + import yaml + + from docgen.image_generate import collect_alignment_corpus + from docgen.scene_spec import image_prompt_alignment_violations, iter_image_elements + + try: + raw = yaml.safe_load(spec_path.read_text(encoding="utf-8")) + except (OSError, yaml.YAMLError) as exc: + return CheckResult( + "image_prompt_alignment", + False, + [f"could not load {spec_path.name}: {exc}"], + ) + if not isinstance(raw, dict): + return CheckResult( + "image_prompt_alignment", + False, + [f"{spec_path.name}: root must be a mapping"], + ) + if not iter_image_elements(raw): + return CheckResult( + "image_prompt_alignment", + True, + ["No image elements (skipped)"], + ) + + corpus = collect_alignment_corpus(self.config, raw) + issues = image_prompt_alignment_violations(raw, corpus_text=corpus) + if issues: + return CheckResult("image_prompt_alignment", False, issues) + if not corpus.strip(): + return CheckResult( + "image_prompt_alignment", + True, + ["No narration/source corpus yet — image prompts not checked"], + ) + return CheckResult( + "image_prompt_alignment", + True, + ["Image prompts share documented terms with narration/source"], + ) + + def _check_image_asset_alignment(self, seg_id: str) -> CheckResult: + """OCR (and optional vision) of generated scene-spec PNGs vs documentation.""" + iaa = self.config.image_asset_alignment_config + if not iaa.get("enabled", True): + return CheckResult( + "image_asset_alignment", + True, + ["validation.image_asset_alignment disabled in config (skipped)"], + ) + + seg_name = self.config.resolve_segment_name(seg_id) + spec_path = self.config.animations_dir / "specs" / f"{seg_name}.scene.yaml" + if not spec_path.is_file(): + return _missing_manim_spec_result("image_asset_alignment", spec_path.name) + + import yaml + + from docgen.image_align import ocr_image_text, review_image_against_docs + from docgen.image_generate import collect_alignment_corpus + from docgen.scene_spec import image_ocr_alignment_violations, iter_image_elements + + try: + raw = yaml.safe_load(spec_path.read_text(encoding="utf-8")) + except (OSError, yaml.YAMLError) as exc: + return CheckResult( + "image_asset_alignment", + False, + [f"could not load {spec_path.name}: {exc}"], + ) + if not isinstance(raw, dict): + return CheckResult( + "image_asset_alignment", + False, + [f"{spec_path.name}: root must be a mapping"], + ) + elements = iter_image_elements(raw) + if not elements: + return CheckResult( + "image_asset_alignment", + True, + ["No image elements (skipped)"], + ) + + corpus = collect_alignment_corpus(self.config, raw) + if not corpus.strip(): + return CheckResult( + "image_asset_alignment", + True, + ["No narration/source corpus yet — image pixels not checked"], + ) + + issues: list[str] = [] + scanned_any = False + for el in elements: + rel = str(el.get("image") or "").strip() + if not rel: + continue + asset = self.config.base_dir / rel + if not asset.is_file(): + continue + scanned_any = True + if iaa.get("ocr", True): + unavail = _tesseract_unavailable_detail() + if unavail: + return CheckResult("image_asset_alignment", False, [unavail]) + ocr_text = ocr_image_text(asset) + if ocr_text is None: + issues.append(f"{rel}: could not OCR image (unreadable file)") + elif ocr_text: + issues.extend( + image_ocr_alignment_violations( + ocr_text, corpus_text=corpus, relpath=rel + ) + ) + if iaa.get("review"): + verdict = review_image_against_docs( + asset, + corpus_text=corpus, + authored_prompt=str(el.get("prompt") or ""), + label=str(el.get("label") or ""), + cfg=self.config, + ) + if not verdict.passed: + issues.append(f"{rel}: vision review FAIL — {verdict.reason}") + + if issues: + return CheckResult("image_asset_alignment", False, issues) + if not scanned_any: + return CheckResult( + "image_asset_alignment", + True, + ["Image assets not on disk yet (skipped) — run `docgen image-generate`"], + ) + return CheckResult( + "image_asset_alignment", + True, + ["Generated image pixels match documented terms"], + ) + def _check_scene_assets(self, seg_id: str) -> CheckResult: """Pre-render stuck / overlap / font / compile-sync gate (no video required).""" sa_cfg = self.config.scene_assets_config diff --git a/src/docgen/yaml_generate.py b/src/docgen/yaml_generate.py index c8163ac..82272ef 100644 --- a/src/docgen/yaml_generate.py +++ b/src/docgen/yaml_generate.py @@ -220,6 +220,8 @@ def merge_defaults( raw["image_generation"] = { "model": "gpt-image-1", "size": "1536x1024", + "align_with_docs": True, + "align_review": True, } changes.append( "image_generation: added model gpt-image-1 (Cursor/OpenAI Images; override with --model)" diff --git a/tests/test_ai_client.py b/tests/test_ai_client.py index aece8ae..e9a7757 100644 --- a/tests/test_ai_client.py +++ b/tests/test_ai_client.py @@ -15,6 +15,7 @@ DEFAULT_GROK_IMAGE_MODEL, GROK_BASE_URL, chat_completion, + chat_completion_with_image, openai_client, resolve_ai_settings, resolve_chat_model, @@ -137,6 +138,41 @@ class _Resp: assert captured["model"] == DEFAULT_GROK_CHAT_MODEL +def test_chat_completion_with_image_sends_data_uri( + tmp_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + monkeypatch.setenv("CURSOR_API_KEY", "sk-proj-cursor") + captured: dict = {} + + class _Msg: + content = "PASS\nok" + + class _Choice: + message = _Msg() + + class _Resp: + choices = [_Choice()] + + fake = MagicMock() + fake.chat.completions.create.side_effect = lambda **kw: captured.update(kw) or _Resp() + + with patch("docgen.ai_client.openai_client", return_value=fake): + out = chat_completion_with_image( + system_prompt="sys", + user_message="review this", + image_bytes=b"png-bytes", + media_type="image/png", + model="gpt-4o", + temperature=0.0, + cfg=_cfg(tmp_path, {}), + ) + assert out == "PASS\nok" + content = captured["messages"][1]["content"] + assert content[0]["type"] == "text" + assert content[1]["type"] == "image_url" + assert content[1]["image_url"]["url"].startswith("data:image/png;base64,") + + def test_openai_empty_chat_content_raises( tmp_path: Path, monkeypatch: pytest.MonkeyPatch, clear_ai_env: None ) -> None: @@ -583,6 +619,34 @@ def _http(url: str, *, data: bytes, headers: dict, **_kwargs) -> bytes: assert captured["payload"]["model"] == DEFAULT_ANTHROPIC_CHAT_MODEL +def test_anthropic_vision_chat_posts_image_block( + tmp_path: Path, monkeypatch: pytest.MonkeyPatch, clear_ai_env: None +) -> None: + monkeypatch.setenv("ANTHROPIC_API_KEY", "sk-ant-test") + cfg = _cfg(tmp_path, {"ai": {"provider": "anthropic"}}) + captured: dict = {} + + def _http(url: str, *, data: bytes, headers: dict, **_kwargs) -> bytes: + captured["payload"] = json.loads(data.decode()) + return json.dumps({"content": [{"type": "text", "text": "PASS\nok"}]}).encode() + + with patch("docgen.ai_client._http_with_retries", side_effect=_http): + out = chat_completion_with_image( + system_prompt="sys", + user_message="review", + image_bytes=b"png-bytes", + media_type="image/png", + model="claude-sonnet-4-5", + temperature=0.0, + cfg=cfg, + ) + assert out == "PASS\nok" + content = captured["payload"]["messages"][0]["content"] + assert content[0]["type"] == "image" + assert content[0]["source"]["type"] == "base64" + assert content[1]["type"] == "text" + + def test_anthropic_chat_honors_base_url( tmp_path: Path, monkeypatch: pytest.MonkeyPatch, clear_ai_env: None ) -> None: diff --git a/tests/test_config.py b/tests/test_config.py index cb79082..bc88250 100644 --- a/tests/test_config.py +++ b/tests/test_config.py @@ -1136,7 +1136,9 @@ def test_from_yaml_validation_enable_bools_allowed(tmp_path: Path) -> None: " scene_assets:\n enabled: true\n" " story_end:\n enabled: false\n" " layout:\n check_overlap: false\n" - " subject_beat_coverage:\n enabled: false\n", + " subject_beat_coverage:\n enabled: false\n" + " image_prompt_alignment:\n enabled: false\n" + " image_asset_alignment:\n enabled: false\n ocr: true\n review: false\n", encoding="utf-8", ) c = Config.from_yaml(p) @@ -1148,3 +1150,53 @@ def test_from_yaml_validation_enable_bools_allowed(tmp_path: Path) -> None: assert c.story_end_config["enabled"] is False assert c.layout_config["check_overlap"] is False assert c.subject_beat_coverage_enabled is False + assert c.image_prompt_alignment_enabled is False + assert c.image_asset_alignment_config["enabled"] is False + assert c.image_asset_alignment_config["ocr"] is True + assert c.image_asset_alignment_config["review"] is False + + +def test_from_yaml_int_image_prompt_alignment_enabled_raises(tmp_path: Path) -> None: + p = tmp_path / "docgen.yaml" + p.write_text( + "validation:\n image_prompt_alignment:\n enabled: 0\n", + encoding="utf-8", + ) + with pytest.raises( + ConfigError, + match="validation.image_prompt_alignment.enabled must be a YAML boolean", + ): + Config.from_yaml(p) + + +def test_from_yaml_int_image_asset_alignment_enabled_raises(tmp_path: Path) -> None: + p = tmp_path / "docgen.yaml" + p.write_text( + "validation:\n image_asset_alignment:\n enabled: 0\n", + encoding="utf-8", + ) + with pytest.raises( + ConfigError, + match="validation.image_asset_alignment.enabled must be a YAML boolean", + ): + Config.from_yaml(p) + + +def test_from_yaml_int_image_align_review_raises(tmp_path: Path) -> None: + p = tmp_path / "docgen.yaml" + p.write_text("image_generation:\n align_review: 1\n", encoding="utf-8") + with pytest.raises( + ConfigError, + match="image_generation.align_review must be a YAML boolean", + ): + Config.from_yaml(p) + + +def test_from_yaml_int_image_align_with_docs_raises(tmp_path: Path) -> None: + p = tmp_path / "docgen.yaml" + p.write_text("image_generation:\n align_with_docs: 1\n", encoding="utf-8") + with pytest.raises( + ConfigError, + match="image_generation.align_with_docs must be a YAML boolean", + ): + Config.from_yaml(p) diff --git a/tests/test_image_align.py b/tests/test_image_align.py new file mode 100644 index 0000000..5842b8a --- /dev/null +++ b/tests/test_image_align.py @@ -0,0 +1,73 @@ +"""Pixel-level image alignment: OCR verdicts and vision PASS/FAIL parsing.""" + +from __future__ import annotations + +from pathlib import Path + +from docgen.image_align import ( + build_review_user_message, + parse_review_verdict, + review_image_against_docs, +) + + +def test_parse_review_verdict_pass_and_fail() -> None: + ok = parse_review_verdict("PASS\nDepicts the checkout service lock.") + assert ok.passed is True + assert "checkout" in ok.reason + bad = parse_review_verdict("FAIL\nShows a WidgetX console instead.") + assert bad.passed is False + assert "WidgetX" in bad.reason + + +def test_parse_review_verdict_unparseable_fails_closed() -> None: + out = parse_review_verdict("looks fine to me") + assert out.passed is False + assert "PASS/FAIL" in out.reason + + +def test_parse_review_verdict_empty_fails_closed() -> None: + out = parse_review_verdict(" ") + assert out.passed is False + + +def test_build_review_user_message_includes_docs() -> None: + msg = build_review_user_message( + corpus_text="The checkout service owns the cart lock.", + authored_prompt="clean diagram of the checkout service", + label="checkout service", + ) + assert "DOCUMENTED SUBJECT" in msg + assert "cart lock" in msg + assert "AUTHORED IMAGE PROMPT" in msg + assert "ON-SCREEN LABEL" in msg + + +def test_review_image_against_docs_uses_chat_fn(tmp_path: Path) -> None: + png = tmp_path / "x.png" + png.write_bytes(b"\x89PNG\r\n\x1a\nnot-a-real-png") + captured: dict = {} + + def _chat(**kwargs: object) -> str: + captured.update(kwargs) + return "PASS\nMatches the checkout service." + + verdict = review_image_against_docs( + png, + corpus_text="The checkout service owns the cart lock.", + authored_prompt="diagram of checkout", + chat_fn=_chat, + ) + assert verdict.passed is True + assert captured["image_bytes"] == png.read_bytes() + assert "checkout service" in str(captured["user_message"]) + + +def test_review_empty_file_fails(tmp_path: Path) -> None: + png = tmp_path / "empty.png" + png.write_bytes(b"") + verdict = review_image_against_docs( + png, corpus_text="checkout service", chat_fn=lambda **_k: "PASS\nok" + ) + assert verdict.passed is False + assert "empty" in verdict.reason diff --git a/tests/test_image_generate.py b/tests/test_image_generate.py index ea0df85..6e52fb0 100644 --- a/tests/test_image_generate.py +++ b/tests/test_image_generate.py @@ -8,8 +8,11 @@ import yaml from docgen.config import Config +from docgen.image_align import ImageReviewResult from docgen.image_generate import ( + DEFAULT_IMAGE_STYLE, ImageGenerationError, + build_aligned_image_prompt, generate_images_for_spec, generate_missing_images_for_bundle, spec_files_for_bundle, @@ -82,6 +85,8 @@ def test_dry_run_reports_without_writing(cfg: Config) -> None: results = generate_images_for_spec(cfg, spec, dry_run=True, image_fn=lambda p: _PNG_BYTES) assert [r.status for r in results] == ["dry-run"] assert results[0].prompt == "a diagram" + assert DEFAULT_IMAGE_STYLE.split(".")[0] in results[0].effective_prompt + assert "Illustration request:" in results[0].effective_prompt assert not (cfg.base_dir / "images" / "arch.png").exists() @@ -125,6 +130,175 @@ def test_no_specs_dir_is_noop(cfg: Config) -> None: assert generate_missing_images_for_bundle(cfg, image_fn=lambda p: _PNG_BYTES) == [] +def test_aligned_prompt_is_what_the_provider_sees(cfg: Config) -> None: + spec = _write_spec( + cfg.animations_dir / "specs" / "01-x.scene.yaml", + prompt="clean diagram of the bootstrap pipeline", + ) + (cfg.narration_dir).mkdir(parents=True, exist_ok=True) + (cfg.narration_dir / "1.md").write_text( + "The bootstrap pipeline seeds the cluster.\n", encoding="utf-8" + ) + seen: list[str] = [] + + def _capture(prompt: str) -> bytes: + seen.append(prompt) + return _PNG_BYTES + + generate_images_for_spec(cfg, spec, image_fn=_capture) + assert seen + assert "bootstrap pipeline" in seen[0] + assert "Documented subject" in seen[0] + assert seen[0] != "clean diagram of the bootstrap pipeline" + + +def test_unaligned_prompt_fails_before_provider(cfg: Config) -> None: + spec = _write_spec( + cfg.animations_dir / "specs" / "01-x.scene.yaml", + prompt="isometric render of the WidgetX orchestrator", + ) + cfg.narration_dir.mkdir(parents=True, exist_ok=True) + (cfg.narration_dir / "1.md").write_text( + "The bootstrap pipeline seeds the cluster.\n", encoding="utf-8" + ) + with pytest.raises(ImageGenerationError, match="image prompt alignment"): + generate_images_for_spec(cfg, spec, image_fn=lambda p: _PNG_BYTES) + assert not (cfg.base_dir / "images" / "arch.png").exists() + + +def test_align_with_docs_false_sends_authored_prompt(tmp_path: Path) -> None: + (tmp_path / "docgen.yaml").write_text( + yaml.dump( + { + "segments": {"all": ["1"]}, + "image_generation": {"align_with_docs": False}, + } + ), + encoding="utf-8", + ) + cfg = Config.from_yaml(tmp_path / "docgen.yaml") + spec = _write_spec( + cfg.animations_dir / "specs" / "01-x.scene.yaml", + prompt="isometric render of the WidgetX orchestrator", + ) + cfg.narration_dir.mkdir(parents=True, exist_ok=True) + (cfg.narration_dir / "1.md").write_text( + "The bootstrap pipeline seeds the cluster.\n", encoding="utf-8" + ) + seen: list[str] = [] + generate_images_for_spec(cfg, spec, image_fn=lambda p: seen.append(p) or _PNG_BYTES) + assert seen == ["isometric render of the WidgetX orchestrator"] + + +def test_source_docs_ground_prompt_without_narration(tmp_path: Path) -> None: + (tmp_path / "docs").mkdir() + (tmp_path / "docs" / "arch.md").write_text( + "The checkout service owns the cart lock.\n", encoding="utf-8" + ) + (tmp_path / "docgen.yaml").write_text( + yaml.dump( + { + "repo_root": ".", + "segments": {"all": ["1"]}, + "manim_scene_generation": {"context": {"paths": ["docs/arch.md"]}}, + } + ), + encoding="utf-8", + ) + cfg = Config.from_yaml(tmp_path / "docgen.yaml") + spec = _write_spec( + cfg.animations_dir / "specs" / "01-x.scene.yaml", + prompt="clean diagram of the checkout service", + ) + seen: list[str] = [] + generate_images_for_spec(cfg, spec, image_fn=lambda p: seen.append(p) or _PNG_BYTES) + assert seen + assert "checkout service" in seen[0] + assert "cart lock" in seen[0] + + +def test_ocr_invented_text_deletes_asset(cfg: Config) -> None: + spec = _write_spec( + cfg.animations_dir / "specs" / "01-x.scene.yaml", + prompt="clean diagram of the bootstrap pipeline", + ) + cfg.narration_dir.mkdir(parents=True, exist_ok=True) + (cfg.narration_dir / "1.md").write_text( + "The bootstrap pipeline seeds the cluster.\n", encoding="utf-8" + ) + with pytest.raises(ImageGenerationError, match="pixel alignment"): + generate_images_for_spec( + cfg, + spec, + image_fn=lambda p: _PNG_BYTES, + ocr_fn=lambda _path: "WidgetX Orchestrator console", + ) + assert not (cfg.base_dir / "images" / "arch.png").exists() + + +def test_vision_fail_retries_then_keeps_pass(cfg: Config) -> None: + spec = _write_spec( + cfg.animations_dir / "specs" / "01-x.scene.yaml", + prompt="clean diagram of the bootstrap pipeline", + ) + cfg.narration_dir.mkdir(parents=True, exist_ok=True) + (cfg.narration_dir / "1.md").write_text( + "The bootstrap pipeline seeds the cluster.\n", encoding="utf-8" + ) + seen: list[str] = [] + reviews = iter( + [ + ImageReviewResult(False, "shows a generic city skyline"), + ImageReviewResult(True, "now shows the bootstrap pipeline"), + ] + ) + + def _review(*_a: object, **_k: object) -> ImageReviewResult: + return next(reviews) + + generate_images_for_spec( + cfg, + spec, + image_fn=lambda p: seen.append(p) or _PNG_BYTES, + review_fn=_review, + ocr_fn=lambda _path: "", + ) + assert len(seen) == 2 + assert "PIXEL REVIEW FAILED" in seen[1] + assert (cfg.base_dir / "images" / "arch.png").read_bytes() == _PNG_BYTES + + +def test_vision_fail_exhausted_deletes_asset(cfg: Config) -> None: + spec = _write_spec( + cfg.animations_dir / "specs" / "01-x.scene.yaml", + prompt="clean diagram of the bootstrap pipeline", + ) + cfg.narration_dir.mkdir(parents=True, exist_ok=True) + (cfg.narration_dir / "1.md").write_text( + "The bootstrap pipeline seeds the cluster.\n", encoding="utf-8" + ) + with pytest.raises(ImageGenerationError, match="vision review FAIL"): + generate_images_for_spec( + cfg, + spec, + image_fn=lambda p: _PNG_BYTES, + review_fn=lambda *_a, **_k: ImageReviewResult(False, "wrong subject"), + ocr_fn=lambda _path: "", + ) + assert not (cfg.base_dir / "images" / "arch.png").exists() + + +def test_build_aligned_image_prompt_includes_corpus_and_label() -> None: + out = build_aligned_image_prompt( + "clean diagram of the bootstrap pipeline", + corpus_text="The bootstrap pipeline seeds the cluster.", + label="bootstrap", + ) + assert "bootstrap pipeline" in out + assert "On-screen timing label" in out + assert "Illustration request:" in out + + def test_cli_image_generate_all_fails_when_manim_has_no_specs(tmp_path: Path) -> None: from click.testing import CliRunner diff --git a/tests/test_scene_spec.py b/tests/test_scene_spec.py index 6471196..f8f00bc 100644 --- a/tests/test_scene_spec.py +++ b/tests/test_scene_spec.py @@ -17,6 +17,8 @@ disk_spec_with_merged_wait_words, cluster_subject_beats, count_spec_labels, + image_ocr_alignment_violations, + image_prompt_alignment_violations, iter_paced_label_anchors, last_paced_reveal_time, layout_budget_violations, @@ -569,6 +571,54 @@ def test_layout_stack_budget_decreases_with_larger_title_font() -> None: assert b_small > b_large +def test_image_prompt_alignment_requires_documented_terms() -> None: + spec = { + "title": {"text": "T", "font_size": 36, "color": "C_WHITE"}, + "rows": [ + { + "run_time": 1.0, + "boxes": [ + { + "image": "images/arch.png", + "width": 4.0, + "height": 2.5, + "prompt": "clean flat diagram of the bootstrap pipeline", + } + ], + } + ], + } + narr = "The bootstrap pipeline seeds the cluster." + assert image_prompt_alignment_violations(spec, corpus_text=narr) == [] + + spec["rows"][0]["boxes"][0]["prompt"] = "a clean flat illustration" + issues = image_prompt_alignment_violations(spec, corpus_text=narr) + assert issues + assert any("only visual style" in msg for msg in issues) + + spec["rows"][0]["boxes"][0]["prompt"] = "isometric render of the WidgetX orchestrator" + issues = image_prompt_alignment_violations(spec, corpus_text=narr) + assert issues + assert any("shares no documented terms" in msg for msg in issues) + + assert image_prompt_alignment_violations(spec, corpus_text="") == [] + + +def test_image_ocr_alignment_flags_invented_on_image_text() -> None: + narr = "The bootstrap pipeline seeds the cluster." + assert image_ocr_alignment_violations( + "bootstrap pipeline", corpus_text=narr, relpath="images/arch.png" + ) == [] + assert image_ocr_alignment_violations("", corpus_text=narr) == [] + issues = image_ocr_alignment_violations( + "WidgetX Orchestrator console", + corpus_text=narr, + relpath="images/arch.png", + ) + assert issues + assert any("OCR" in msg for msg in issues) + + def test_subject_beat_coverage_allows_dwell_rejects_missed_topics() -> None: narr = ( "The bootstrap pipeline seeds the cluster. " diff --git a/tests/test_scene_spec_generate.py b/tests/test_scene_spec_generate.py index 32f0a10..199c0d7 100644 --- a/tests/test_scene_spec_generate.py +++ b/tests/test_scene_spec_generate.py @@ -370,6 +370,61 @@ def test_user_message_includes_computed_layout_stack_budgets() -> None: assert "13.22" in msg # horizontal safe width (FRAME_WIDTH - 1.0) +def test_user_message_includes_source_snippets() -> None: + msg = build_scene_spec_user_message( + seg_id="01", + seg_name="01-x", + class_name="XScene", + narration_text="The bootstrap pipeline seeds the cluster.", + timing_enrichment="(no timing)", + hints=[], + extra_hints=[], + reference_scenes="", + source_snippets=[("docs/architecture.md", "The bootstrap pipeline writes a lockfile.")], + ) + assert "SOURCE DOCUMENTATION" in msg + assert "docs/architecture.md" in msg + assert "writes a lockfile" in msg + assert "Do not invent product names" in msg + + +def test_generate_rejects_unaligned_image_prompt(tmp_path: Path) -> None: + cfg = _bundle(tmp_path) + + def fake_llm(**_kwargs: object) -> str: + return """```yaml +segment_id: "08" +class_name: ExtrasScene +title: + text: "Synthetic" + font_size: 40 + color: C_WHITE +rows: + - run_time: 1.2 + boxes: + - label: "Hello" + color: C_ORANGE + width: 4.0 + height: 1.0 + font_size: 20 + - image: images/hello.png + width: 4.0 + height: 2.0 + label: "Hello" + prompt: "isometric render of the WidgetX orchestrator" +```""" + + with pytest.raises(SceneGenerationError, match="image prompt alignment"): + generate_scene_spec( + cfg, + "08", + extra_paths=[], + extra_hints=[], + dry_run=False, + llm=fake_llm, + ) + + def test_scene_spec_generate_all_uses_default_when_all_missing(tmp_path: Path) -> None: from click.testing import CliRunner diff --git a/tests/test_validate.py b/tests/test_validate.py index e2fb80e..0ca1dd0 100644 --- a/tests/test_validate.py +++ b/tests/test_validate.py @@ -623,6 +623,159 @@ def test_validate_passes_covered_beats(self, cfg_dir: Path) -> None: assert check["passed"], check["details"] +class TestImagePromptAlignmentValidate: + def test_validate_flags_unaligned_image_prompt(self, cfg_dir: Path) -> None: + cfg_raw = yaml.safe_load((cfg_dir / "docgen.yaml").read_text(encoding="utf-8")) + cfg_raw["visual_map"]["01"] = {"type": "manim", "source": "Scene01.mp4"} + cfg_raw.setdefault("segment_names", {})["01"] = "01-test" + (cfg_dir / "docgen.yaml").write_text(yaml.dump(cfg_raw), encoding="utf-8") + (cfg_dir / "narration" / "01-test.md").write_text( + "The bootstrap pipeline seeds the cluster.\n", + encoding="utf-8", + ) + specs = cfg_dir / "animations" / "specs" + specs.mkdir(parents=True, exist_ok=True) + (specs / "01-test.scene.yaml").write_text( + yaml.dump( + { + "segment_id": "01", + "class_name": "DemoScene", + "title": {"text": "T", "font_size": 36, "color": "C_WHITE"}, + "rows": [ + { + "run_time": 1.0, + "boxes": [ + { + "label": "bootstrap pipeline", + "color": "C_ORANGE", + "width": 4.0, + "height": 1.0, + "font_size": 18, + }, + { + "image": "images/arch.png", + "width": 4.0, + "height": 2.0, + "prompt": "isometric render of the WidgetX orchestrator", + }, + ], + } + ], + } + ), + encoding="utf-8", + ) + config = Config.from_yaml(cfg_dir / "docgen.yaml") + report = Validator(config).validate_segment("01") + check = next(c for c in report["checks"] if c["name"] == "image_prompt_alignment") + assert not check["passed"] + assert any("documented terms" in d for d in check["details"]) + + def test_validate_passes_grounded_image_prompt(self, cfg_dir: Path) -> None: + cfg_raw = yaml.safe_load((cfg_dir / "docgen.yaml").read_text(encoding="utf-8")) + cfg_raw["visual_map"]["01"] = {"type": "manim", "source": "Scene01.mp4"} + cfg_raw.setdefault("segment_names", {})["01"] = "01-test" + (cfg_dir / "docgen.yaml").write_text(yaml.dump(cfg_raw), encoding="utf-8") + (cfg_dir / "narration" / "01-test.md").write_text( + "The bootstrap pipeline seeds the cluster.\n", + encoding="utf-8", + ) + specs = cfg_dir / "animations" / "specs" + specs.mkdir(parents=True, exist_ok=True) + (specs / "01-test.scene.yaml").write_text( + yaml.dump( + { + "segment_id": "01", + "class_name": "DemoScene", + "title": {"text": "T", "font_size": 36, "color": "C_WHITE"}, + "rows": [ + { + "run_time": 1.0, + "boxes": [ + { + "label": "bootstrap pipeline", + "color": "C_ORANGE", + "width": 4.0, + "height": 1.0, + "font_size": 18, + }, + { + "image": "images/arch.png", + "width": 4.0, + "height": 2.0, + "prompt": "clean diagram of the bootstrap pipeline", + }, + ], + } + ], + } + ), + encoding="utf-8", + ) + config = Config.from_yaml(cfg_dir / "docgen.yaml") + report = Validator(config).validate_segment("01") + check = next(c for c in report["checks"] if c["name"] == "image_prompt_alignment") + assert check["passed"], check["details"] + + def test_validate_flags_invented_ocr_on_asset(self, cfg_dir: Path, monkeypatch) -> None: + cfg_raw = yaml.safe_load((cfg_dir / "docgen.yaml").read_text(encoding="utf-8")) + cfg_raw["visual_map"]["01"] = {"type": "manim", "source": "Scene01.mp4"} + cfg_raw.setdefault("segment_names", {})["01"] = "01-test" + (cfg_dir / "docgen.yaml").write_text(yaml.dump(cfg_raw), encoding="utf-8") + (cfg_dir / "narration" / "01-test.md").write_text( + "The bootstrap pipeline seeds the cluster.\n", + encoding="utf-8", + ) + specs = cfg_dir / "animations" / "specs" + specs.mkdir(parents=True, exist_ok=True) + (specs / "01-test.scene.yaml").write_text( + yaml.dump( + { + "segment_id": "01", + "class_name": "DemoScene", + "title": {"text": "T", "font_size": 36, "color": "C_WHITE"}, + "rows": [ + { + "run_time": 1.0, + "boxes": [ + { + "label": "bootstrap pipeline", + "color": "C_ORANGE", + "width": 4.0, + "height": 1.0, + "font_size": 18, + }, + { + "image": "images/arch.png", + "width": 4.0, + "height": 2.0, + "prompt": "clean diagram of the bootstrap pipeline", + }, + ], + } + ], + } + ), + encoding="utf-8", + ) + asset = cfg_dir / "images" / "arch.png" + asset.parent.mkdir(parents=True, exist_ok=True) + asset.write_bytes(b"\x89PNG\r\n\x1a\nfake") + monkeypatch.setattr( + "docgen.image_align.ocr_image_text", + lambda _path: "WidgetX Orchestrator console", + ) + monkeypatch.setattr( + "docgen.validate._tesseract_unavailable_detail", + lambda: None, + ) + config = Config.from_yaml(cfg_dir / "docgen.yaml") + report = Validator(config).validate_segment("01") + check = next(c for c in report["checks"] if c["name"] == "image_asset_alignment") + assert not check["passed"] + assert any("OCR" in d for d in check["details"]) + + # ── ffprobe JSON probes honor returncode ────────────────────────────── class _FakeProbe: @@ -686,7 +839,7 @@ def test_drift_pass_when_ffprobe_succeeds(self, config, tmp_path, monkeypatch): @pytest.mark.parametrize( "check_name", - ("av_sync", "subject_beat_coverage", "ocr_scan", "layout", "freeze_ratio"), + ("av_sync", "subject_beat_coverage", "image_prompt_alignment", "image_asset_alignment", "ocr_scan", "layout", "freeze_ratio"), ) def test_run_pre_push_visual_sync_fail_is_hard(check_name: str, capsys) -> None: """Visual-sync FAILs must be FAIL + SystemExit, not WARN (leftover #2).""" diff --git a/tests/test_validate_timing_sync.py b/tests/test_validate_timing_sync.py index 2775ca4..eb75544 100644 --- a/tests/test_validate_timing_sync.py +++ b/tests/test_validate_timing_sync.py @@ -435,7 +435,10 @@ def _write_paced_hand_authored_scenes(cfg: Config) -> None: ) -@pytest.mark.parametrize("check_name", ("story_end", "subject_beat_coverage")) +@pytest.mark.parametrize( + "check_name", + ("story_end", "subject_beat_coverage", "image_prompt_alignment", "image_asset_alignment"), +) def test_hand_authored_paced_manim_without_spec_fails(cfg, check_name: str) -> None: """Manim + paced timed_play + no *.scene.yaml must fail, not skip-PASS.""" _write_paced_hand_authored_scenes(cfg) @@ -446,6 +449,10 @@ def test_hand_authored_paced_manim_without_spec_fails(cfg, check_name: str) -> N v = Validator(cfg) if check_name == "story_end": check = v._check_story_end("01") + elif check_name == "image_prompt_alignment": + check = v._check_image_prompt_alignment("01") + elif check_name == "image_asset_alignment": + check = v._check_image_asset_alignment("01") else: check = v._check_subject_beat_coverage("01") assert check.name == check_name diff --git a/tests/test_yaml_generate.py b/tests/test_yaml_generate.py index 8041e4f..0bf7781 100644 --- a/tests/test_yaml_generate.py +++ b/tests/test_yaml_generate.py @@ -85,6 +85,8 @@ def test_merge_defaults_adds_archive_exclude(tmp_path: Path) -> None: assert raw["ai"]["provider"] == "openai" assert raw["image_generation"]["model"] == "gpt-image-1" assert raw["image_generation"]["size"] == "1536x1024" + assert raw["image_generation"]["align_with_docs"] is True + assert raw["image_generation"]["align_review"] is True def test_merge_defaults_idempotent_archive(tmp_path: Path) -> None: