Skip to content

Ground image prompts in documentation and fail unaligned artwork - #169

Draft
jmjava wants to merge 3 commits into
mainfrom
cursor/image-doc-alignment-4f6d
Draft

jmjava wants to merge 3 commits into
mainfrom
cursor/image-doc-alignment-4f6d

Conversation

@jmjava

@jmjava jmjava commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Image quality was a caption-only hop: scene-spec-generate collected source snippets and never sent them, and image-generate forwarded the authored prompt: to the Images API with no check that it named anything from the docs — and no check that the PNG pixels matched either.

This change closes both gaps the same way subject-beat coverage already works for box labels.

Prompt alignment (first commit)

  • scene-spec-generate now includes --- SOURCE DOCUMENTATION --- in the user message (the snippets were collected and dropped).
  • Image prompt: text must share documented terms with narration/source. Style-only prompts and invented subjects fail closed.
  • image-generate wraps each Images API call with narration, source snippets, the box label, and an educational-diagram style prefix (image_generation.align_with_docs, default true).

Pixel alignment (follow-up)

  • After the PNG is written, OCR rejects invented on-image labels (WidgetX orchestrator vs “checkout service”).
  • A vision review (OpenAI / Grok / Claude chat_completion_with_image) answers PASS/FAIL against the same corpus. FAIL retries once with the critique appended to the prompt, then deletes the asset (image_generation.align_review, default true).
  • validate / --pre-push add image_asset_alignment (OCR on by default; vision opt-in via validation.image_asset_alignment.review so CI stays offline).

Evidence

Scratch-bundle: invented OCR terms fail and the file is not kept; vision FAIL then PASS retries; validate OCR of matching terms passes.

pixel OCR fail, vision retry, validate pass

pytest tests/: 919 passed, 1 skipped.

How to use it

Keep manim_scene_generation.context.paths pointed at the real docs. Image prompts should name those terms. Opt out of vision (still keep prompt + OCR) with image_generation.align_review: false. Opt out of all alignment with image_generation.align_with_docs: false.

To show artifacts inline, enable in settings.

Open in Web Open in Cursor 

cursoragent and others added 3 commits September 26, 2026 23:11
…work.

Image elements now share the same fail-closed contract as subject-beat
coverage: scene-spec-generate feeds source snippets into the LLM, image
prompts must use documented terms, and image-generate wraps the Images
API call with narration/source plus a validate gate.

Co-authored-by: jmjava <jmjava@gmail.com>
The image stem was treated as an invented subject-beat label and failed
coverage before the new prompt-alignment gate could run.

Co-authored-by: jmjava <jmjava@gmail.com>
Prompt grounding only checked the caption. image-generate now OCRs the
PNG for invented labels and vision-reviews it (OpenAI / Grok / Claude),
retries once with the critique, and deletes a failing asset. validate
adds image_asset_alignment (OCR on by default; vision opt-in).

Co-authored-by: jmjava <jmjava@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants