Related external issue:
openai/codex#40755
@xormania — your reproduction is interesting to CounterProof for a reason slightly different from the Codex bug itself.
You already did the hard work: the reported review can cite a stable-looking Git SHA as command evidence even though that object does not resolve in the repository, while the review's own Reviewed commit: pin does resolve.
CounterProof currently carries exact BASE / HEAD SHAs and a digest in its evidence artifacts, but it mostly treats those as provenance metadata around behavioral evidence.
#40755 suggests provenance may need to be a gate, not decoration.
The evidence state I am trying to name
Something like:
claim:
"commit violates attribution invariant"
evidence citation:
git cat-file -p <sha>
provenance:
cited SHA resolves in reviewed repository? NO
cited SHA == reviewed commit? NO
reviewed commit itself resolves? YES
claim status:
INCONCLUSIVE / provenance-invalid
In other words, CounterProof should not even reach “is the claim witnessed / contradicted?” if the object named as evidence cannot be tied to the reviewed history.
Why I am asking instead of implementing
We just learned from external reviewers that evidence axes are useful only if the vocabulary stays mechanical and auditable. I do not want to add another generic “trust score.”
A possible provenance contract would be narrowly deterministic:
- cited commit/object resolves;
- object belongs to the repository under review;
- object identity matches the declared HEAD / BASE / reviewed commit when that relationship is claimed;
- command evidence names the same object its prose conclusion describes.
Question
From the failure mode you measured in #40755, would a compact provenance-valid / provenance-invalid / unverified gate have made the bad findings immediately actionable for you, or is the more useful artifact simply “show the exact command + stdout + reviewed SHA” and let the reviewer compare them?
Either answer helps decide whether CounterProof should treat provenance as a first-class evidence axis or leave it as raw receipt metadata.
Related external issue:
openai/codex#40755
@xormania — your reproduction is interesting to CounterProof for a reason slightly different from the Codex bug itself.
You already did the hard work: the reported review can cite a stable-looking Git SHA as command evidence even though that object does not resolve in the repository, while the review's own
Reviewed commit:pin does resolve.CounterProof currently carries exact BASE / HEAD SHAs and a digest in its evidence artifacts, but it mostly treats those as provenance metadata around behavioral evidence.
#40755 suggests provenance may need to be a gate, not decoration.
The evidence state I am trying to name
Something like:
In other words, CounterProof should not even reach “is the claim witnessed / contradicted?” if the object named as evidence cannot be tied to the reviewed history.
Why I am asking instead of implementing
We just learned from external reviewers that evidence axes are useful only if the vocabulary stays mechanical and auditable. I do not want to add another generic “trust score.”
A possible provenance contract would be narrowly deterministic:
Question
From the failure mode you measured in #40755, would a compact provenance-valid / provenance-invalid / unverified gate have made the bad findings immediately actionable for you, or is the more useful artifact simply “show the exact command + stdout + reviewed SHA” and let the reviewer compare them?
Either answer helps decide whether CounterProof should treat provenance as a first-class evidence axis or leave it as raw receipt metadata.