Translation bench - #2831
Conversation
Empty expectedActions is scored as zero actions. Stop minting contrastive adjacent commands and refuse-then-alternate forms as empty-gold negatives; format_checker now hard-rejects them as BAD_NEGATIVE. Allowed negatives: pure refusal, non-action status/howto questions, and missing-info clarifications. Align synthesizer and quality-verifier prompts with the zero-action contract; add unit coverage including adversarial holes.
- Drop ACTION_VP / refusal regex classifier (unmaintainable verb lists). - Require semantic_checker negativeAssessments (kind + fairEmptyGold). - Code hard-fails unfair assessments as BAD_NEGATIVE; format stays structural. - Tests cover assessment parse/enforce and mock-LLM semantic gate.
- Single zod schema; derive OpenAI JSON schema via z.toJSONSchema - Match assessments by order (equal count); drop path-index maps - Simplify force-reject merge; strip explanatory comments
- Drop dead exports; parse via shared parseWithZod - Derive OpenAI schema from zod without void/\$schema clutter - Add test that assessments pair to negatives by order
- Join assessments to negatives by exact genCase path (bijective set); reject unknown/duplicate/missing paths as BAD_NEGATIVE instead of silent index pairing. - Parse reviewer decision before assessments so structured reject issues survive missing/invalid negativeAssessments. - Share TRANSLATION_BENCH_NEGATIVE_FAIRNESS_RULE in synthesizer context; tighten howto/soft-solicit unfair rules and untrusted payload framing. - Extend reviewerDecision mocks with path-keyed fair assessments; replace order-only path tests with multi-negative path-join coverage.
- Fixed short PATH_MSG/FIX only; use LLM reason as issue message - Drop rewrite-hint and per-case string assembly - Keep path 1:1 cover + fairEmptyGold/kind gate
- Honor concurrency with a worker pool and serialized checkpoint commits
- Keep parameters:{} after stripEmpty for required empty-object schemas
- Inject constant string-union fields (e.g. settings id) during gold validate
- Isolate per-slot generation failures so other workers keep committing
- Fair empty gold is pure_refusal/leave-alone only - Reject definition/status/Q&A and missing_info as empty gold - Synthesizer + quality-verifier prompts match any-tool abstain test - Extend unit coverage for definition Q and soft-solicit rejects
Join assessments to genCases via a path→case map built from the same keys the checker emits, then read dimensions.negativeKind directly.
…params - Allow incomplete case sets when requireCompleteCoverage is false - Prompt: nested objects must carry required schema fields (e.g. timeRange)
…ial gen - Partial-gen branch previously emitted planned caseCount/coverage, so validateTranslationBenchBenchmark always threw and the branch was unreachable dead code - Recompute scheduledActionCount, complete, and caseCount from the cases actually emitted; happy-path (complete) output is unchanged - Makes requireCompleteCoverage=false produce a valid draft directly
Derive FAIR_KINDS from TRANSLATION_BENCH_FAIR_EMPTY_GOLD_KINDS so the allowlist cannot drift from the exported constant.
There was a problem hiding this comment.
Pull request overview
Adds a fairness contract ensuring empty-gold translation-bench negatives represent only pure refusals.
Changes:
- Adds path-keyed LLM fairness assessments and enforcement.
- Updates generation/verifier prompts and gold-action validation.
- Supports concurrent generation and accurate partial-result metadata.
Reviewed changes
Copilot reviewed 13 out of 13 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
README.AUTOGEN.md |
Refreshes generated documentation metadata. |
actionValidation.ts |
Normalizes gold actions for schema validation. |
benchmark.ts |
Uses normalized gold-action validation. |
dataQualityVerifier.ts |
Enforces semantic negative-fairness assessments. |
datasetGenerator.ts |
Adds concurrency and partial-generation handling. |
generationCandidate.ts |
Preserves required parameter objects during cleanup. |
index.ts |
Exports the fairness module. |
negativeFairness.ts |
Implements fairness parsing and enforcement. |
quality-verifier.prompt.yaml |
Defines full-catalog fairness review requirements. |
synthesizer.prompt.yaml |
Restricts empty-gold negatives to pure refusals. |
utteranceDisambiguation.ts |
Documents separate negative handling. |
translationBench.datasetGenerator.spec.ts |
Updates generation fixtures for fairness assessments. |
translationBench.negativeFairness.spec.ts |
Tests fairness validation and semantic enforcement. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…s, ambiguous-route guard Row-by-row 3-model review of the 1k eval showed most all-models "failures" were dataset/scoring fairness issues, not real misses. Address the generator-side ones: - Wire parameterScore specs into every generated case at finalize (`parameterScoreSpecsForExpectedActions`). Free-text echo params such as `originalRequest` and setting free-text now score `nonempty` instead of exact, matching the grader's own classification. Canonical payload hash excludes parameterScore, so dataset identity is preserved. - Add `HARDCODED_NON_EVAL_ACTION_IDS` (`chat.generateResponse`, `utility.claudeTask`) as a single source of truth, unioned into the packaged exclusion set so they are never targeted. - Drop cross-schema duplicate action names from targeting: when the same bare action name is owned by more than one schema (e.g. `deleteWebFlow`), the single gold route is ambiguous, so exclude every sibling. Tests: parameterScore wiring + canonical-hash stability, non-eval exclusion set, and ambiguous cross-schema guard. Full benchmarks suite green (138).
…eckpoint, param-spec cleanup Code-review + Copilot follow-ups on TB generation fairness: - negativeFairness: require dimensions.negativeKind to exactly equal the accepted empty-gold assessment kind. Previously a fair pure_refusal assessment was accepted even when negativeKind was missing/numeric/boolean, letting unlabeled empty-gold rows through. Now the label must match. - datasetGenerator: persist the checkpoint row BEFORE mutating casesBySlot / usedUtterances in commitAccepted, so an I/O failure can no longer leave an uncheckpointed case that the partial-coverage path would return. - Unify the duplicated param-field-mode union: grader now imports TranslationBenchParamFieldMode / TranslationBenchParameterScoreSpec from benchmark instead of re-declaring RunnerParamFieldMode; collapse the spec derivation to Object.fromEntries. - benchmark: enforce parameterScore aligns 1:1 with expectedActions in the probe payload schema (validateProbePayload). - Extract ambiguousCrossSchemaActionIds helper; drop the size===0 ternary. Tests: new integration coverage for generateTranslationBenchBenchmark (concurrent full run + partial run past a failed slot, asserting checkpoint contents and coverage); negativeKind-missing rejection; fixtures now label negatives pure_refusal. Full suite green (141).
- Seed 24 empirically-mined cross-schema collision pairs into KNOWN_CONFUSABLE_PAIRS (browser tab routes, code/utility file ops, visualStudio/code-debug, desktop/display settings, player queues, etc.) - Add generic cross-schema near-duplicate detector in findTranslationBenchConfusableSiblings: flags equivalent actions in different schemas when BOTH action-name token overlap >=0.5 AND description overlap >=0.34, so shared generic verbs alone do not over-flag unrelated actions - Add significantTokensFromText helper for description-token overlap - Extend unit tests with cross-schema newTextFile<->writeFile case plus a readFile negative control Root cause: synthesizer scheduled every catalog action as a gold target but only compared same-schema siblings for disambiguation, so cross-schema near-synonyms produced no disambiguation constraint and yielded ambiguous-gold cases the models unanimously routed elsewhere.
- Move catalog/grader generation into translationBench/policy - Add action-eligibility hard bans + LLM quality picker (model required) - Ship eligible-gold-actions.generated.json; fail-closed load/integrity - Schedule lattice + eval pin allowlist hash; copyAssets requires assets - Tests for picker, policy, nested llmAsAJudge, schedule allowlist-on
…ypeAgentQuestion from eligible actions
Declare each model once with tpmLimit. Synthesizer and classifier scenarios pick from that list; the limiter enforces the same quota regardless of scenario.
Empty-gold negatives are judged by structured LLM flags, not utterance regex. Generate concurrency comes from TPM; probes fan out on isolated contexts and the limiter is the throttle.
- eval.models is { model, concurrency? }[]
- run one eval model at a time; fan out cases inside that model
- keep only design principles in AGENTS.md
…-fairness # Conflicts: # ts/packages/benchmarks/package.json # ts/packages/benchmarks/scripts/copyAssets.mjs # ts/packages/benchmarks/src/translationBench/config.schema.json # ts/packages/benchmarks/src/translationBench/config/run-config.example.json # ts/packages/benchmarks/src/translationBench/synthesizer/actionValidation.ts # ts/packages/benchmarks/src/translationBench/synthesizer/index.ts
Eval chosenActions now drop anything off the packaged gold allowlist so synthesis targets and scored fires share one list.
- Store approved 100-row dataset, eval results, and HTML report under eval-runs/ as Git LFS objects - Commit the actual run-config used for generate/eval, not only the example file
100-row generate + eval (this push)Artifacts (Git LFS):
Models: As stored in
|
| Metric | Overall | 4.1 | 4.1-mini | 5.4-nano |
|---|---|---|---|---|
| Pass | 76.3% (453/594) | 76.3% (151/198) | 76.3% (151/198) | 76.3% (151/198) |
| Exact pass | 71.0% (422/594) | 71.2% | 71.2% | 70.7% |
| Tool (positives) | 97.3% (289/297) | 97.0% | 98.0% | 97.0% |
| Param / routed | 92.4% (267/289) | 91.7% | 92.8% | 92.7% |
| FNR | 2.7% (8/297) | 3.0% | 2.0% | 3.0% |
| FPR | 37.4% (111/297) | 36.4% | 38.4% | 37.4% |
Positive pass 90.0% (267/297). Negative pass 62.6% (186/297). Schema-valid 100%. Errors 0.
Scorer change on this branch
Eval now scores only the packaged gold allowlist (same list synthesis can target). chat.generateResponse and other off-allowlist fires stay in rawChosenActions but are not scored.
If this run is rescored under that rule:
- 27 chat-only empty-gold fires drop
- 30 other off-allowlist-only fires drop
- FPR 111/297 → 54/297 = 18.2%
- those 57 negative fails become passes → overall pass 510/594 = 85.9%
Typical remaining FPs are real gold tools on a hard abstain, e.g. saveCurrentFile on “Don't close the code editor…”, removeBreakpoint/cancelReminder as the opposite tool, or firing the refused action itself (AdjustColorTemperature).
Semantic field meaning goes through the Zod LLM classifier/verifier. Deterministic rules stay type-only.
Reuse gold validation and runner alignment; keep simple-mode prompts and generator untouched so concurrent single-action work does not collide. tb-generate --mode multi uses new packs, coverage quotas, and 2-3 call gold checks.
Name-table cases now fail closed or go through the mock LLM+verifier path.
Ambiguity is LLM-judged; format_checker no longer substring- matches action cues. Unions walk arms. Dead overlay helpers gone.
Verifier was fail-closing startUrl for lacking a url create policy that does not exist.
Fail-closed rebuilds died on record+exact for free-text fragment objects. Verifier now accepts nonempty.
Pin nine catalog URL fields to nonempty so fail-closed classify does not emit free_text+exact. Packaged rulesFingerprint now matches live policy.
Parameter overrides changed the policy hash; drop visualStudio.addBreakpoint now marked llmAsAJudge.
Force rebuild: 558 actions, 318 structural, 600 LLM-verified.
--require-complete-coverage now fills caseCount; full catalog coverage stays on coverage.complete for the 1k run.
Round-001 accepted 0/100: the verifier scored anchorFidelity as topic match. Isolated prompts now require scheduled gold, not the anchor's calls.
Slot cycling reused a few favorite pairs. Least-used unique combos keep 100-row gold on many distinct actions.
100-slot catalog must use >=150 actions and keep 3-call cross slots on two schemas.
OpenAI tool JSON schemas list every property in required[]. Reviewers then demanded search_filters, and the generator dumped optional createIfNotExists/fallbackToActiveFile:false into gold. - Rewrite gold-facing required[] to TypeAgent optionality - Strip optional false booleans after empty-placeholder hygiene - Bind negativeAssessments 1:1 to $.genCases[N].utterance - Treat probe unknown-throws as abstention and retry translate blips
Round-003 died 0/100: TypeAgent probe errors were treated as bad gold. Only unscheduled extras reject.
Least-used picks joined Bluetooth with editor split. Cross gold now prefers one product surface.
Probe errors and omitted calls stay valid gold.
Round-004 died 99/100: slot 74 named exactMatch off but omitted it from setSearchParameters.
Probes posted model=None because createChatModel("azure/gpt-5.5")
ignored the name on the OpenAI-compatible path. Slash ids are now
the request model. Stop stripping optional false booleans so
enable:false survives when the utterance asks to turn something off.
Round-005 review failed: history replayed SystemThemeMode/findMusic; any-order on start/stop.
Summary
Empty-gold TB negatives (
expectedActions: []) are scored as zero actions across the full loaded tool catalog. Label-fair Q&A/definition negatives were not zero-action-scorer-fair: models correctly firechat/help/history, tanking negative pass.This PR ships a pipeline fairness contract so empty gold is only accepted for hard abstains:
pure_refusalonly (leave-alone / don't / never / cancel of the target; no questions, no alternate task).negativeFairnessmodule: path-keyed 1:1negativeAssessments(no regex path parse); fail-closed on missing/duplicate paths;fairEmptyGoldonly when kind is in the allowlist (TRANSLATION_BENCH_FAIR_EMPTY_GOLD_KINDS).caseCount/ coverage metadata match the cases actually produced (so validation can succeed).Measured 3-model eval (honest)
Regenerated draft under the new policy (998/1000 rows; 2 slots exhausted data-quality, unrelated to fairness), then evaluated
azure/gpt-5.6-{sol,terra,luna}.pure_refusalFinding: fairness regen fixed the dataset (kind mix + unfair rate), but did not produce a large negative-pass lift. Checkpoint inspection shows models often emit the refused action itself on clear hard-abstains (negation-following failure). Translator/prompt work is out of scope here.
Test plan
pnpm run buildints/packages/benchmarkspnpm run jest-esm --testPathPattern='(negativeFairness|datasetGenerator)\\.spec\\.js'(43/43)Infra: extract cross-process TPM limiter + run config to
coreFollow-on refactor moving the run harness's rate limiter and config loader into
@typeagent/benchmarksas tested, reusable modules.Interface (one owner per concern):
createRateLimiter(tpmLimits, { dbPath, estTokensPerCall?, maxWaitMs?, onWait? })→{ disabledFor, run, close }.dbPathis required so callers own persistence; no hidden default DB. Optional knobs are orthogonal.runConfig.tsis pure: reads a JSON file (schema-validated), merges batch over base, derives per-model concurrency. No env, no process globals — runtime overrides are commander flags prop-drilled by the local runners (TB_*env removed).OPENAI_*/AZURE_*) is intentionally kept as theaiclientcontract; documented inAGENTS.md.Correctness:
BEGIN IMMEDIATE, WAL): concurrent awaitedrun()calls reserve then settle to actual token usage, staying within quota across processes.STALE_MS(purge guarded onpending = 0); settle falls back to insert if the claim row was purged.Tests:
translationBench.rateLimiter.spec+translationBench.runConfig.spec(17 cases) — passthrough for unquota'd models, budget admit/throttle, cross-instance shared budget, estimate fallbacks, batch merge, concurrency derivation, malformed-json handling.Test plan (infra)
pnpm run build+pnpm run prettierclean ints/packages/benchmarkspnpm run jest-esm --testPathPattern='translationBench.(rateLimiter|runConfig).spec.js'(17/17)resolveRunnerConfigsmoke against realconfig.local.json