Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
345 changes: 345 additions & 0 deletions .agents/logs/tone_model_eval/2026-09-10-report.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,345 @@
{
"scope_boundary": "## What this eval can and cannot claim\nThis eval supports a **relative** claim: which candidate model most closely follows this repo's AGENTS.md tone/concision guidance when every model edits the identical \"before\" text for a fixture. The per-model, per-fixture and aggregate scores below are comparable to each other.\n\nThis eval does **not** support an **absolute** \"how much better than the original\" claim. Each fixture's \"before\" text has an unknown or uncontrolled authorship model, so no candidate's edit can be scored as an absolute improvement over it \u2014 only relative to the other candidates scored on the same fixture.",
"judge_model_id": "gemini-3.1-pro",
"fable_model_id": "claude-5-1-fable-high",
"default_model_id": "claude-4-5-sonnet",
"rows": [
{
"fixture_id": "cli-agent-conversations-resume-menu-label",
"model_id": "claude-5-1-fable-high",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 1395,
"after_words": 1324,
"delta": -71,
"delta_pct": -0.05089605734767025
},
"judge": {
"concision": 5,
"avoids_over_explaining": 5,
"technical_fidelity": 4
},
"judge_composite": 4.666666666666667,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "byollm-gemini-enterprise-google-cloud-setup",
"model_id": "claude-5-1-fable-high",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 2630,
"after_words": 2555,
"delta": -75,
"delta_pct": -0.028517110266159697
},
"judge": {
"concision": 4,
"avoids_over_explaining": 4,
"technical_fidelity": 2
},
"judge_composite": 3.3333333333333335,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "quickstart-synthetic-verbose-seed",
"model_id": "claude-5-1-fable-high",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 235,
"after_words": 107,
"delta": -128,
"delta_pct": -0.5446808510638298
},
"judge": {
"concision": 5,
"avoids_over_explaining": 5,
"technical_fidelity": 5
},
"judge_composite": 5.0,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "cli-agent-conversations-resume-menu-label",
"model_id": "claude-4-5-sonnet",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 1395,
"after_words": 1264,
"delta": -131,
"delta_pct": -0.0939068100358423
},
"judge": {
"concision": 4,
"avoids_over_explaining": 4,
"technical_fidelity": 2
},
"judge_composite": 3.3333333333333335,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "byollm-gemini-enterprise-google-cloud-setup",
"model_id": "claude-4-5-sonnet",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 2630,
"after_words": 2595,
"delta": -35,
"delta_pct": -0.013307984790874524
},
"judge": {
"concision": 4,
"avoids_over_explaining": 4,
"technical_fidelity": 4
},
"judge_composite": 4.0,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "quickstart-synthetic-verbose-seed",
"model_id": "claude-4-5-sonnet",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 235,
"after_words": 98,
"delta": -137,
"delta_pct": -0.5829787234042553
},
"judge": {
"concision": 5,
"avoids_over_explaining": 5,
"technical_fidelity": 5
},
"judge_composite": 5.0,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "cli-agent-conversations-resume-menu-label",
"model_id": "claude-4-5-haiku",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 1395,
"after_words": 1251,
"delta": -144,
"delta_pct": -0.1032258064516129
},
"judge": {
"concision": 4,
"avoids_over_explaining": 4,
"technical_fidelity": 3
},
"judge_composite": 3.6666666666666665,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "byollm-gemini-enterprise-google-cloud-setup",
"model_id": "claude-4-5-haiku",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 2630,
"after_words": 2579,
"delta": -51,
"delta_pct": -0.019391634980988594
},
"judge": {
"concision": 4,
"avoids_over_explaining": 4,
"technical_fidelity": 1
},
"judge_composite": 3.0,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "quickstart-synthetic-verbose-seed",
"model_id": "claude-4-5-haiku",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 235,
"after_words": 88,
"delta": -147,
"delta_pct": -0.625531914893617
},
"judge": {
"concision": 5,
"avoids_over_explaining": 5,
"technical_fidelity": 5
},
"judge_composite": 5.0,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "cli-agent-conversations-resume-menu-label",
"model_id": "gpt-5-mini",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 1395,
"after_words": 1188,
"delta": -207,
"delta_pct": -0.14838709677419354
},
"judge": {
"concision": 4,
"avoids_over_explaining": 4,
"technical_fidelity": 1
},
"judge_composite": 3.0,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "byollm-gemini-enterprise-google-cloud-setup",
"model_id": "gpt-5-mini",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 2630,
"after_words": 2048,
"delta": -582,
"delta_pct": -0.22129277566539923
},
"judge": {
"concision": 4,
"avoids_over_explaining": 4,
"technical_fidelity": 3
},
"judge_composite": 3.6666666666666665,
"judge_model_id": "gemini-3.1-pro"
},
{
"fixture_id": "quickstart-synthetic-verbose-seed",
"model_id": "gpt-5-mini",
"mechanical": {
"tone_buzzword": 0,
"tone_meta_opener": 0,
"combined": 0
},
"word_count": {
"before_words": 235,
"after_words": 117,
"delta": -118,
"delta_pct": -0.502127659574468
},
"judge": {
"concision": 5,
"avoids_over_explaining": 5,
"technical_fidelity": 1
},
"judge_composite": 3.6666666666666665,
"judge_model_id": "gemini-3.1-pro"
}
],
"aggregates": {
"claude-5-1-fable-high": {
"fixture_count": 3,
"dimension_scores": {
"concision": 4.666666666666667,
"avoids_over_explaining": 4.666666666666667,
"technical_fidelity": 3.6666666666666665
},
"composite_judge_score": 4.333333333333333,
"combined_mechanical_violations": 0
},
"claude-4-5-sonnet": {
"fixture_count": 3,
"dimension_scores": {
"concision": 4.333333333333333,
"avoids_over_explaining": 4.333333333333333,
"technical_fidelity": 3.6666666666666665
},
"composite_judge_score": 4.111111111111112,
"combined_mechanical_violations": 0
},
"claude-4-5-haiku": {
"fixture_count": 3,
"dimension_scores": {
"concision": 4.333333333333333,
"avoids_over_explaining": 4.333333333333333,
"technical_fidelity": 3.0
},
"composite_judge_score": 3.888888888888889,
"combined_mechanical_violations": 0
},
"gpt-5-mini": {
"fixture_count": 3,
"dimension_scores": {
"concision": 4.333333333333333,
"avoids_over_explaining": 4.333333333333333,
"technical_fidelity": 1.6666666666666667
},
"composite_judge_score": 3.444444444444444,
"combined_mechanical_violations": 0
}
},
"recommendation": {
"adopt_fable_guidance": {
"passed": false,
"concision_margin": 0.3333333333333339,
"mechanical_violation_reduction_pct": 0.0
},
"cheaper_model_candidates": {
"claude-4-5-sonnet": {
"passed": false,
"judge_gap": 0.22222222222222143,
"mechanical_violations": 0,
"fable_mechanical_violations": 0,
"technical_fidelity": 3.6666666666666665
},
"claude-4-5-haiku": {
"passed": false,
"judge_gap": 0.4444444444444442,
"mechanical_violations": 0,
"fable_mechanical_violations": 0,
"technical_fidelity": 3.0
},
"gpt-5-mini": {
"passed": false,
"judge_gap": 0.8888888888888888,
"mechanical_violations": 0,
"fable_mechanical_violations": 0,
"technical_fidelity": 1.6666666666666667
}
}
},
"calibration_warning": null
}
51 changes: 51 additions & 0 deletions .agents/logs/tone_model_eval/2026-09-10-report.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# Tone/concision model eval report

## What this eval can and cannot claim
This eval supports a **relative** claim: which candidate model most closely follows this repo's AGENTS.md tone/concision guidance when every model edits the identical "before" text for a fixture. The per-model, per-fixture and aggregate scores below are comparable to each other.

This eval does **not** support an **absolute** "how much better than the original" claim. Each fixture's "before" text has an unknown or uncontrolled authorship model, so no candidate's edit can be scored as an absolute improvement over it — only relative to the other candidates scored on the same fixture.

**Judge model:** gemini-3.1-pro (check for a same-family match against any candidate before trusting its score)

## Per-model scores
### claude-5-1-fable-high
- Fixtures scored: 3
- Composite judge score: 4.33/5
- concision: 4.67/5
- avoids_over_explaining: 4.67/5
- technical_fidelity: 3.67/5
- Combined mechanical violations (tone-buzzword + tone-meta-opener): 0

### claude-4-5-sonnet
- Fixtures scored: 3
- Composite judge score: 4.11/5
- concision: 4.33/5
- avoids_over_explaining: 4.33/5
- technical_fidelity: 3.67/5
- Combined mechanical violations (tone-buzzword + tone-meta-opener): 0

### claude-4-5-haiku
- Fixtures scored: 3
- Composite judge score: 3.89/5
- concision: 4.33/5
- avoids_over_explaining: 4.33/5
- technical_fidelity: 3.00/5
- Combined mechanical violations (tone-buzzword + tone-meta-opener): 0

### gpt-5-mini
- Fixtures scored: 3
- Composite judge score: 3.44/5
- concision: 4.33/5
- avoids_over_explaining: 4.33/5
- technical_fidelity: 1.67/5
- Combined mechanical violations (tone-buzzword + tone-meta-opener): 0

## Recommendation
### Adopt Fable-5.1-derived guidance
**No meaningful difference found** — neither the ≥1.0-point concision-dimension margin (0.33 observed) nor the ≥30% mechanical-violation reduction (0% observed) was met.

### Recommend a cheaper model for production copy passes
**No candidate met the threshold** — every candidate either fell outside the 0.5-point judge tolerance, exceeded Fable 5.1's mechanical violation count, or scored below 4/5 on technical fidelity.
- claude-4-5-sonnet: Fail — judge gap 0.22 (tolerance ≤0.5), mechanical violations 0 vs Fable 5.1's 0, technical fidelity 3.67/5 (min 4.0)
- claude-4-5-haiku: Fail — judge gap 0.44 (tolerance ≤0.5), mechanical violations 0 vs Fable 5.1's 0, technical fidelity 3.00/5 (min 4.0)
- gpt-5-mini: Fail — judge gap 0.89 (tolerance ≤0.5), mechanical violations 0 vs Fable 5.1's 0, technical fidelity 1.67/5 (min 4.0)
Loading
Loading