Conversation
The trajectory reached only the offline evaluations judge. The two online
paths built their own message_history and neither included it, so the
same judge grading the same response saw a different conversation
depending on which path reached it -- and a trajectory rubric silently
degraded to grading prose when run online.
They had already drifted before the trajectory made it visible:
offline row input + trajectory + output + format block
inline user input + + output + format block
deferred + output + format block
The deferred path carried no input at all, so a background judge graded
a response with no request beside it.
judge_scoring.build_message_history is now the only place a history is
built, in the module that already owns the {score, reasoning} contract
for the same reason. All three paths call it, and a test asserts the
inline and deferred paths produce byte-identical output for one row.
Capture online happens in execute_and_track and execute_and_stream,
which return the rendered trajectory alongside response and track_data.
client.py and the two per-node graph.py judge runs thread it through.
JudgeTask gains user_input and trajectory -- plain strings, since every
field on it has to survive pickling to a worker thread.
Recording is composed *inside* wrap_tool_handlers, on the original tool
map, so the recorder still sees a NativeTool as a NativeTool and skips
it. Wrapping the tracked map instead would have recorded the sync
callable stub that wrapper substitutes for a native tool, showing a
judge a call with an empty result while the provider's real result
stayed invisible. Both paths now treat natives identically, and
$ld:ai:tool_call still fires underneath -- both asserted.
trajectory.py moves from evaluations/ to the package root: it is no
longer evaluations-specific.
A graph-level judge deliberately gets no trajectory. It grades a final
answer produced across several nodes, and splicing their trajectories
would describe a conversation that never happened.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
donei003
marked this pull request as ready for review
September 18, 2026 00:04
donei003
merged commit Sep 18, 2026
bbc59c7
into
feature/ld-judges-tool-trajectory
7 of 8 checks passed
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
Reviewed by Cursor Bugbot for commit 398dbad. Configure here.
| recorder = TrajectoryRecorder() | ||
| tracked_tool_handlers = wrap_tool_handlers( | ||
| recorder.wrap(tool_handlers or {}), ld_ctx, track_data | ||
| ) |
There was a problem hiding this comment.
Handoff tools pollute judge trajectories
Medium Severity
execute_and_track now records every callable into the trajectory, including synthetic __handoff_* tools that route() injects. Per-node judges therefore see those names under tools available and as real calls, and a node with no customer tools still gets a trajectory block. wrap_tool_handlers already omits the same tools from $ld:ai:tool_call as non-work.
Additional Locations (2)
Reviewed by Cursor Bugbot for commit 398dbad. Configure here.
donei003
added a commit
that referenced
this pull request
Sep 28, 2026
Now carries two commits: the trajectory capture, and #98 (merged in here) which made every judge path share one `message_history`. ## Problem A judge can only grade what it is shown. Handler packages record tool traffic onto OpenTelemetry spans and return only `{output, usage}`, so by the time a judge ran, the calls made on the way to that output were gone. "Did the agent call the right tool, in the right order, with the right arguments?" was an unaskable question of an SDK that had just run the agent that answered it. Underneath that, the three judge paths had already drifted. Each built `message_history` with its own inline join: | Path | `message_history` was | | --- | --- | | Offline evals | row input + output + format block | | Online inline (`run_judges`) | user input + output + format block | | Online deferred (`run_judge`) | output + format block — **no input at all** | A deferred judge was grading a response with no request beside it. That is a pre-existing bug, fixed here as a consequence of unifying. ## Approach **Capture by wrapping the tool map**, which both paths already own — so every handler package is covered without changing any of them, and a custom caller-supplied handler is covered too. A handler still resolves a tool by the key the model named and calls it. - Offline: once per row, in `evaluations/runner.py`. - Online: once per invocation, in `tracking.execute_and_track` / `execute_and_stream`, which now return the rendered trajectory alongside `response` and `track_data`. **`judge_scoring.build_message_history` is the only place a history is built** — in the module that already owns the `{score, reasoning}` contract, for exactly the same reason. It orders the conversation the way it happened: input → trajectory → output → format block, skipping empty parts. All three paths call it. What a judge now sees: ```text Where is order A1? Tools available: lookup_order, issue_refund Tool calls made while producing the response, in order: 1. lookup_order arguments: {"id":"A1"} result: order A1 shipped 2026-08-02 2. issue_refund arguments: {"id":"A1","amount":19.99} error: refund window closed That order shipped on Aug 2 and is outside the refund window. Your response MUST be in valid JSON format with the following structure: ... ``` Verified live against a real LaunchDarkly judge: the judge's reasoning cited the call, its arguments, **and** its result — all three only visible via the trajectory. ## Properties pinned by tests - **The recorder observes; it never intervenes.** A wrapped tool returns exactly what the original returned and raises exactly what the original raised. Calls past the 50-call cap **still execute** and are only counted — truncation drops the record, never the work, because a harness that changed the agent's behaviour would no longer be evaluating the agent. - **A recorder belongs to one invocation.** Rows generate concurrently against one shared tool map, so a shared recorder would splice one row's calls into another's and hand the judge a conversation that never happened. Tested with two rows whose calls interleave. - **Order is call-start order, not completion order.** A judge asked whether the agent searched before it refunded is reading a sequence. - **Inline and deferred produce byte-identical history** for the same row. This is the property that makes a rubric portable between a production sample and a dataset replay. - **A tool result stays literal in the judge prompt.** A tool result is a *new* injection surface alongside generated output, closed by the existing rule: the judge config is passed unrendered and the handler makes exactly one template pass. Tested with a tool returning the literal text `{{expected_output}}`. - **Back-compatible.** A run with no observable tools adds no block, so a judge authored before this reads exactly the `message_history` it read before. - **Nothing is added to any event payload.** The trajectory reaches LaunchDarkly only inside the prompt a judge was shown, never as a wire field the backend has not specified. ## Design notes - **No standalone trajectory variable.** An earlier revision exposed `{{tool_trajectory}}` too; it overlapped `message_history` and bought nothing, while inviting a rubric to interpolate both and pay for the trajectory twice. Confirmed live — a real judge config's own scaffolding already interpolates `{{message_history}}`. A test pins its absence. Since `message_history` is what judges cloned from the AI Library's default templates read, **an existing judge becomes a trajectory judge by editing its rubric text alone.** - **Native provider tools are skipped in both paths, and left out of "Tools available".** Online, `wrap_tool_handlers` substitutes a callable tracking stub, so a native call *is* locally observable — but the stub returns nothing, so recording it would show a judge a call with an empty result while the provider's real result stayed invisible. Recording is therefore composed *inside* that wrapper, on the original map. Tests assert both the exclusion and that `$ld:ai:tool_call` still fires underneath. - **`JudgeTask` gains `user_input` and `trajectory`** as plain strings — every field on it has to survive pickling to a worker thread; a test pins that. - **Graph-level judges get no trajectory,** deliberately: `graph_judge` grades an answer produced across several nodes, and splicing their trajectories would describe a conversation that never happened. Per-node judges get their own node's. - **Bounds:** 50 recorded calls per invocation, 2000 characters per rendered argument bag or result. A trajectory goes into a judge prompt, so an agent looping over a large result set would otherwise spend the judge's context window — and budget — on a tail no judge reads. - `trajectory.py` sits at the package root, not under `evaluations/`, since it is no longer evaluations-specific. ## Deliberately out of scope **Scorers cannot see the trajectory.** `Scorer.fn(row, output)` is the contract, and the trajectory is not dataset-owned so it does not belong on `DatasetRow`. A deterministic check like "called `lookup_order` exactly once" is a natural follow-up but needs a contract change, not a quiet signature widening. ## Validation - `uv run pytest -q` — **1282 passed**, 11 skipped - `uv run mypy packages/client/src/launchdarkly_ai_server` — clean; `ruff check` / `format --check` — clean - Exercised end-to-end in `ai-sdk-evaluations-example` (launchdarkly-labs/ai-sdk-evaluations-example#5), including a live run against a real judge Language-agnostic spec: launchdarkly/ai-sdks-monorepo#13 — being restructured to describe the shared flow rather than the offline phase alone, now that this shape is settled. Submodule pointer: launchdarkly/ai-sdks-monorepo#14. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Overview** > Judges can now grade **how** an agent answered, not only the final text. The SDK **records tool calls** by wrapping the caller’s tool map (online in `execute_and_track` / `execute_and_stream`, offline per dataset row), renders them into a bounded transcript, and splices that into **`{{message_history}}`** between user input and model output—no new judge template variable and no new event payload fields. > > **`judge_scoring.build_message_history`** becomes the single builder for inline `run_judges`, deferred `run_judge` (`JudgeTask` now carries `user_input` and a picklable `trajectory` string), and offline eval scoring—fixing the deferred path that previously omitted the user request. Graph per-node judges receive the streamed/captured trajectory; graph-level judges still do not. > > Recording rules exclude synthetic `__handoff_*` tools and provider **`NativeTool`** calls from the transcript, cap volume (50 calls, 2k chars per value), and keep tool behavior and `$ld:ai:tool_call` tracking unchanged. Docs/READMEs describe trajectory rubrics; OTel graph spans rename from `ld.ai.graph` to **`launchdarkly.graph`** (attribute `launchdarkly.graph.key`). Package versions bump to **0.2.3** / **0.1.7**. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 245a0a5. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Stacked on #89 — base is
feature/ld-judges-tool-trajectory, so the diff is just this change. Retarget tomainwhen #89 merges.Problem
#89 gave the trajectory to the offline judge only. The two online paths built their own
message_historyand neither included it — so a trajectory rubric silently degraded to grading prose when run online, and a judge grading the same response saw a different conversation depending on which path reached it.They had already drifted before the trajectory made it visible:
message_historywasrun_judges)run_judge)A deferred judge was grading a response with no request beside it. That is a pre-existing bug this PR also fixes.
Change
judge_scoring.build_message_historyis now the only place a history is built — in the module that already owns the{score, reasoning}contract, for exactly the same reason. All three paths call it, and a test asserts the inline and deferred paths produce byte-identical output for one row.Online capture happens in
execute_and_track/execute_and_stream, which now return the rendered trajectory alongsideresponseandtrack_data.client.pyand the two per-nodegraph.pyjudge runs thread it through.JudgeTaskgainsuser_inputandtrajectory— plain strings, since every field on it has to survive pickling to a worker thread; a test pins that.trajectory.pymoves fromevaluations/to the package root, since it is no longer evaluations-specific.The
NativeTooldecision you asked aboutRecording is composed inside
wrap_tool_handlers, on the original tool map, so the recorder still sees aNativeToolas aNativeTooland skips it — identically to offline.Wrapping the tracked map instead was the tempting option, because native calls are locally observable online:
wrap_tool_handlerssubstitutes a callable tracking stub. But that stub returns nothing, so recording it would show a judge a tool call with an empty result while the provider's real result stayed invisible — worse than not showing it. Tests assert the native tool is absent from the online trajectory and that$ld:ai:tool_callstill fires underneath the recorder.Tell me if you'd rather natives appear online with an explicit "result not observable" marker; it's a small change now that one function owns the rendering.
Graph-level judges get no trajectory, deliberately
graph_judgegrades a final answer produced across several nodes. Splicing their trajectories together would describe a conversation that never happened, so it gets"". Per-node judges inside a graph do get their own node's.Validation
uv run pytest -q— 1282 passed, 11 skippeduv run mypy packages/client/src/launchdarkly_ai_server— clean;ruff check/format --check— cleantest_judge_message_history.py: the builder's ordering and skipping, the trajectory reaching both online paths, inline-vs-deferred agreement,JudgeTaskpicklability, online capture through the realexecute_and_track, native-tool exclusion, and$ld:ai:tool_callsurviving the compositionSpec follow-up for
ai-sdks-monorepo§3.13/§3.14 to come once this shape is agreed.🤖 Generated with Claude Code
Note
Overview
Unifies what judges see across inline online, deferred (
JudgeTask), and offline evaluation paths by introducingjudge_scoring.build_message_historyas the only place{{message_history}}is assembled: user input, rendered tool trajectory, model output, then JSON formatting instructions (empty parts skipped).Online invocations now capture tool trajectories in
execute_and_track/execute_and_stream(per-invocationTrajectoryRecorder, composed insidewrap_tool_handlerson the original tool map soNativeToolstays unrecorded). The renderedtrajectorystring is returned with the generation result and threaded throughclient.py,graph.py,run_judges, andbuild_judge_tasks.Deferred judges gain
JudgeTask.user_inputandJudgeTask.trajectory(picklable strings), fixing the prior bug where deferred grading omitted the request entirely. Offline scoring inevaluations/runner.pyuses the same builder instead of inline joins.trajectorymoves fromevaluations/to the package root for shared online/offline use. Docs andtest_judge_message_history.pyassert inline vs deferred byte-identical history and online capture behavior.Reviewed by Cursor Bugbot for commit 398dbad. Bugbot is set up for automated code reviews on this repo. Configure here.