Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 3 additions & 4 deletions packages/ai/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,14 +54,13 @@ Never raises. Returns `{"enabled": bool, "config": dict | None, "meta": dict | N

## Evaluations from code

`init_evaluations`, the criterion types, and the evaluations result types are all re-exported:
`init_evaluations`, the criterion types, `Tool`, and the evaluations result types are all re-exported:

```python
from launchdarkly_ai_python import Judge, Scorer, init_evaluations

evals = init_evaluations()
evals = init_evaluations(project_key="my-project")
result = await evals.run(
project_key="my-project",
key="unique-evaluation-key",
dataset="golden-dataset",
handler=my_handler,
Expand All @@ -73,7 +72,7 @@ result = await evals.run(
)
```

`LD_API_TOKEN` is required. Configure `LD_SDK_KEY` — or initialize your own client with `init_client(client=...)` — to emit one `$ld:ai:offline-evals:generation` event per generated row, plus one `$ld:ai:offline-evals:criterion` event per `(row, criterion)` when `criteria` are supplied, through the standard SDK event transport. The SDK reports scores; LaunchDarkly rules on them at ingest. A judge served by a different provider than `generation` needs a handler for it in `judge_handlers`. Each row's tool calls are recorded during generation and rendered into the judge's `{{message_history}}`, between the row input and the generated output, so a rubric can grade the tool trajectory as well as the final answer. Use `LD_API_BASE_URI` for staging or local management API traffic; it is separate from the SDK delivery setting `LD_BASE_URI`. Evaluation-run links use the explicit `ui_base_uri` option or `LD_UI_BASE_URI`, defaulting to `https://app.launchdarkly.com`; set it when the project is not in production, or a run created elsewhere still links to the production app. See the [core evaluations guide](../client/README.md#run-an-evaluation-from-code).
`LD_API_TOKEN` is required. `LD_SDK_KEY` is also required, because the harness needs an initialized client to send your results to LaunchDarkly. Supply your own client with `init_client(client=...)` if you prefer. The SDK reports a score for each row and criterion, and LaunchDarkly decides whether each one passes. A judge served by a different provider than `generation` needs a handler for it in `judge_handlers`. Each row's tool calls go into the judge's `{{message_history}}`, between the row input and the generated output, so a rubric can grade the tool calls and the final answer. `tools` is a list of `Tool`. Construct one to define a tool in code, or call `evals.tools.get()` for a tool that already exists in LaunchDarkly. See the [core evaluations guide](../client/README.md#run-an-evaluation-from-code).

---

Expand Down
59 changes: 33 additions & 26 deletions packages/client/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ No code changes are required — `init_client()` detects the packages at runtime
| `LD_SERVICE_NAME` | No | OTel `service.name` resource attribute (default: `python-sdk`) |
| `LD_ENVIRONMENT` | No | `deployment.environment` resource attribute attached to telemetry |
| `OTEL_EXPORTER_OTLP_ENDPOINT` | No | OTLP endpoint override (default: LaunchDarkly Observability backend) |
| `LD_API_TOKEN` | For evaluations | API access token used by the evaluations management API |
| `LD_API_TOKEN` | For evaluations | API key used by the evaluations management API |
| `LD_SDK_KEY` | For evaluations | SDK key whose event transport carries generation results to LaunchDarkly |
| `LD_API_BASE_URI` | No | Evaluations management API host override; intentionally separate from `LD_BASE_URI` |
| `LD_UI_BASE_URI` | No | LaunchDarkly application host for evaluation-run links (default: `https://app.launchdarkly.com`). Set it for a non-production project, or its runs still link to the production app |
Expand Down Expand Up @@ -165,40 +165,47 @@ A tool result is now judge-prompt input. It stays literal for the same reason th

Judges are resolved through flag delivery, and handlers are matched to them, **before** any evaluation records are created — a missing judge or one no handler covers fails the run up front rather than after the generation spend. After that point a criterion failure never aborts the run: an unparseable judge response, an out-of-range score, a raising handler or scorer, and a row whose generation errored each become a per-criterion `ERROR` event with a cause code (`invalid_judge_output`, `invalid_score`, `handler_raised`, `scorer_raised`, `generation_incomplete`) and a top-level `errorMessage`. Event *delivery* is different: the backend needs one result per `(row, criterion)` to finish row accounting, so if tracking a criterion event fails, every remaining result is still attempted and flushed and then `run()` raises — rather than polling to its timeout with the cause hidden.

The client uses **lazy initialization**: importing the package does not connect to LaunchDarkly. The singleton is created automatically on the first API call that needs it (`config().invoke()`, `graph().invoke()`, `resolve_graph()`, etc.), as long as `LD_SDK_KEY` is set in the environment.
### Give the evaluation tools

Call `init_client()` explicitly when you want to:
- Pass SDK or telemetry options programmatically (overriding env vars)
- Initialize at startup before the first AI call (e.g. to avoid latency on the first request)
- Fail fast at boot if `LD_SDK_KEY` is missing
Pass `tools` to `run()` as a list of `Tool`. Construct one to define a tool in code. Call `evals.tools.get()` to use a tool that already exists in your project's AI library, which reads the tool and pins its version at that point. One list can hold both kinds.

```python
import asyncio
from launchdarkly_ai_server import init_client, shutdown
from launchdarkly_ai_server import Tool, init_evaluations

async def main():
# Standard path — auto-discovers launchdarkly-server-sdk.
client = await init_client({
"sdkKey": "sdk-...",
"serviceName": "my-service",
"environment": "production",
})

# Or skip init_client() and let the first model/graph call initialize lazily.
def lookup_order(order_id: str) -> str:
return f"order {order_id} shipped"

# Flush telemetry, flush LD events, and close the client.
await shutdown()

asyncio.run(main())
evals = init_evaluations(project_key="my-project")

# Read from the AI library now. Pins the version. Raises now if the tool is absent.
search_docs_tool = evals.tools.get("search_docs", implementation=search_docs)

# Defined here. Needs no tool in LaunchDarkly.
lookup_order_tool = Tool(
key="lookup_order",
implementation=lookup_order,
schema={
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
description="Look up an order by id",
)

result = await evals.run(
key="support-qa-2026-08-20",
dataset="support-golden",
handler=create_openai_messages_handler(),
generation={"provider": "OpenAI", "model": "gpt-4o"},
tools=[search_docs_tool, lookup_order_tool],
)
```

| Export | Description |
|---|---|
| `init_client(options?)` | Auto-discover and initialize `launchdarkly-server-sdk`. Optional — the first AI API call triggers lazy init when `LD_SDK_KEY` is set. Returns `Awaitable[LDClientInterface]`. |
| `init_client(client=...)` | **BYOC overload** — accept a pre-initialized `LDClientInterface`. Skips SDK auto-discovery. |
| `get_client()` | Return the initialized `LDClientInterface`. Raises if `init_client` has not completed. |
| `shutdown()` | Flush all events and telemetry, then close the client. Await before process exit. |
| `inspect_config(key, context)` | Read an AI Config variation without invoking the model. Never raises. Returns `{"enabled", "config", "meta"}`. |
`run()` reads no tool from the API. `tools.get()` blocks until its read completes, so call it while you set a run up, not inside a running event loop. A constructed `Tool` is always inline, because `source` and `version` are not constructor arguments. Only `tools.get()` produces a library tool. Handlers receive the same `{key: callable}` map whichever kind a tool is, so handler code needs no change.

**The list is checked before any network I/O.** A blank key, an uppercase key, a `schema` that is not a JSON object, a schema that is not JSON-serializable (including a `NaN` or `Infinity` value), and a non-callable implementation each fail with zero requests issued. A repeated key fails too, and keys are compared without case. A `NativeTool` is valid only for a library tool, because the provider supplies its schema.

### `config(**args)`

Expand Down
2 changes: 1 addition & 1 deletion packages/client/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -131,7 +131,7 @@ Handlers may return any of these — the client normalizes them before emitting

`init_evaluations()` creates an evaluations harness using `LD_API_TOKEN` and the management API host `LD_API_BASE_URI`. Do not reuse `LD_BASE_URI`: that variable configures SDK delivery and may point at a relay proxy. Evaluation-run links use the separate `ui_base_uri` option, then `LD_UI_BASE_URI`, then `https://app.launchdarkly.com`; do not derive their host from `LD_API_BASE_URI`. An event transport is resolved in `init_evaluations()`, which raises before any network I/O when it finds neither an SDK key (`sdk_key` or `LD_SDK_KEY`) nor an already-initialized event-capable client: generation events are the only ingest path for row results, so a run without a transport could never complete. The lifecycle module's bring-your-own-client path (`init_client(client=...)`) therefore satisfies the check on its own, and `run()` reuses that singleton through `_resolve_client`; `run()` raises if the client disappears before it emits. Both polling arguments reject NaN, which would otherwise never compare past a deadline and hang the run. The harness always queues one `$ld:ai:offline-evals:generation` custom event per row through the standard SDK event transport and flushes before returning. No feature flag gates event emission. The harness polls the run summary endpoint until a nonzero `total_rows` has `pending_rows == 0` and `passed + failed + error` rows accounting for the total, polling every `poll_interval_seconds` (default 2s) until `poll_timeout_seconds` (default 180s); both are `run()` arguments so large datasets can widen them. The summary endpoint does not return run state, so `RunSummary` exposes row counts only.

`await EvaluationsModule.run(...)` takes `project_key` per call. Dataset lookup/row pagination, evaluation creation, and run creation are private helpers; only `run()` is public. Each call creates a new evaluation with `POST` and a run with `source="api"`, so its key must be unique. The harness directly invokes the supplied handler once per row and never retries it — event delivery is never a reason to rerun a handler because that would repeat tool side effects; retries apply only to management API requests. A 429 is replayed for any method, but 5xx responses and transport failures are replayed only for `GET`/`HEAD`, so an evaluation or run `POST` that the server may already have applied is never duplicated. Management API calls run in a worker thread (`asyncio.to_thread`) because the client is synchronous; the caller's event loop stays free. Generation events go through the already-initialized SDK client when the application has one — `init_client` is idempotent, so an existing singleton wins and the evaluations SDK key is ignored with a warning. Dataset-owned `input`, `expected_output`, `metadata`, and `variables` are deliberately excluded from the event payload. The harness flushes events, polls the run summary endpoint until row accounting is complete (`total_rows > 0`, `pending_rows == 0`, and `passed + failed + error == total_rows`), and raises a timeout once `poll_timeout_seconds` elapses if the backend never reaches one. `RunSummary` includes row counts only, and `EvalRunResult.passed` is true only when error and pending row counts are both zero.
`init_evaluations()` takes `project_key`; `run()` does not. Dataset lookup/row pagination, evaluation creation, and run creation are private helpers; `run()` and `tools.get()` are the public surface. `tools` is a `list[Tool]`. `evals.tools.get(key, implementation=...)` issues `GET projects/<p>/ai-tools/<key>`, pins the returned version, and builds a library tool through `Tool._library`; a constructed `Tool` is always inline because `source` and `version` are `init=False`. `run()` issues no tool request at all: a library tool was already read by `tools.get()`. Library tools are sent as `{key, version, source: "library"}` and inline tools as `{key, schema, description, source: "inline"}`. `_validate_tools` runs in `run()` before any I/O, so a bad list — a non-`Tool` entry, a blank or uppercase key, a non-object or non-serializable `schema` (`allow_nan=False`), a non-callable implementation, a repeated key (compared without case), or a `NativeTool` on an inline tool — fails with zero requests recorded. `_validate_tool_key` is shared by `tools.get()` and `_validate_tools`, so the key rules are identical on both paths. Handler config synthesis is unforked: `config["tools"][tool.key] = {"description", "parameters"}` is fed from whichever kind the tool is, so handlers cannot tell them apart, and `_tool_handlers` maps each tool to its executable. Each call creates a new evaluation with `POST` and a run with `source="api"`, so its key must be unique. The harness directly invokes the supplied handler once per row and never retries it — event delivery is never a reason to rerun a handler because that would repeat tool side effects; retries apply only to management API requests. A 429 is replayed for any method, but 5xx responses and transport failures are replayed only for `GET`/`HEAD`, so an evaluation or run `POST` that the server may already have applied is never duplicated. Management API calls run in a worker thread (`asyncio.to_thread`) because the client is synchronous; the caller's event loop stays free. Generation events go through the already-initialized SDK client when the application has one — `init_client` is idempotent, so an existing singleton wins and the evaluations SDK key is ignored with a warning. Dataset-owned `input`, `expected_output`, `metadata`, and `variables` are deliberately excluded from the event payload. The harness flushes events, polls the run summary endpoint until row accounting is complete (`total_rows > 0`, `pending_rows == 0`, and `passed + failed + error == total_rows`), and raises a timeout once `poll_timeout_seconds` elapses if the backend never reaches one. `RunSummary` includes row counts only, and `EvalRunResult.passed` is true only when error and pending row counts are both zero.

---

Expand Down
2 changes: 2 additions & 0 deletions packages/client/src/launchdarkly_ai_server/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@
Judge,
RunSummary,
Scorer,
Tool,
init_evaluations,
)
from .graph import GraphInstance, graph, resolve_graph
Expand Down Expand Up @@ -194,6 +195,7 @@
"EvaluationsError",
"EvaluationsModule",
"GenerationConfig",
"Tool",
"Judge",
"RunSummary",
"Scorer",
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@
)
from .criteria import Criterion, Judge, Scorer, SuccessDirection
from .module import EvaluationsModule, init_evaluations
from .tools import Tool, ToolsClient
from .types import (
AIConfig,
DatasetRow,
Expand All @@ -36,6 +37,8 @@
"RunSummary",
"Scorer",
"SuccessDirection",
"Tool",
"ToolsClient",
"Transport",
"Usage",
"init_evaluations",
Expand Down
32 changes: 28 additions & 4 deletions packages/client/src/launchdarkly_ai_server/evaluations/api.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
import urllib.error
import urllib.parse
import urllib.request
from collections.abc import Callable
from collections.abc import Callable, Mapping
from dataclasses import dataclass, field
from datetime import UTC, datetime
from email.utils import parsedate_to_datetime
Expand Down Expand Up @@ -84,15 +84,15 @@ class LDApiClient:

def __init__(
self,
api_token: str,
api_key: str,
base_uri: str = DEFAULT_BASE_URI,
transport: Transport = urllib_transport,
timeout: float = 30.0,
max_retries: int = 3,
sleep: Callable[[float], None] = time.sleep,
random_value: Callable[[], float] = random.random,
) -> None:
self.api_token = api_token
self.api_key = api_key
self.base_uri = base_uri.rstrip("/")
self._transport = transport
self._timeout = timeout
Expand Down Expand Up @@ -135,7 +135,7 @@ def request(
params: dict[str, Any] | None = None,
) -> Any:
headers = {
"Authorization": self.api_token,
"Authorization": self.api_key,
"Accept": "application/json",
"LD-API-Version": "20240415",
"User-Agent": "launchdarkly-ai-evaluations-python",
Expand Down Expand Up @@ -192,3 +192,27 @@ def get(self, path: str, params: dict[str, Any] | None = None) -> Any:

def post(self, path: str, body: Any = None) -> Any:
return self.request("POST", path, body=body)


def segment(value: str) -> str:
"""Percent-encode one path segment."""
return urllib.parse.quote(value, safe="")


def require_mapping(value: Any, *, description: str) -> Mapping[str, Any]:
"""Return ``value`` as a mapping. Raises ``EvaluationsError``."""
if not isinstance(value, Mapping):
raise EvaluationsError(
f"LaunchDarkly returned an invalid {description} response"
)
return value


def require_string(data: Mapping[str, Any], key: str, description: str) -> str:
"""Return the non-empty string at ``key``. Raises ``EvaluationsError``."""
value = data.get(key)
if not isinstance(value, str) or not value:
raise EvaluationsError(
f"LaunchDarkly {description} response is missing string field {key!r}"
)
return value
Loading
Loading