Datasets, diffs, checks and traces for LLM features

Four tools that share one set of real cases. Use them together, or start with the one that hurts most today.

Cases from real traffic, not a spreadsheet

Pull conversations from production traces or a read replica. Tag them, keep a golden set, and let it refresh every night so your tests track what customers ask this month.

Sources

Sources

Production traces, CSV or JSON, a Postgres query, or cases typed by hand.

Golden sets

Golden sets

Pin the cases that must always pass. They run first and can block on their own.

Snapshots

Snapshots

Every refresh keeps a snapshot, so a run from June replays exactly.

app.fenwick.run/datasets

Fenwick dataset view listing real customer messages with tags

app.fenwick.run/datasets

Fenwick dataset view listing real customer messages with tags

app.fenwick.run/prompts

Fenwick prompt diff comparing version 13 and version 14

app.fenwick.run/prompts

Fenwick prompt diff comparing version 13 and version 14

See which answers changed, and why

Put two prompt versions side by side. Every case shows both outputs, the grader scores and a one-line note on what moved.

Side by side

Side by side

Both outputs for every case, with the changed sentences marked.

Graders

Graders

Exact match, JSON schema, regex, or a model grader that writes a one-line note.

Cost and latency

Cost and latency

Tokens, dollars and time per case, next to the quality score.

A red check when quality drops

Fenwick runs the suite on every pull request that touches a prompt. Set the bar once and the merge waits when pass rate or cost crosses it.

Gate rules

Gate rules

Pass rate, golden set, cost and latency limits, each with its own threshold.

The comment

The comment

Ten flipped cases in the pull request comment, with a link to the rest.

Overrides

Overrides

A reviewer can approve a failing check with a written reason, kept in the log.

app.fenwick.run/checks

Fenwick checks list with pull requests blocked or passed

app.fenwick.run/checks

Fenwick checks list with pull requests blocked or passed

app.fenwick.run/traces

Fenwick trace timeline for one production answer

app.fenwick.run/traces

Fenwick trace timeline for one production answer

Every bad answer becomes a test

Production traces show each retrieval, model call and grader with its time and cost. Add any trace to a dataset in one click.

Capture

Capture

SDKs for TypeScript and Python, plus OpenTelemetry. About 2 ms of overhead.

Search

Search

Filter by customer, prompt version, grader score or cost.

To a dataset

To a dataset

Any trace becomes a tagged test case in one click.

Run your next prompt change against 200 real cases this afternoon.

Free for 500 runs a month. Team from $39 a seat.

Use this free template

Create a free website with Framer, the website builder loved by startups, designers and agencies.