Changelog
What shipped in Fenwick, most weeks, written for the people who run the evals.
v2.14
Gate rules on cost, not just quality
Pull request checks can now fail when cost per request rises more than a limit you set.
Read the entry
New
v2.13
Datasets refresh on a schedule
Point a dataset at a query or a trace filter and it refreshes every night.
Read the entry
New
v2.12
Faster side by side diffs
Prompt diffs now load in under a second on runs with 2,000 cases.
Read the entry
Improved
v2.11
Linear integration
Send failing cases to Linear with both outputs and the grader note attached.
Read the entry
New
v2.10
Grader notes explain the score
Model graders now write one sentence on why a case passed or failed.
Read the entry
Improved
v2.9.3
Fixed: traces over 1 MB
Traces with large retrieval payloads no longer fail to save.
Read the entry
Fixed
v2.9
Self-hosted runner
Run evals inside your own cloud so cases and outputs never leave your network.
Read the entry
New
v2.8
Cost per case in every run
Runs show tokens and dollars per case, per grader and per model.
Read the entry
Improved
Run your next prompt change against 200 real cases this afternoon.
Free for 500 runs a month. Team from $39 a seat.