Notes on testing LLM features
Short, practical writing from the Fenwick team on evals, graders, cost and the review process around prompt changes.

Evals
Your golden set is lying to you
A dataset you built in March does not describe what customers ask in September. Here is how we keep ours honest.
By Priya Raman

Workflow
What a prompt pull request should show a reviewer
A diff of the prompt text is not enough. Reviewers need the cases that changed.
By Dana Okafor

Evals
Model graders need their own tests
If a model decides whether an answer is correct, something has to check that model.
By Marcus Bell

Cost
Cheaper models, same answers: a test plan
How one team cut tagging cost to a third without losing accuracy, and how to know when you can.
By Priya Raman

Workflow
Traces are test cases you have not written yet
Every bad answer in production is a free test case. Most teams throw them away.
By Dana Okafor

Company
Why we price per seat and cap runs
Usage pricing punishes the teams that test the most. We would rather they test more.
By Priya Raman
Run your next prompt change against 200 real cases this afternoon.
Free for 500 runs a month. Team from $39 a seat.