A small team building the test bench for LLM features
We started Fenwick in 2024 after shipping a support bot that got quietly worse for three weeks before anyone noticed.

In 2023 we ran the AI team at a help desk company. A small prompt change made our support bot tell some customers to email the legal team about refunds. Nothing crashed. It took three weeks and a lot of angry tickets to find it.
We had tests for every line of code and none for the sentences our product said to customers. The tools that existed were notebooks and spreadsheets that nobody reran.
Fenwick is what we wanted then: real cases, a clear diff, and a check that stops the merge. We use it on every prompt we ship, including the ones inside Fenwick.
Your customers ask things no public benchmark covers. Test on what they actually say.
Every score in Fenwick links to the cases and the grader note behind it.
Checks block merges, but a reviewer can always override with a written reason.
We never train on customer data and every export is one click.


The people behind it
Fourteen of us, most of whom have shipped an LLM feature to paying customers and been paged about it.
Founder. Ran the AI team at a help desk company for four years.
Engineer. Built the pull request check and the diff view.
ML engineer. Owns the graders and how we measure them.
Head of customers. Onboards every Business account.
Designer. Makes large tables readable.
Engineer. Runs the eval runtime and its on-call.
Run your next prompt change against 200 real cases this afternoon.
Free for 500 runs a month. Team from $39 a seat.