Test every prompt change on real cases before it ships.

Fenwick runs each new prompt against hundreds of real customer messages, shows which answers got better or worse, and blocks the merge when quality drops.

Fenwick runs each new prompt against hundreds of real customer messages, shows which answers got better or worse, and blocks the merge when quality drops.

500 runs a month free. No card.

app.fenwick.run/runs/support-reply
Fenwick
Support bot
Runs31
Datasets12
Prompts8
Graders6
Checks4
Traces
Projects
Refund classifier
Onboarding emails
Search answers
Runs / support-reply v14 vs v13RunningSharePromote v14
CasesPrompt diffGradersCost
Pass rate
90.4%
+0.0 pts vs v13
Cases
240
from production, Sep 12
Cost per 1k replies
$1.84
-$0.22 vs v13
Median latency
1.9s
+0.2s vs v13
CaseInputv13v14Score
c-0142My refund still hasn’t shown up after 9 days, order 55120PassRunning0.92
c-0143Can I switch my plan to yearly without losing credits?FailRunning0.88
c-0147I was charged twice for the same seat in AugustPassRunning0.95
c-0151How do I export tickets older than 90 days?PassRunning0.90
c-0158Your bot told me to email legal. Why?FailRunning0.81
c-0160Cancel my account and delete everything todayPassRunning0.54
c-0163Does the Team plan include SSO?PassRunning0.97
c-0166Refund the unused months on my annual invoiceFailRunning0.86
c-0171Where do I change the reply-to address?PassRunning0.93
c-0174The invoice PDF shows the wrong company namePassRunning0.91

Runs on the models you already call

Runs on the models you already call

From a prompt edit to a safe release

From a prompt edit to a safe release

Most LLM bugs are not crashes. They are answers that got a little worse for one group of customers. Fenwick is built to find those before anyone else does.

Most LLM bugs are not crashes. They are answers that got a little worse for one group of customers. Fenwick is built to find those before anyone else does.

Cases from real traffic, not a spreadsheet

Pull conversations from production traces or a read replica. Tag them, keep a golden set, and let it refresh every night so your tests track what customers ask this month.

app.fenwick.run/datasets

Fenwick dataset view listing real customer messages with tags

app.fenwick.run/prompts

Fenwick prompt diff comparing version 13 and version 14

See which answers changed, and why

Put two prompt versions side by side. Every case shows both outputs, the grader scores and a one-line note on what moved.

A red check when quality drops

Fenwick runs the suite on every pull request that touches a prompt. Set the bar once and the merge waits when pass rate or cost crosses it.

app.fenwick.run/checks

Fenwick checks list with pull requests blocked or passed

app.fenwick.run/traces

Fenwick trace timeline for one production answer

Every bad answer becomes a test

Production traces show each retrieval, model call and grader with its time and cost. Add any trace to a dataset in one click.

Set up in an afternoon

Most teams have their first comparison running before lunch and the pull request check on by the end of the day.

Connect your model keys and repo

Add an API key for each provider you call and install the GitHub app on the repo that holds your prompts. Keys stay encrypted in your workspace.

10 minutes

10 minutes

Import 200 real cases

Pull recent conversations from traces, a CSV or a Postgres query. Tag the ones that matter and mark a golden set.

30 minutes

30 minutes

Run it, then turn on the check

Compare your current prompt with the next one. When the numbers look right, set a pass bar and every pull request gets checked.

One run

One run

Priced per seat, with runs to spare

Priced per seat, with runs to spare

Start free on your own. Move to Team when a second person needs to review runs. Every plan includes every model provider.

Start free on your own. Move to Team when a second person needs to review runs. Every plan includes every model provider.

Monthly

Yearly

Yearly billing takes two months off.

Hobby

For one engineer trying Fenwick on a side project

$0

free, no card

500 runs a month

3 datasets, 1 seat

Prompt diffs and graders

Community forum support

Team

For teams shipping AI features every week

$39

per seat a month, billed yearly

25,000 runs a month

Unlimited datasets

Pull request checks

Production traces kept 30 days

Email support within one business day

Business

For companies with a security review

$99

per seat a month, billed yearly

250,000 runs a month

SSO, SCIM and audit log

Traces kept for a year

Self-hosted runner in your cloud

A named support engineer

Monthly

Yearly

Yearly billing takes two months off.

Hobby

For one engineer trying Fenwick on a side project

$0

free, no card

500 runs a month

3 datasets, 1 seat

Prompt diffs and graders

Community forum support

Team

For teams shipping AI features every week

$39

per seat a month, billed yearly

25,000 runs a month

Unlimited datasets

Pull request checks

Production traces kept 30 days

Email support within one business day

Business

For companies with a security review

$99

per seat a month, billed yearly

250,000 runs a month

SSO, SCIM and audit log

Traces kept for a year

Self-hosted runner in your cloud

A named support engineer

What changed for teams that test first

“We caught a refund answer that would have told customers the wrong policy on 11 of 240 cases. The check blocked the merge and the diff showed the one sentence that did it.”

Hannah Brooks

ML lead at Tallow, help desk software, Denver

“We caught a refund answer that would have told customers the wrong policy on 11 of 240 cases. The check blocked the merge and the diff showed the one sentence that did it.”

Hannah Brooks

ML lead at Tallow, help desk software, Denver

“Moving our ticket tagger to a smaller model took one afternoon of runs instead of two weeks of arguing. The bill for tagging dropped to a third.”

Luis Ortega

Staff engineer at Parcelwise, shipping software, Columbus

“Reviewers actually read prompt pull requests now, because the comment shows ten changed cases instead of a wall of text.”

Amara Nwosu

Engineering manager at Ledgerline, billing for clinics, Atlanta

Questions before you connect real data

Something else? Write to hello@fenwick.run and an engineer answers within one business day.

Do you train models on our data?

Can we keep everything inside our own cloud?

What counts as a run?

Which models and frameworks work?

What happens if we go over our runs?

Can we leave and take our data with us?

Run your next prompt change against 200 real cases this afternoon.

Free for 500 runs a month. Team from $39 a seat.

Use this free template

Create a free website with Framer, the website builder loved by startups, designers and agencies.