Integrations for the stack you already run
Source control, alerts, data and every major model provider. Most take a few minutes from settings and none need a new SDK.
GitHub
Source control
Run your eval suite on every pull request and block the merge when scores drop.
Setup: 5 minutes
GitLab
Source control
Merge request checks for teams on GitLab.com or a self-managed instance.
Setup: 10 minutes
Vercel
Deploys
Test preview deployments against real cases before they get promoted.
Setup: 5 minutes
Slack
Alerts
Post run results, blocked merges and drifting production scores to a channel.
Setup: 2 minutes
Linear
Alerts
Turn failing cases into issues with the input, output and grader note attached.
Setup: 3 minutes
Datadog
Observability
Send grader scores and cost per request as metrics next to your latency graphs.
Setup: 10 minutes
PostgreSQL
Data
Pull cases straight from a read replica with a SQL query you control.
Setup: 15 minutes
LangChain
Frameworks
Capture traces from LangChain apps with a callback handler, no code rewrite.
Setup: 5 minutes
Sentry
Observability
Link model errors and timeouts to the prompt version that caused them.
Setup: 5 minutes
OpenAI
Model providers
Run any OpenAI chat or reasoning model in your suites with your own key.
Setup: 2 minutes
Anthropic
Model providers
Test Claude models side by side with your current provider.
Setup: 2 minutes
Mistral AI
Model providers
Try open-weight and hosted Mistral models against your golden set.
Setup: 2 minutes
Missing a tool you use?
Run your next prompt change against 200 real cases this afternoon.
Free for 500 runs a month. Team from $39 a seat.