Most teams build a golden set once, in a spreadsheet, the week before launch. It passes. Six months later it still passes, and support tickets say the bot is getting worse.
The set did not change. The questions did. New plans, new features and a pricing change all moved what people ask.
We refresh ours every night from production traces, keep the last ten snapshots and look at which cases are new each week. The new cases are where the regressions hide.
