Pick OpenAI, Anthropic, Google, or an open model. Change it per scorecard if you want cheaper binary checks and a heavier model on compliance.
Every row comes with a why. Click it, jump the transcript, and coach from the sentence that actually failed.