Defines a new call-quality criterion (yes/no question like "Did the agent mention pricing?"). After every call, the LLM evaluates the transcript and returns SUCCESS/FAILURE/UNKNOWN with a rationale. Use to track agent performance metrics and build quality dashboards.