When to use it
Use when comparing two LLM versions, system prompts, or RAG pipelines on identical evaluation benchmark test suites.
Required Inputs
- Model A scores / pass-fail array
- Model B scores / pass-fail array
- Evaluation metric type (Binary Pass/Fail or Continuous 1-5 scale)