AI evaluation toolkit â measure inter-rater agreement (Fleiss' κ, Kendall's W) across multiple LLM providers
conkurrence