Skip to content
#

human-evaluation

Here are 41 public repositories matching this topic...

开源 AI 应用评测平台,支持 RAG、AI Agent、多轮对话、LLM-as-Judge、接口评测、评测报告和人工盲测。Open-source AI evaluation platform for RAG, AI Agents, multi-turn conversations, LLM-as-Judge, endpoint evaluation

  • Updated Jul 20, 2026
  • Python

Concept-Guided Chain-of-Thought (CGCoT) pairwise annotation tool for systematic text evaluation using LLMs. Generate breakdowns, compare items, compute scores, and validate against human judgments. Supports Ollama, Hugging Face, Google Gemini, OpenAI, and Anthropic models.

  • Updated Apr 18, 2026
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the human-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more