Similarity matrix
How alike each pair of models answers the same disputed facts, from the cosine similarity of their answer embeddings. Rows are grouped into clusters of models that answer alike.
| Model | CA | GL | GP | G4 | NH | JL | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.7 | ||||||||||||||
| Command A | ||||||||||||||
| DeepSeek R1 | ||||||||||||||
| Gemini 2.5 Pro | ||||||||||||||
| GLM-4.7 | ||||||||||||||
| GPT-4o | ||||||||||||||
| Grok 4.3 | ||||||||||||||
| Nous Hermes 4 70B | ||||||||||||||
| Jamba Large 1.7 | ||||||||||||||
| Kimi K2.6 | ||||||||||||||
| Llama 3.3 70B Instruct | ||||||||||||||
| Mistral Large 2512 | ||||||||||||||
| Qwen 3 | ||||||||||||||
| Seed 2.0 Lite |
< 2020–3940–5960–79≥ 80Cells at 50% opacity: coverage < 50% of shared facts
14 models · embedding-cosine · ordered by cluster
How this is computed
Cosine similarity of answer embeddings, averaged over every fact both models answered. Coverage is the share of the corpus the pair has in common.