322 evaluations of multi-model ensemble strategies. The synthesized answer beat every individual model 39% of the time, landed between them 40% of the time, and fell below all of them 21% of the time.
| Strategy | Runs | Beat the best model | Average score |
|---|---|---|---|
| Debate Tournament | 26 | 81% | 9.0 |
| Chain-of-Thought | 69 | 54% | 8.7 |
| Competitive Refinement | 58 | 33% | 8.4 |
| Expert Panel | 73 | 30% | 8.5 |
| Collaborative Synthesis | 73 | 29% | 8.5 |
| Red Team / Blue Team | 23 | 26% | 8.9 |