Benchmark Results

322 evaluations of multi-model ensemble strategies. The synthesized answer beat every individual model 39% of the time, landed between them 40% of the time, and fell below all of them 21% of the time.

Strategy performance

StrategyRunsBeat the best modelAverage score
Debate Tournament2681%9.0
Chain-of-Thought6954%8.7
Competitive Refinement5833%8.4
Expert Panel7330%8.5
Collaborative Synthesis7329%8.5
Red Team / Blue Team2326%8.9

Read the benchmark analysis