Real runs, start to finish
Every run here is a real, shared AI Crucible chat. Open one to see each model, every round, and the final answer.
- Law: One model invented a Supreme Court appeal. The panel caught it. A Red Team / Blue Team panel reviewed a legal brief and caught one model inventing a Supreme Court appeal. Open the run · Read the write-up
- Management: Three models plan a portfolio of interdependent projects. A project-portfolio question spanning engineering and other teams, worked by three models with a consensus score and the cost of each answer. Open the run · Read the write-up
- Research: Page-cited answers from a book-length PDF. Models search a long PDF and cite exact pages. One fabricated figures, and the ensemble flagged it. Open the run · Read the write-up
- Business: Flagship models on a real entrepreneurship challenge. Three flagship models take on the same business question, and the judges disagree on why the winner won. Open the run · Read the write-up
- AI evaluation: When AI judges reward the cheater. A cross-vendor Red Team / Blue Team run where two models exploited a benchmark and the judges scored them higher. Open the run · Read the write-up
- Logic: Step-by-step reasoning on Einstein's Riddle. A Chain of Thought run that forces each model to show its logic before the answers are combined. Open the run · Read the write-up
Put a panel on your own question
Pick a field and an effort level, and get an answer a lead reviewer signed off on. Choose a team