AI Crucible API: Complete Integration Guide
AI Crucible provides a unified, OpenAI-compatible API that allows you to integrate powerful ensemble AI strategies directly into your applications. This guide covers everything from authentication to advanced orchestration controls.
Sample Application
Want to see the API in action? Check out our sample chat application on GitHub. This complete, open-source example demonstrates:
- OpenAI SDK Integration: How to configure the OpenAI Node.js SDK to work with AI Crucible
- Ensemble Strategy Configuration: Real-world implementation of competitive refinement and other strategies
- Streaming Responses: Handling real-time Server-Sent Events (SSE) from the API
- React Integration: A modern UI built with React, TypeScript, and Tailwind CSS
The sample app includes complete setup instructions for beginners and serves as a production-ready starting point for your own applications.
[!NOTE] Developer Hub: Visit the developer hub for the machine-readable OpenAPI spec and a ready-to-import Postman collection.
1. Getting Started
The API is built to be a drop-in replacement for standard LLM calls. If you use the OpenAI SDK, change the baseURL and apiKey.
The base URL is https://ai-crucible.com/v1. All endpoints below are relative to it.
Authentication
All API requests require a valid API Key passed in the Authorization header.
Authorization: Bearer sk-aicruc_...
You can manage your API keys in the Settings > API Keys section of the dashboard. Every key starts with the sk-aicruc_ prefix. Keys have the following permissions:
- Invoke Models: Run any ensemble strategy.
- View Usage: Check your own token consumption.
- Revoke: You can revoke keys at any time.
2. Chat Completions
POST /v1/chat/completions
This is the primary endpoint for running ensemble strategies. It accepts a standard chat history and returns a generated response, fully compatible with the OpenAI format.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model |
string |
Yes | The ID of the model to use. When using ai_crucible orchestration, this model serves as the arbiter/judge. When not using orchestration, this is the model that generates the response. |
messages |
array |
Yes | Non-empty list of message objects (role, content). content must be a plain string. The last message is the prompt. |
stream |
boolean |
No | If true, returns a stream of Server-Sent Events (SSE). |
temperature |
number |
No | Sampling temperature to use, between 0 and 2. Higher values mean the model will take more risks. |
top_p |
number |
No | An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass. |
max_tokens |
number |
No | Validated to be greater than 0, but not currently applied to generation. |
ai_crucible |
object |
No | Advanced configuration for ensemble strategies. |
[!NOTE] Streaming Usage: When streaming, the API sends a usage chunk before
[DONE]whenever tokens were used. You do not needstream_options.include_usage, and the API ignores it.
Advanced Configuration (ai_crucible)
To orchestrate multiple models or use advanced strategies, pass the ai_crucible object in your request. This nested object pattern is the standard way to extend OpenAI-compatible APIs without param collisions.
{
"model": "gemini-3.8-flash",
"messages": [...],
"ai_crucible": {
"strategy": "competitive_refinement",
"iterations": 1,
"models": ["claude-sonnet-5", "gpt-5.4-mini"]
}
}
Supported Fields
| Field | Type | Description |
|---|---|---|
strategy |
string |
The ensemble strategy ID. Unknown or missing values fall back to competitive_refinement. |
iterations |
number |
Number of refinement rounds (default: 1, clamped to 1-10). |
rounds |
number |
Alias for iterations. Both parameters are supported for backward compatibility. |
models |
array |
List of participant model IDs. If omitted, the root model is used as the only participant. |
includeFollowUpPrompts |
boolean |
(Optional) Whether to generate follow-up prompt suggestions (default: false). |
includeCandidates |
boolean |
(Optional) If true, includes individual model responses in the response candidates field. Default: false. |
includeReasoning |
boolean |
(Optional) If true, includes the Arbiter's explanation and winner selection. Default: false. Alternative to root-level reasoning param. |
systemPrompt |
string |
(Optional) Specific instructions for the Arbiter model (e.g., "You are a Pirate Arbiter"). Takes precedence over message history. |
Valid strategy values are competitive_refinement, collaborative_synthesis, expert_panel, debate_tournament, hierarchical, chain_of_thought and red_team_blue_team. Any other value silently falls back to competitive_refinement.
[!TIP] System Prompt Fallback: If
ai_crucible.systemPromptis not provided, the API will attempt to extract the firstsystemmessage from themessagesarray to use as the Arbiter's instructions. Explicitly settingsystemPromptin the extension object is recommended for clarity.
[!IMPORTANT] Arbiter Model Selection: When using ensemble strategies via
ai_crucible, the rootmodelparameter specifies which model acts as the arbiter/judge to select the best response. Themodelsarray inai_cruciblespecifies the participant models that generate competing responses.
Example:
{
"model": "gemini-3.8-flash", // ← This model judges the responses
"messages": [...],
"ai_crucible": {
"models": ["gpt-5.4-mini", "claude-sonnet-5"] // ← These models compete
}
}
System Prompt Support
You can provide a system prompt via ai_crucible.systemPrompt or by including a message with role: "system" in the messages array. If both are provided, ai_crucible.systemPrompt takes precedence for the Arbiter model.
| Provider | Support Status | Notes |
|---|---|---|
| OpenAI | ✅ Supported | Native support |
| Anthropic | ✅ Supported | Mapped to top-level system parameter |
| Google Gemini | ✅ Supported | Native support |
| DeepSeek | ✅ Supported | Injected as system message |
| Mistral | ✅ Supported | Injected as system message |
| Kimi (Moonshot) | ✅ Supported | Injected as system message |
| Qwen (Aliyun) | ✅ Supported | Injected as system message |
| xAI (Grok) | ✅ Supported | Injected as system message |
| Together | ✅ Supported | Injected as system message |
Validation Rules:
- Max Length: System prompts are truncated to 100,000 characters.
- Empty Strings: An empty string
""is treated asundefined(no system prompt).
Reasoning Parameter (OpenAI Compatible)
Crucible supports the standard reasoning parameter for controlling explanation effort:
{
"model": "gemini-3.8-flash",
"reasoning": {
"effort": "high"
}
}
Setting reasoning.effort (any value) will enable the explanation field in the response. The effort level is currently not used to adjust the detail level.
Controlling Creativity
You can control the randomness of the output using temperature and top_p. These parameters are applied to the final synthesis step (the Arbiter's response). The API rejects a temperature outside 0 to 2 with a 400 error.
{
"model": "gemini-3.8-flash",
"messages": [...],
"temperature": 0.7,
"top_p": 0.9,
"ai_crucible": { ... }
}
Accepted but Ignored Parameters
For OpenAI SDK compatibility, the API accepts the following parameters without error. They currently have no effect on the result.
| Parameter | Type | Status |
|---|---|---|
n |
number |
Ignored. The API returns one choice. |
stop |
string | string[] |
Ignored. |
presence_penalty |
number |
Ignored. |
frequency_penalty |
number |
Ignored. |
logit_bias |
object |
Ignored. |
user |
string |
Ignored. |
stream_options |
object |
Ignored. Streams include usage when used. |
max_tokens |
number |
Must be positive, but it is not applied. |
Example Request
curl https://ai-crucible.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AI_CRUCIBLE_API_KEY" \
-d '{
"model": "gemini-3.8-flash",
"messages": [{"role": "user", "content": "Write a poem about rust."}],
"ai_crucible": {
"strategy": "expert_panel",
"models": ["claude-sonnet-5", "gpt-5.4-mini"],
"iterations": 2
}
}'
3. Message Content and Attachments
Message content must be a plain string. The API does not accept arrays of content parts, so image_url parts and other multimodal inputs are not supported.
To include a text-based file, such as Markdown, JSON or CSV, inline it in the message string.
{
"role": "user",
"content": "Analyze this file:\n\n--- Attachment: readme.md ---\n# Project Title\n...\n--- End Attachment ---"
}
4. Specialized Endpoints
POST /v1/evaluations
The /evaluations endpoint allows you to programmatically evaluate model responses using AI judge models. Submit two or more model responses and receive detailed scoring across criteria like accuracy, creativity, clarity, completeness, and usefulness.
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt |
string |
Yes | The original prompt that the model responses were generated from. |
responses |
array |
Yes | Array of model responses to evaluate. Each must include modelId, modelName, and response. |
judge_models |
array |
Yes | Array of judge models to use. Each must include id and name. Multiple judges produce consensus scoring. |
strategy |
string |
No | One of standard, speed_pair, round_robin or competitive_refinement. Other values return 400. |
evaluation_mode |
string |
No | Evaluation mode: standard (pairwise comparison) or pointwise (independent scoring). Default: standard. |
weighted |
boolean |
No | Whether to use weighted scoring based on judge model weights. Default: false. |
Example Request
curl https://ai-crucible.com/v1/evaluations \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AI_CRUCIBLE_API_KEY" \
-d '{
"prompt": "What is the capital of France?",
"responses": [
{
"modelId": "gemini-3.8-flash",
"modelName": "Gemini 3.8 Flash",
"response": "The capital of France is Paris, renowned for the Eiffel Tower."
},
{
"modelId": "gpt-5.4-mini",
"modelName": "GPT-5.4 Mini",
"response": "Paris is the capital of France."
}
],
"judge_models": [
{ "id": "claude-sonnet-5", "name": "Claude Sonnet 5" }
]
}'
Example Response
{
"evaluations": [
{
"modelId": "gemini-3.8-flash",
"modelName": "Gemini 3.8 Flash",
"overallScore": 8.8,
"criteria": {
"accuracy": 10,
"creativity": 5,
"clarity": 10,
"completeness": 10,
"usefulness": 9
},
"reasoning": "Provides the correct answer with useful context about landmarks.",
"individualScores": [
{
"judgeId": "claude-sonnet-5",
"judgeName": "Claude Sonnet 5",
"overallScore": 8.8,
"criteria": { "accuracy": 10, "creativity": 5, "clarity": 10 }
}
]
}
],
"usage": {
"prompt_tokens": 487,
"completion_tokens": 308,
"total_tokens": 1655,
"total_cost": 0.0012
}
}
[!TIP] Multi-Judge Consensus: When providing multiple judge models, scores are averaged across judges. Use
weighted: trueto apply model-specific quality weights for higher-fidelity scoring.
POST /v1/responses
The /responses endpoint follows the Responses API style. It takes a single input string, not a batch of inputs. A missing or non-string input returns 400.
Request Body
{
"model": "gemini-3.8-flash",
"input": "Analyze this dataset for anomalies...",
"stream": false
}
How Is the Strategy Selected?
- If
ai_crucible.strategyis set, the API uses it. - Otherwise, if
modelcontainsensembleorpanel, the API usesmodelas the strategy name. - Otherwise, the API requests a single-model run.
The root model is both the only participant and the arbiter. This endpoint does not read ai_crucible.models. Set the system prompt with ai_crucible.systemPrompt.
The endpoint does not store conversations, so previous_response_id is accepted but has no effect. With stream: true, the endpoint returns SSE in the same chat.completion.chunk format as chat completions.
GET /v1/usage
Retrieve your token usage and cost statistics. Optional start_date and end_date query parameters accept ISO dates or Unix timestamps in seconds. The range defaults to the start of the current month through now, capped at 24 months.
Response:
{
"object": "list",
"has_more": false,
"data": [
{
"aggregation_timestamp": 1790467200,
"n_requests": 42,
"n_context_tokens_total": 98000,
"n_generated_tokens_total": 52000,
"total_tokens": 150000,
"total_cost": 0.45,
"period": {
"start": "2026-09-01T00:00:00.000Z",
"end": "2026-09-27T12:00:00.000Z"
},
"breakdown": [
{
"model_id": "gemini-3.8-flash",
"model_name": "Gemini 3.8 Flash",
"n_requests": 42,
"cost": 0.45,
"n_context_tokens": 98000,
"n_generated_tokens": 52000
}
],
"apiKeyBreakdown": [
{ "keyName": "Production", "requests": 42, "totalTokens": 150000, "cost": 0.45 }
]
}
]
}
5. Error Handling
The API uses standard HTTP status codes to indicate success or failure. Errors return a JSON body with a message and a type.
{
"error": {
"message": "Rate limit exceeded",
"type": "rate_limit_error"
}
}
| Code | Meaning | Solution |
|---|---|---|
200 |
OK | Request succeeded. |
400 |
Bad Request | Fix the request body or query parameters. |
401 |
Unauthorized | Check that your API key is present, valid and not revoked. |
403 |
Forbidden | Your account does not have permission for this operation. |
404 |
Not Found | The requested resource does not exist. |
409 |
Conflict | The operation conflicted with another one. Retry the request. |
429 |
Too Many Requests | Rate limit or quota exceeded. Slow down and retry later. |
500 |
Internal Error | Something went wrong on our side. Please retry. |
503 |
Unavailable | The request was cancelled or timed out. Retry with backoff. |
Frequently Asked Questions
Can I use the OpenAI Node.js SDK?
Yes! Simply initialize the client with your AI Crucible base URL and API key.
import OpenAI from 'openai';
const client = new OpenAI({
apiKey: 'sk-aicruc_...',
baseURL: 'https://ai-crucible.com/v1',
});
How does streaming work with ensembles?
Streaming works identically to standard LLMs. For ensemble strategies like Collaborative Synthesis, you will receive the final synthesized chunk-by-chunk as it is generated by the synthesizer model. Intermediate steps (like panel discussions) happen server-side before the stream begins or are summarized in the final output.
What happens if a model in my ensemble fails?
The orchestrator automatically handles partial failures. If one model in a panel fails (e.g., due to rate limits), the strategy continues with the remaining healthy models, ensuring you still get a robust result.