A multi-model consensus engine that queries several frontier LLMs in parallel and synthesizes their responses into a single, superior answer.
Instead of relying on a single model, Wave runs your request through a consortium of independent AI systems. Each model analyzes your prompt in isolation — no shared context, no cross-contamination. A fast synthesizer then compares all expert responses, corrects errors, and delivers a refined final answer.
Architecture
Your request is dispatched simultaneously to every model in the selected tier. All models receive the full conversation context — system prompt, history, and user message.
Each model processes the request in complete isolation. No shared state, no influence between models. This independence is critical — it ensures genuine diversity of perspective.
DeepSeek V4 Flash receives all expert responses alongside the original request. It compares outputs, identifies contradictions, corrects hallucinations, and produces a single refined answer.
The synthesized response streams back to the client as a standard OpenAI-compatible completion. The caller sees a single model response — the complexity is fully abstracted.
Request Flow
Models
Seven models power the Wave engine. Each brings different strengths — different architectures, training data, and reasoning approaches.
ali/minimax-m2.5MiniMax M2.5 — strong analytical reasoning with excellent multilingual capabilities.
xiaomi/mimo-v2.5-proMiMo V2.5 Pro — high-performance reasoning with advanced tool-use support.
z/glm-5.2GLM-5.2 — frontier reasoning model with strong analytical capabilities.
ali/qwen3.7-plusQwen 3.7 Plus — latest generation with 1M context window and multilingual strength.
ali/qwen3.6-plusQwen 3.6 Plus — proven workhorse with excellent cost-to-quality ratio.
xiaomi/mimo-v2.5MiMo V2.5 — efficient model balancing speed and intelligence.
ds/deepseek-v4-proDeepSeek V4 Pro — powerful reasoning model with 1M context and thinking mode.
ds/deepseek-v4-flashDeepSeek V4 Flash — ultra-fast synthesizer. Handles all Wave tiers.
Performance
Measured response times and output throughput across all models in the Wave consortium.
Response Latency
Time from request to first token (seconds)
Output Throughput
Tokens generated per second
End-to-End Latency
Expert processing + synthesis overhead per tier
Tiers
Each tier defines which models participate. More models means broader coverage and higher quality — at a proportional cost increase.
wave/ghostMulti-StageMulti-stage deep analysis pipeline.
Input / 1M
$3.00
Output / 1M
$9.50
Stage 1 — Initial Analysis
ali/minimax-m2.5xiaomi/mimo-v2.5-proStage 2 — Review (with Stage 1 history)
z/glm-5.2ali/qwen3.7-plusds/deepseek-v4-flash (synth)Latency
20–35s
Output Speed
~35–55 tok/s
Context
1M
wave/maxMaximum intelligence. No compromises.
Input / 1M
$2.50
Output / 1M
$7.50
Expert Models
z/glm-5.2ali/qwen3.7-plusxiaomi/mimo-v2.5-prods/deepseek-v4-flash (synth)Latency
10–20s
Output Speed
~40–60 tok/s
Context
1M
wave/mediumBalanced performance and cost.
Input / 1M
$1.00
Output / 1M
$3.50
Expert Models
ali/qwen3.6-plusxiaomi/mimo-v2.5ds/deepseek-v4-flash (synth)Latency
8–15s
Output Speed
~50–80 tok/s
Context
1M
wave/fastSingle expert with synthesis pass.
Input / 1M
$0.50
Output / 1M
$1.00
Expert Models
ds/deepseek-v4-prods/deepseek-v4-flash (synth)Latency
5–10s
Output Speed
~60–100 tok/s
Context
1M
Philosophy
When one model hallucinates, others catch it. The synthesizer identifies contradictions between expert outputs and produces the most accurate composite answer.
Different architectures, different training corpora, different strengths. Wave combines reasoning from Zhipu, Alibaba, Xiaomi, and DeepSeek into one coherent response.
If one provider goes down or hits rate limits, others compensate. The consensus degrades gracefully — you always get an answer, even if fewer experts respond.
A fast, cheap synthesizer (DeepSeek V4 Flash at $0.14/1M input) handles the final pass. The heavy lifting is done by specialized models only where needed.
Accuracy
Based on our internal testing with logic puzzles, lateral thinking tasks, and multi-step reasoning benchmarks. Wave tiers show measurable improvement in reasoning accuracy — the synthesizer catches errors, fills gaps, and produces more complete answers than individual models.
Logic & Reasoning Tasks
Correct answer rate on benchmark puzzles
Individual models occasionally miss subtle logical traps or provide incomplete explanations. Wave's synthesis layer identifies contradictions between expert outputs and produces more robust answers.
The improvement is most visible on tasks requiring spatial reasoning, lateral thinking, or multi-step logic. Simple factual queries show less variance.
+13%
Accuracy gain from individual models to Wave/Max
3x
More complete explanations in synthesized responses
0
Hallucinations when all experts agree on facts
Multi-Model Consensus Engine
Powered by Helium AI Gateway · 2026