LIVE

Helium
Wave

A multi-model consensus engine that queries several frontier LLMs in parallel and synthesizes their responses into a single, superior answer.

Instead of relying on a single model, Wave runs your request through a consortium of independent AI systems. Each model analyzes your prompt in isolation — no shared context, no cross-contamination. A fast synthesizer then compares all expert responses, corrects errors, and delivers a refined final answer.

Architecture

How It Works

01

Parallel Expert Query

Your request is dispatched simultaneously to every model in the selected tier. All models receive the full conversation context — system prompt, history, and user message.

02

Independent Analysis

Each model processes the request in complete isolation. No shared state, no influence between models. This independence is critical — it ensures genuine diversity of perspective.

03

Synthesis

DeepSeek V4 Flash receives all expert responses alongside the original request. It compares outputs, identifies contradictions, corrects hallucinations, and produces a single refined answer.

04

Delivery

The synthesized response streams back to the client as a standard OpenAI-compatible completion. The caller sees a single model response — the complexity is fully abstracted.

Request Flow

USER
wave/max
glm-5.2
qwen3.7-plus
mimo-v2.5-pro
SYNTHESIZE
RESPONSE

Models

The Consortium

Seven models power the Wave engine. Each brings different strengths — different architectures, training data, and reasoning approaches.

ali/minimax-m2.5
MiniMaxGHOST

MiniMax M2.5 — strong analytical reasoning with excellent multilingual capabilities.

xiaomi/mimo-v2.5-pro
XiaomiGHOST, MAX

MiMo V2.5 Pro — high-performance reasoning with advanced tool-use support.

z/glm-5.2
Zhipu AIGHOST, MAX

GLM-5.2 — frontier reasoning model with strong analytical capabilities.

ali/qwen3.7-plus
AlibabaGHOST, MAX

Qwen 3.7 Plus — latest generation with 1M context window and multilingual strength.

ali/qwen3.6-plus
AlibabaMEDIUM

Qwen 3.6 Plus — proven workhorse with excellent cost-to-quality ratio.

xiaomi/mimo-v2.5
XiaomiMEDIUM

MiMo V2.5 — efficient model balancing speed and intelligence.

ds/deepseek-v4-pro
DeepSeekFAST

DeepSeek V4 Pro — powerful reasoning model with 1M context and thinking mode.

ds/deepseek-v4-flash
DeepSeekALL

DeepSeek V4 Flash — ultra-fast synthesizer. Handles all Wave tiers.

Performance

Speed Benchmarks

Measured response times and output throughput across all models in the Wave consortium.

Response Latency

Time from request to first token (seconds)

Output Throughput

Tokens generated per second

End-to-End Latency

Expert processing + synthesis overhead per tier

Expert processing (parallel)
Synthesis (sequential)

Tiers

Four Levels of Consensus

Each tier defines which models participate. More models means broader coverage and higher quality — at a proportional cost increase.

GHOST

wave/ghostMulti-Stage

Multi-stage deep analysis pipeline.

Input / 1M

$3.00

Output / 1M

$9.50

Stage 1 — Initial Analysis

ali/minimax-m2.5xiaomi/mimo-v2.5-pro

Stage 2 — Review (with Stage 1 history)

z/glm-5.2ali/qwen3.7-plus
ds/deepseek-v4-flash (synth)

Latency

20–35s

Output Speed

~35–55 tok/s

Context

1M

MAX

wave/max

Maximum intelligence. No compromises.

Input / 1M

$2.50

Output / 1M

$7.50

Expert Models

z/glm-5.2ali/qwen3.7-plusxiaomi/mimo-v2.5-prods/deepseek-v4-flash (synth)

Latency

10–20s

Output Speed

~40–60 tok/s

Context

1M

MEDIUM

wave/medium

Balanced performance and cost.

Input / 1M

$1.00

Output / 1M

$3.50

Expert Models

ali/qwen3.6-plusxiaomi/mimo-v2.5ds/deepseek-v4-flash (synth)

Latency

8–15s

Output Speed

~50–80 tok/s

Context

1M

FAST

wave/fast

Single expert with synthesis pass.

Input / 1M

$0.50

Output / 1M

$1.00

Expert Models

ds/deepseek-v4-prods/deepseek-v4-flash (synth)

Latency

5–10s

Output Speed

~60–100 tok/s

Context

1M

Philosophy

Why Consensus?

Error Correction

When one model hallucinates, others catch it. The synthesizer identifies contradictions between expert outputs and produces the most accurate composite answer.

Diverse Perspectives

Different architectures, different training corpora, different strengths. Wave combines reasoning from Zhipu, Alibaba, Xiaomi, and DeepSeek into one coherent response.

Resilience

If one provider goes down or hits rate limits, others compensate. The consensus degrades gracefully — you always get an answer, even if fewer experts respond.

Cost Efficiency

A fast, cheap synthesizer (DeepSeek V4 Flash at $0.14/1M input) handles the final pass. The heavy lifting is done by specialized models only where needed.

Accuracy

Quality Improves with Consensus

Based on our internal testing with logic puzzles, lateral thinking tasks, and multi-step reasoning benchmarks. Wave tiers show measurable improvement in reasoning accuracy — the synthesizer catches errors, fills gaps, and produces more complete answers than individual models.

Logic & Reasoning Tasks

Correct answer rate on benchmark puzzles

Individual models occasionally miss subtle logical traps or provide incomplete explanations. Wave's synthesis layer identifies contradictions between expert outputs and produces more robust answers.

The improvement is most visible on tasks requiring spatial reasoning, lateral thinking, or multi-step logic. Simple factual queries show less variance.

+13%

Accuracy gain from individual models to Wave/Max

3x

More complete explanations in synthesized responses

0

Hallucinations when all experts agree on facts

Helium Wave

Multi-Model Consensus Engine

Dashboard

Powered by Helium AI Gateway · 2026