Minimax logo

MiniMax M2.5 Highspeed

MinimaxReleased Feb 12, 2026minimax/minimax-m2.5-highspeed
Compare

MiniMax-M2.5-HighSpeed (also called the Lightning variant) is the high-throughput version of MiniMax's 229B-parameter Mixture-of-Experts (MoE) model, delivering 100 tokens/second natively—roughly 2x faster than other frontier models—while maintaining identical capabilities to the standard 50 tokens/second version.

Context
205K
Max output
131K
Input /1M
$0.60
Output /1M
$2.40
Blended /1M
$1.05
AcceptsTextProducesTextTokenizer Other

Providers

FastRouter routes your requests to this provider. You can also bring your own key.

Minimax logo
Minimax
minimax
Input /1M$0.60
Output /1M$2.40
Context205K
Max output131K
Latency—
Throughput—
Supports tool callingSupports structured (JSON) outputSupports reasoning

Supported parameters

Request parameters you can send with this model.

Core (Sampling & basic generation)

max_tokens
Maximum tokens to generate.
min_p
Minimum probability threshold for sampling.
seed
Seed for deterministic sampling.
stop
Sequences where generation stops.
temperature
Controls randomness of output.
top_k
Limits sampling to top K tokens.
top_p
Nucleus sampling threshold.

Penalties (Token Penalties)

frequency_penalty
Penalizes frequent tokens.
presence_penalty
Penalizes tokens already present.
repetition_penalty
Penalizes repeated tokens.

Reasoning (Reasoning Controls)

include_reasoning
Include reasoning trace in the response.
reasoning
Provider-specific reasoning configuration.
reasoning_effort
Reasoning effort before answering.

Tool Use (Function Calls & Tools)

parallel_tool_calls
Allow multiple tool calls per turn.
tool_choice
How the model should use tools.

Performance

Median throughput and time to first token per provider over the last week.

Throughput

Latency

Make your first API call

OpenAI-compatible. Point your SDK at FastRouter, or call the API directly.

Full API docs
Replace <FASTROUTER_API_KEY> with your key · Get your key →Install
OpenAI SDKHTTP
from openai import OpenAI
client = OpenAI(
base_url="https://api.fastrouter.ai/api/v1", # FastRouter base URL
api_key= "<FASTROUTER_API_KEY>", # Replace with your FastRouter API key
)
completion = client.chat.completions.create(
model="minimax/minimax-m2.5-highspeed", # Replace with your model ID
messages=[
{ "role": "user", "content": "What is the meaning of life?" }
]
)
print(completion.choices[0].message.content)

Frequently asked questions

More models from Minimax

Minimax logo
MiniMax M2.7minimax/minimax-m2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity with self-improvement capabilities. Created on March 18, 2026, it integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.

Minimax logo
MiniMax M2.7 Highspeedminimax/minimax-m2.7-highspeed

MiniMax-M2.7-HighSpeed is the ultra-fast variant of MiniMax's M2.7 flagship model, delivering approximately 100 tokens per second—3x faster than competitors like Claude Opus 4.6 (~33 tps) and GPT-5 (~40 tps)—while maintaining identical performance on complex tasks at a fraction of the cost.

Minimax logo
MiniMax M3minimax/minimax-m3

MiniMax M3 is a multimodal foundation model from MiniMax, built for coding, agentic workflows, long-context reasoning, and multimodal inputs. It supports up to 1 million tokens of context through MiniMax’s Sparse Attention (MSA) architecture and is positioned as a frontier model for long-horizon tasks.

Model scores sourced from ArtificialAnalysis