New AI Models October 2026: Pricing, Comparison and API
Between late September and early October 2026 a new wave of models was added to the Onysoft catalog: Claude Opus 5.5 and Sonnet 5.5 from Anthropic, GPT-6 Sol, GPT-6 Luna and GPT-6.1 Sol from OpenAI, Grok 4.7 from xAI, the MiMo V2.6 family from Xiaomi, Command A+ from Cohere, Qwen3.8 Max Prime and Omni Flash from Alibaba, and GLM 5.3 Prime from Z.ai.
This guide puts each of them on one page: the API ID on Onysoft, context window, USD and TL list price, and which way the price moved relative to the previous version. The article is based on the Onysoft catalog as of October 5, 2026, not on the makers' announcement dates; model descriptions follow the makers' catalog descriptions and no benchmark scores are quoted. For the previous wave, see the September 2026 guide.
New Models in the Catalog at the Start of October 2026
1. Claude Opus 5.5 and Sonnet 5.5 — Anthropic
Opus 5.5 (anthropic/claude-opus-5.5) is, per the maker's description, the model for demanding reasoning, coding and long-horizon agentic work that succeeds Opus 5. Sonnet 5.5 (anthropic/claude-sonnet-5.5) is the direct upgrade to Sonnet 5 for everyday work. Both offer a 1,000,000-token context. List price (1M input / 1M output): Opus 5.5 $6 / $30, Sonnet 5.5 $3 / $15. Details: Claude Opus 5.5 and Sonnet 5.5 API guide.
2. GPT-6 Sol, GPT-6 Luna and GPT-6.1 Sol — OpenAI
The GPT-6 series is now in the catalog in three price tiers: below the flagship Astra sit Sol (openai/gpt-6-sol, $3 / $15) and, in the speed and cost tier, Luna (openai/gpt-6-luna, $0.15 / $0.75). GPT-6.1 Sol (openai/gpt-6.1-sol), an upgrade to Sol, is the same price as Sol. All three offer a 1,050,000-token context; for requests whose input reaches 272,000 tokens, input is billed at 2 times and output at 1.5 times the list price. Details: GPT-6 Sol, Luna and GPT-6.1 Sol API guide.
3. Grok 4.7 — xAI
Grok 4.7 (x-ai/grok-4.7) is, per the maker's description, the flagship for coding, agentic tasks and knowledge work that succeeds Grok 4.6; the description highlights long-running software engineering tasks and verifying its own work. Context is 500,000 tokens and the list price is $3 / $9 — the same as Grok 4.6. For requests whose input reaches 200,000 tokens, input and output are billed at 2 times the list price. The ~x-ai/grok-latest alias redirects to the latest model in the family. Family page: Grok models.
4. The MiMo V2.6 family — Xiaomi
Per the maker's descriptions, MiMo V2.6 Pro (xiaomi/mimo-v2.6-pro, $0.6525 / $1.305) is the flagship model at a scale of over 1 trillion parameters; MiMo V2.6 Flash (xiaomi/mimo-v2.6-flash, $0.21 / $0.42) is an open-source mixture-of-experts model with 309 billion total parameters and 15 billion activated per token; and MiMo V2.6 Pro UltraSpeed (xiaomi/mimo-v2.6-pro-ultraspeed, $6.525 / $13.05) is the speed edition built from the same model weights as Pro, listed at 10 times the Pro price. All three offer a context of roughly 1 million tokens.
5. Command A+ — Cohere
Command A+ (cohere/command-a-plus) is, per the maker's description, the flagship for enterprise agentic workflows: it accepts text and image inputs, offers a 192,000-token context, and supports native tool calling with strict tool schemas. Its list price is $0.45 / $2.25, which is 88% below the Command A list price ($3.75 / $15) on input and 85% below on output. Family page: Cohere models.
6. Qwen3.8 Max Prime and Qwen3.8 Omni Flash — Alibaba
Qwen3.8 Max Prime (qwen/qwen3.8-max-prime, $6 / $18) is, per the maker's description, a higher-throughput variant of Qwen3.8 Max served as a separate, higher-priced version; it accepts text, image and video input. Its list price is 2 times that of Qwen3.8 Max 0902 ($3 / $9). Qwen3.8 Omni Flash (qwen/qwen3.8-omni-flash, $0.225 / $0.705) is an omni-modal reasoning model described with native audio and video understanding. Both offer a 1,000,000-token context. Family page: Qwen models.
7. GLM 5.3 Prime — Z.ai
GLM 5.3 Prime (z-ai/glm-5.3-prime, $4.20 / $13.20) is, per the maker's description, the high-speed variant of GLM-5.3: it keeps the same capabilities while delivering 1.5–2 times the output throughput. Text in and out, 1,000,000-token context. Its list price is 2 times that of GLM-5.3 ($2.10 / $6.60), so you pay extra for the extra speed.
8. Solar Mini 4 — Upstage
Solar Mini 4 (upstage/solar-mini4, $0.075 / $0.30) is, per the maker's description, a compact, cost-efficient mixture-of-experts model with 35 billion parameters and 3 billion active; it is listed in the catalog with a 524,288-token context. Among the paid chat models in this article it has the lowest input price.
9. On the image side: Seedream 5 Flash — ByteDance
Seedream 5 Flash was added under three IDs: seedream/5-flash-text-to-image, seedream/5-flash-image-to-image and seedream/5-flash-layer-decomposition. All three are billed at a list price of $0.0243 (1.19 TL) per image.
10. Free options
Two new IDs joined the free models whose ID ends in :tr: deepseek/deepseek-v4.1-flash:tr and xiaomi/mimo-v2.6-pro:tr. Nothing is deducted from your balance for these models; a limit of 100 requests per API key in the last 24 hours applies, and it is counted jointly across all :tr models. Details are in the models section of the documentation.
Price Comparison Table: Onysoft List Prices (USD + TL)
List prices per 1 million tokens for the newly added chat models:
| Model | API ID | Context (tokens) | 1M input | 1M output | 1M input (TL) | 1M output (TL) |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 | anthropic/claude-opus-5.5 | 1,000,000 | $6.00 | $30.00 | 294.21 TL | 1,471.04 TL |
| Claude Sonnet 5.5 | anthropic/claude-sonnet-5.5 | 1,000,000 | $3.00 | $15.00 | 147.10 TL | 735.52 TL |
| GPT-6 Sol | openai/gpt-6-sol | 1,050,000 | $3.00 | $15.00 | 147.10 TL | 735.52 TL |
| GPT-6.1 Sol | openai/gpt-6.1-sol | 1,050,000 | $3.00 | $15.00 | 147.10 TL | 735.52 TL |
| GPT-6 Luna | openai/gpt-6-luna | 1,050,000 | $0.15 | $0.75 | 7.36 TL | 36.78 TL |
| Grok 4.7 | x-ai/grok-4.7 | 500,000 | $3.00 | $9.00 | 147.10 TL | 441.31 TL |
| Qwen3.8 Max Prime | qwen/qwen3.8-max-prime | 1,000,000 | $6.00 | $18.00 | 294.21 TL | 882.63 TL |
| Qwen3.8 Omni Flash | qwen/qwen3.8-omni-flash | 1,000,000 | $0.225 | $0.705 | 11.03 TL | 34.57 TL |
| GLM 5.3 Prime | z-ai/glm-5.3-prime | 1,000,000 | $4.20 | $13.20 | 205.95 TL | 647.26 TL |
| MiMo V2.6 Pro | xiaomi/mimo-v2.6-pro | 1,050,000 | $0.6525 | $1.305 | 32.00 TL | 63.99 TL |
| MiMo V2.6 Pro UltraSpeed | xiaomi/mimo-v2.6-pro-ultraspeed | 1,048,576 | $6.525 | $13.05 | 319.95 TL | 639.90 TL |
| MiMo V2.6 Flash | xiaomi/mimo-v2.6-flash | 1,050,000 | $0.21 | $0.42 | 10.30 TL | 20.59 TL |
| Command A+ | cohere/command-a-plus | 192,000 | $0.45 | $2.25 | 22.07 TL | 110.33 TL |
| Solar Mini 4 | upstage/solar-mini4 | 524,288 | $0.075 | $0.30 | 3.68 TL | 14.71 TL |
Prices are Onysoft list prices as of October 5, 2026; see /models for live pricing. TL equivalents use 1 USD = 49.0348 TL (Central Bank of the Republic of Türkiye rate for October 2, 2026); billing uses the current central bank rate at the time of the request.
The Claude models in the table, as well as GPT-6 Sol and Luna, also have a half-price :batch ID (for example anthropic/claude-opus-5.5:batch, openai/gpt-6-luna:batch); the current list is on the /models page. To see rankings by actual usage, visit the model rankings page.
Is the New Version More Expensive? Price Change Version by Version
A new version is not always more expensive. Based on catalog list prices, this wave looks like this:
| Previous | New | Previous price (input / output) | New price (input / output) | Change |
|---|---|---|---|---|
| Claude Opus 5 | Claude Opus 5.5 | $7.50 / $37.50 | $6 / $30 | Down 20% |
| Claude Sonnet 5 | Claude Sonnet 5.5 | $3 / $15 | $3 / $15 | Same |
| GPT-5.6 Luna | GPT-6 Luna | $0.30 / $1.80 | $0.15 / $0.75 | Input down 50%, output down 58% |
| GPT-5.6 Sol | GPT-6 Sol | $3 / $15 | $3 / $15 | Same |
| Grok 4.6 | Grok 4.7 | $3 / $9 | $3 / $9 | Same |
| MiMo V2.5 Pro | MiMo V2.6 Pro | $0.6525 / $1.305 | $0.6525 / $1.305 | Same |
| Command A | Command A+ | $3.75 / $15 | $0.45 / $2.25 | Input down 88%, output down 85% |
The two exceptions are the "Prime" versions, which charge extra for speed: Qwen3.8 Max Prime and GLM 5.3 Prime are listed at 2 times the price of their base models. Choose them only when response speed has business value; otherwise the base model does the same job at half the price.
The takeaway: if you use Opus 5, GPT-5.6 Luna or Command A, moving to the new version lowers the list price; moving from Sonnet 5, GPT-5.6 Sol, Grok 4.6 or MiMo V2.5 Pro leaves the budget unchanged. In every case token consumption can differ between versions, so track the actual amount through the usage and cost fields of the response.
Decision Guide: Which New Model for Which Budget and Task?
The mapping below is based on how the makers describe their models and on the unit prices in the table; treat it as a starting point and verify on your own workload.
| Need | Where to start | List price (input / output) |
|---|---|---|
| High-volume chat, classification, short summaries | openai/gpt-6-luna | $0.15 / $0.75 |
| Lowest input price, simple agent steps | upstage/solar-mini4 | $0.075 / $0.30 |
| Large context (about 1M tokens) at a low price | xiaomi/mimo-v2.6-flash | $0.21 / $0.42 |
| Tool-calling enterprise agent workflows | cohere/command-a-plus | $0.45 / $2.25 |
| Day-to-day development, coding assistance, RAG answers | anthropic/claude-sonnet-5.5 or openai/gpt-6-sol | $3 / $15 |
| Long-running software engineering tasks | x-ai/grok-4.7 or anthropic/claude-opus-5.5 | $3 / $9 and $6 / $30 |
| Work that needs audio and video understanding | qwen/qwen3.8-omni-flash | $0.225 / $0.705 |
| Trials and prototypes without spending balance | xiaomi/mimo-v2.6-pro:tr | Free (daily request limit) |
Rule of thumb: start with the cheapest candidate that fits your budget and move the cases where quality falls short up a tier. Use the Playground to try the same prompt on several models side by side, and OnyRouter (onysoft/auto) to have the model picked automatically per request.
Try Them All with One Key: Python and curl
Every model in this article is called through the same endpoint with the same sk-ony- key; only the model field changes. After you create an account and add balance, the script below runs the same prompt on three of the new models in turn:
from openai import OpenAI
client = OpenAI(
base_url="https://api.onysoft.com/v1",
api_key="sk-ony-YOUR_KEY",
)
CANDIDATES = [
"openai/gpt-6-luna",
"anthropic/claude-sonnet-5.5",
"x-ai/grok-4.7",
]
PROMPT = "Suggest a rate limiting strategy for a REST API in three bullet points."
for model in CANDIDATES:
stream = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": PROMPT}],
max_tokens=400,
stream=True,
)
print("\n==", model)
for chunk in stream:
if not chunk.choices:
continue
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="")A single non-streaming request with curl:
curl https://api.onysoft.com/v1/chat/completions \
-H "Authorization: Bearer sk-ony-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "cohere/command-a-plus",
"messages": [{"role": "user", "content": "Summarize this support ticket and suggest the next step."}]
}'Non-streaming responses come back in a success/data envelope; the OpenAI-format response object is in the data field. Each response carries the amount for that request in USD in the cost field, so you can compare candidates on actual cost as well as on quality. See the API documentation for supported parameters and the model list for all 743 models in the catalog.
Frequently Asked Questions
Which new AI models were added to the Onysoft catalog in October 2026?
According to the catalog as of October 5, 2026, the main models added in the latest wave are: Claude Opus 5.5 and Sonnet 5.5 (Anthropic), GPT-6 Sol, GPT-6 Luna and GPT-6.1 Sol (OpenAI), Grok 4.7 (xAI), MiMo V2.6 Pro, Flash and Pro UltraSpeed (Xiaomi), Command A+ (Cohere), Qwen3.8 Max Prime and Omni Flash (Alibaba), GLM 5.3 Prime (Z.ai), Solar Mini 4 (Upstage) and, on the image side, Seedream 5 Flash (ByteDance).
How much do the new models cost on Onysoft?
As of October 5, 2026, list prices per 1 million input / output tokens are: Claude Opus 5.5 $6 / $30, Claude Sonnet 5.5 $3 / $15, GPT-6 Sol $3 / $15, GPT-6 Luna $0.15 / $0.75, Grok 4.7 $3 / $9, MiMo V2.6 Pro $0.6525 / $1.305, Command A+ $0.45 / $2.25, Qwen3.8 Max Prime $6 / $18, GLM 5.3 Prime $4.20 / $13.20. Billing is in TL at the central bank rate at the time of the request; live prices are on the /models page.
Is Grok 4.7 available on Onysoft?
Yes. It is active in the catalog as x-ai/grok-4.7: a 500,000-token context and a list price of $3 / $9 per 1 million tokens (input / output). The price is the same as Grok 4.6. For requests whose input reaches 200,000 tokens, input and output are billed at 2 times the list price.
Does moving to a new version raise costs?
In this wave, mostly not. By list price, Claude Opus 5.5 is 20% cheaper than Opus 5; GPT-6 Luna is 50% cheaper on input and 58% cheaper on output than GPT-5.6 Luna; Command A+ is 88% cheaper on input and 85% cheaper on output than Command A. Claude Sonnet 5.5, GPT-6 Sol, Grok 4.7 and MiMo V2.6 Pro are the same price as their previous versions. Qwen3.8 Max Prime and GLM 5.3 Prime are speed-oriented versions listed at 2 times the price of their base models.
Are any of the new models free?
Yes. The IDs deepseek/deepseek-v4.1-flash:tr and xiaomi/mimo-v2.6-pro:tr are free; nothing is deducted from your balance for these requests. There is a limit of 100 requests per API key in the last 24 hours, counted jointly across all models whose ID ends in :tr. For production work that has to run without interruption, prefer the paid models.
How do I access these new models?
Sign up for Onysoft AI Gateway, generate an API key with the sk-ony- prefix in the dashboard, and add balance. Use https://api.onysoft.com/v1 as base_url and put one of the IDs from this article in the model field. No separate contract with the model makers is needed; usage is deducted from your balance in TL at the central bank rate at the time of the request.
Share this article
Related pages
Ready to build?
Access 743+ AI models through a single API. Pay as you go — no subscription.