New AI Models September 2026: Pricing, Comparison, and API Access
The first ten days of September 2026 brought a run of major AI launches, one after another: OpenAI's GPT-6 Astra, Anthropic's Claude Fable 5.1, Google's Gemini 3.8 Flash, Meta's Muse Spark 1.3, the Qwen team's Qwen3.8 Max 0902 update, and DeepSeek V4.1 Flash reaching general availability. This guide puts the newest AI models on a single page: what each one is for, its Onysoft model ID, its price (in USD and TRY), and its context window.
Alongside the September wave, we cover Grok 4.6 and Gemini 3.7 Flash, which arrived in recent weeks, plus Grok 4.20, GLM-5.3, and Mistral Large 3 as current alternatives, and the Kling 3.0, Nano Banana 2, and ElevenLabs V3 media models. We also flag the models that were announced but are not yet in the Onysoft catalog, and close with a budget-based decision guide. Every catalog model covered here is available through one sk-ony- key and one OpenAI-compatible API.
New AI Models Released in September 2026
Launch details come from the providers' own announcements and press coverage. Prices are Onysoft prices as of September 13, 2026; see /models for current pricing.
1. GPT-6 Astra — OpenAI, September 3, 2026
OpenAI positions GPT-6 Astra as "the world's smartest and most aligned model." According to the announcement, it is built for computer and browser use, software engineering, professional work, and science. It can work across collections of files, write and run code, and complete long tasks with less human steering. OpenAI says Astra is the first model to meet the "Critical" cybersecurity threshold in its Preparedness Framework. For that reason, cyber-sensitive capabilities sit behind a trusted-access program, and the rollout is staged.
- Onysoft ID:
openai/gpt-6-astra. Useopenai/gpt-6-astra-profor the pro reasoning mode, or~openai/gpt-astra-latestto always get the current Astra. - Price: $15 per 1M input tokens, $75 per 1M output tokens. The pro mode has the same token price.
- Half price:
openai/gpt-6-astra:batchis offered as a separate model ID at $7.50 / $37.50. - Context: 1,050,000 tokens, 128K max output.
- Full guide: GPT-6 Astra API guide
2. Claude Fable 5.1 — Anthropic, September 1, 2026
Anthropic positions Fable 5.1 among "the world's most advanced models" for coding and knowledge work. According to the announcement, it matches or beats Fable 5 at low and medium effort and performs far better at high effort. In Claude Code, it produces roughly 60% fewer cybersecurity false positives. Anthropic says the model can discover vulnerabilities but does not develop exploits. The press has reported a Terminal-Bench-Science score of 52.6% (MarkTechPost).
Cache pricing note: Anthropic announced a 75% cut to cache-read pricing on its own direct API. The company says this saves about 25% on typical workloads and up to about 45% on heavy agent workloads. That discount applies to cache-read pricing on Anthropic's direct API; when you use the model through Onysoft, the standard input/output price applies.
- Onysoft ID:
anthropic/claude-fable-5.1. The alias~anthropic/claude-fable-latestalso works.anthropic/claude-fable-5remains active at the same price. - Price: $15 per 1M input, $75 per 1M output.
- Context: 1M tokens, 128K max output.
- Full guide: Claude Fable 5.1 API guide
3. Gemini 3.8 Flash — Google, September 2, 2026
Gemini 3.8 Flash is the third Flash release in six weeks. According to Google, it takes text, image, audio, video, and PDF input. It is built on 3.7 Flash and deliberately spends more thinking tokens. Google reports that it beats 3.7 Flash on every benchmark row it published and surpasses Claude Opus 5 on three benchmarks.
Cost note: Onysoft bills each request as input and output token counts multiplied by the model's price. 3.8 Flash has the same token price as 3.7 Flash and 3.6 Flash. Because it thinks more, though, the same task may use more output tokens.
- Onysoft ID:
google/gemini-3.8-flash - Price: $1.125 per 1M input, $5.625 per 1M output.
- Context: 1,048,576 tokens, 65,536 max output tokens.
- Full guide: Gemini 3.8 Flash API guide
4. Meta Muse Spark 1.3 — Meta, September 2, 2026
Muse Spark 1.3 is Meta's multimodal reasoning model. According to its catalog description, it is designed for long-running agent, multi-agent, and coding workflows and keeps track of information across extended tasks. Independent sources on this release are limited, so we stick to the technical details in the catalog.
- Onysoft ID:
meta/muse-spark-1.3 - Price: $1.875 per 1M input, $6.375 per 1M output.
- Context: 1,048,576 tokens.
5. Qwen3.8 Max 0902 — Qwen, early September
The Qwen team shipped a quiet update to Qwen3.8 Max in early September; 0902 is the label of that snapshot. According to the catalog, it is a 2.4-trillion-parameter mixture-of-experts (MoE) model that accepts text, image, and video input and returns text. For teams that prefer open weights, the family's open-weight variant qwen/qwen3.8-2.4t-a95b (2.4T total, 95B active parameters) is in the catalog at the same price.
- Onysoft ID:
qwen/qwen3.8-max-0902 - Price: $3 per 1M input, $9 per 1M output.
- Context: 1M tokens.
6. DeepSeek V4.1 Flash — DeepSeek, September 10, 2026
DeepSeek V4.1 Flash became generally available on September 10, 2026. According to DeepSeek, it is a 552-billion-parameter multimodal MoE model. It uses a new Causal Encoder-Decoder (CED) architecture, offers native visual understanding, and ships with MIT-licensed weights. It has the lowest token price of the six new models in this section.
- Onysoft ID:
deepseek/deepseek-v4.1-flash - Price: $0.225 per 1M input, $0.90 per 1M output.
- Context: 1,048,576 tokens, 384K max output.
- Family guide: DeepSeek API guide
Announced, but not yet in the Onysoft catalog
- Claude Mythos 5.1: an invite-only model Anthropic announced alongside Fable 5.1. Mythos 5 is not in the catalog either.
- Gemini 3.8 Flash Cyber: the restricted-access "Cyber" sibling of Gemini 3.8 Flash.
- Grok 4.7: announced by xAI on September 12, 2026, and not yet added to the Onysoft catalog.
You cannot call these models through Onysoft right now. Every model added to the catalog shows up on the /models page.
Recent Arrivals and Current Alternatives: Grok, Gemini, GLM, Mistral, and Media
Models that joined the catalog just before the September wave, plus models that are still strong options in any comparison:
- Grok 4.6 (xAI, August 12, 2026): the catalog describes it as the family's smartest model, focused on coding, knowledge work, and STEM. xAI says it has 1.5 trillion parameters. ID
x-ai/grok-4.6, price $3 / $9, 500K-token context. Family guide: Grok API guide. - Grok 4.20 (xAI): built for fast reasoning and agentic tool calling. With a 2M-token context and 1.8M-token max output, it has the largest window on this list. ID
x-ai/grok-4.20, price $1.875 / $3.75. Thex-ai/grok-4.20-multi-agentvariant runs parallel agents at the same price. For agentic coding there is alsox-ai/grok-build-0.1($1.50 / $3, 256K context). - Gemini 3.7 Flash (Google): the release 3.8 Flash is built on. According to the catalog, it is designed for fast agentic workflows, coding, and multi-step reasoning. ID
google/gemini-3.7-flash, price $1.125 / $5.625, 1M-token context.google/gemini-3.6-flashis also active at the same price (see the Gemini 3.6 Flash guide). - GLM-5.3 (Z.ai): a large-scale reasoning model for complex software engineering and long-horizon agent tasks. ID
z-ai/glm-5.3, price $1.638 / $5.148, context of about 1.3M tokens. - Mistral Large 3 (Mistral AI): an MoE model with 675B total and 41B active parameters, Apache 2.0 licensed. ID
mistralai/mistral-large-2512, price $0.75 / $2.25, 262K-token context.
Media models are billed per generation. Their prices are in the table in the next section:
- Kling 3.0 (video): video generation with or without audio at 720p, 1080p, and 4K. Example ID:
kling-3.0-video-without-audio-720p-720p-silent. Turbo and motion-control variants are in the catalog too. More in the AI video generation API guide. - Nano Banana 2 (image): 1K, 2K, and 4K image generation, plus a cheaper Lite version. Example ID:
google-nano-banana-2-1k. - ElevenLabs V3 dialogue (audio): turns text into multi-speaker dialogue audio. ID:
elevenlabs-v3-text-to-dialogue.
Price Comparison Table: Onysoft Prices (USD + TRY)
Text model prices are per 1M tokens. Prices are Onysoft prices as of September 13, 2026; see /models for current pricing. TRY amounts use the Central Bank of the Republic of Türkiye (TCMB) USD/TRY rate of September 11, 2026 (48.4941). Usage is deducted from your balance at the current TCMB rate.
| Model (announced) | Onysoft model ID | Input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| GPT-6 Astra (Sep 3) | openai/gpt-6-astra | $15 TRY 727.41 | $75 TRY 3,637.06 | 1.05M |
| GPT-6 Astra Batch | openai/gpt-6-astra:batch | $7.50 TRY 363.71 | $37.50 TRY 1,818.53 | 1.05M |
| Claude Fable 5.1 (Sep 1) | anthropic/claude-fable-5.1 | $15 TRY 727.41 | $75 TRY 3,637.06 | 1M |
| Gemini 3.8 Flash (Sep 2) | google/gemini-3.8-flash | $1.125 TRY 54.56 | $5.625 TRY 272.78 | 1M |
| Meta Muse Spark 1.3 (Sep 2) | meta/muse-spark-1.3 | $1.875 TRY 90.93 | $6.375 TRY 309.15 | 1M |
| Qwen3.8 Max 0902 (early Sep) | qwen/qwen3.8-max-0902 | $3 TRY 145.48 | $9 TRY 436.45 | 1M |
| DeepSeek V4.1 Flash (Sep 10) | deepseek/deepseek-v4.1-flash | $0.225 TRY 10.91 | $0.90 TRY 43.64 | 1M |
| Grok 4.6 (Aug 12) | x-ai/grok-4.6 | $3 TRY 145.48 | $9 TRY 436.45 | 500K |
| Grok 4.20 | x-ai/grok-4.20 | $1.875 TRY 90.93 | $3.75 TRY 181.85 | 2M |
| Gemini 3.7 Flash | google/gemini-3.7-flash | $1.125 TRY 54.56 | $5.625 TRY 272.78 | 1M |
| GLM-5.3 | z-ai/glm-5.3 | $1.638 TRY 79.43 | $5.148 TRY 249.65 | ~1.3M |
| Mistral Large 3 | mistralai/mistral-large-2512 | $0.75 TRY 36.37 | $2.25 TRY 109.11 | 262K |
Media models are priced per generation:
| Media model | Option | Per generation |
|---|---|---|
| Kling 3.0 (video) | 720p, silent | $0.105 / TRY 5.09 |
| Kling 3.0 (video) | 1080p, silent | $0.135 / TRY 6.55 |
| Kling 3.0 (video) | 4K | $0.5025 / TRY 24.37 |
| Nano Banana 2 (image) | 1K / 2K / 4K | $0.06 / $0.09 / $0.135 TRY 2.91 / 4.36 / 6.55 |
| Nano Banana 2 Lite (image) | 1K | $0.03 / TRY 1.45 |
| ElevenLabs V3 dialogue (audio) | Text to dialogue | $0.105 / TRY 5.09 |
How billing works: Onysoft bills each request as the model's input and output token counts multiplied by its price, and the price table has no separate cache tier. That means provider-specific pricing mechanics on their own APIs, such as the cache-read discount announced with Fable 5.1, do not carry over to your Onysoft bill. The batch discount, on the other hand, comes as a separate model ID: openai/gpt-6-astra:batch is half price. For comparisons with other models, see the AI API pricing guide.
Decision Guide: Which New Model for Your Budget and Task?
The newest model is not the right choice for every job. The matrix below shows both when to pick each model and when not to:
| Need | Recommended model | Why | When not to pick it |
|---|---|---|---|
| Long-running agent work, computer/browser use, scientific analysis | GPT-6 Astra | The provider positions it for these task classes; 1.05M context. The batch variant is half price for bulk jobs | For simple, high-volume traffic: it sits in the most expensive per-token tier of the table |
| Coding and knowledge work, high-effort deep analysis | Claude Fable 5.1 | Per Anthropic, far stronger than Fable 5 at high effort; 1M context, 128K output | For cost-sensitive work: Claude Opus 5 ($7.50 / $37.50) or Sonnet 5 ($3 / $15) may often be enough |
| Fast, multimodal, high-volume generation (text, image, audio, video, PDF) | Gemini 3.8 Flash | Broad input support, 1M context, low token price | For short classification or simple answers: extra thinking can raise output tokens, so compare with 3.7 Flash |
| Long context at the lowest cost among the new models | DeepSeek V4.1 Flash | $0.225 / $0.90; 1M context, 384K output; multimodal | If you only process text and cost comes first: deepseek/deepseek-v4-flash ($0.0735 / $0.147) is still active and cheaper |
| Context up to 2M tokens, fast tool calling | Grok 4.20 | 2M context, 1.8M output, $1.875 / $3.75 | If you want the model positioned as the family's smartest: Grok 4.6 |
| Reasoning over image and video input, open-weight option | Qwen3.8 Max 0902 | Text, image, and video input; 1M context; open-weight variant at the same price | If you need audio or PDF input: Gemini 3.8 Flash |
| Long-horizon agents and coding on a mid-range budget | GLM-5.3 or Muse Spark 1.3 | About 1.3M / 1M context, input price between $1.638 and $1.875 | For critical production work, test on your own data first; independent data on these releases is limited |
| Openly licensed, low-cost general-purpose model | Mistral Large 3 | Apache 2.0 license, $0.75 / $2.25 | If you need more than 262K tokens of context |
| I don't want to pick a model | onysoft/auto | OnyRouter picks the model for you, and routing is free | For reproducible workflows that need the same model on every request |
A practical rule: start with the cheapest candidate that fits your budget, and escalate only the cases where quality falls short. Benchmark tables are the providers' own measurements, so before you decide, run the same prompt against two or three models side by side in the Playground. Current prices by family: GPT, Claude, Gemini, Grok, Qwen, DeepSeek.
Step-by-Step Access, with TRY Billing in Türkiye
None of these new models require a separate provider contract. Through Onysoft AI Gateway, you can get started in a few minutes:
- Create a free account.
- Generate an API key with the
sk-ony-prefix from the dashboard. - Top up your balance. Usage is pay-as-you-go, with no subscription or monthly commitment; teams in Türkiye pay in Turkish lira at the TCMB rate.
- Try the models in the Playground before you write any code.
- Set
base_urltohttps://api.onysoft.com/v1in your app and put the ID from the table above in the model field.
One key, one balance, and one invoice give you access to all 754+ models in the catalog; switching models means changing only the model field. Corporate e-invoicing, KVKK (Turkish data protection law) compliance, and 24/7 support in Turkish come standard. For the big picture, see the AI API guide, and to get to know every model family, see the AI models guide.
Code Examples: Python, curl, and Streaming
The Onysoft endpoint follows the OpenAI schema. Non-stream responses come wrapped in a success / data envelope, so the first example reads the content from resp.json()["data"]. The Python code below sends the same prompt to three new models in turn:
import requests
API_KEY = "sk-ony-YOUR_KEY"
MODELS = ["google/gemini-3.8-flash", "deepseek/deepseek-v4.1-flash", "qwen/qwen3.8-max-0902"]
for model in MODELS:
resp = requests.post(
"https://api.onysoft.com/v1/chat/completions",
headers={"Authorization": f"Bearer {API_KEY}"},
json={
"model": model,
"messages": [{"role": "user", "content": "Classify this customer review as positive, negative, or neutral and explain why."}],
"max_tokens": 800,
},
timeout=120,
)
resp.raise_for_status()
data = resp.json()["data"] # non-stream responses come wrapped in success/data
print(f"--- {model} ---")
print(data["choices"][0]["message"]["content"])The same request with curl:
curl https://api.onysoft.com/v1/chat/completions \
-H "Authorization: Bearer sk-ony-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-6-astra",
"messages": [{"role": "user", "content": "Propose a comprehensive testing strategy for this repository layout."}]
}'Streaming responses arrive as raw OpenAI chunks, so the official openai Python package works as-is:
from openai import OpenAI
client = OpenAI(
base_url="https://api.onysoft.com/v1",
api_key="sk-ony-YOUR_KEY",
)
stream = client.chat.completions.create(
model="anthropic/claude-fable-5.1",
messages=[{"role": "user", "content": "Find potential bugs in this function and suggest fixes."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)To stream with curl, add "stream": true to the request body. The -N flag turns off buffering so output appears as it arrives:
curl -N https://api.onysoft.com/v1/chat/completions \
-H "Authorization: Bearer sk-ony-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"stream": true,
"messages": [{"role": "user", "content": "Compare microservices and monolithic architecture."}]
}'OnyRouter tip: if you are not sure which new model fits your workload, set the model field to onysoft/auto. OnyRouter picks the model for you; routing is free, and you pay only the selected model's price. Details are in the OnyRouter guide, and every parameter and error code is in the API documentation.
Frequently Asked Questions
Which new AI models were released in September 2026?
The headline launches in the first ten days of September 2026 were Claude Fable 5.1 (Anthropic, September 1), Gemini 3.8 Flash (Google, September 2), Meta Muse Spark 1.3 (September 2), GPT-6 Astra (OpenAI, September 3), Qwen3.8 Max 0902 (an early-September update), and DeepSeek V4.1 Flash (generally available September 10). All of them are active in the Onysoft catalog. Grok 4.6 came out earlier, on August 12, 2026.
How much do the new models cost on Onysoft?
Onysoft prices per 1M input and 1M output tokens are: GPT-6 Astra $15 / $75 (batch variant $7.50 / $37.50), Claude Fable 5.1 $15 / $75, Gemini 3.8 Flash $1.125 / $5.625, Meta Muse Spark 1.3 $1.875 / $6.375, Qwen3.8 Max 0902 $3 / $9, and DeepSeek V4.1 Flash $0.225 / $0.90. Prices are Onysoft prices as of September 13, 2026; see the /models page for current pricing. Balances in Türkiye are charged in TRY at the TCMB rate.
How do I choose between GPT-6 Astra, Claude Fable 5.1, and Gemini 3.8 Flash?
GPT-6 Astra and Claude Fable 5.1 are top-tier models with the same token price. Their providers position Astra for computer use, long agent tasks, and science, and Fable 5.1 for coding and knowledge work. Gemini 3.8 Flash costs far less and is a fast multimodal model suited to high-volume work. To decide for sure, run the same prompt against all three side by side in the Playground.
Can I use Claude Mythos 5.1, Gemini 3.8 Flash Cyber, or Grok 4.7 on Onysoft?
No, not right now. Claude Mythos 5.1 is invite-only, Gemini 3.8 Flash Cyber is a restricted-access variant, and Grok 4.7 was only announced on September 12, 2026. None of the three are in the Onysoft catalog yet. Models added to the catalog appear on the /models page.
Does the Claude Fable 5.1 cache discount apply through Onysoft?
No. The 75% cache-read discount Anthropic announced applies to cache-read pricing on Anthropic's direct API. Onysoft bills each request as input and output token counts multiplied by the model's price, so Fable 5.1 is charged at its standard price of $15 per 1M input tokens and $75 per 1M output tokens.
How do I access these new models with TRY billing from Türkiye?
Sign up for Onysoft for free, generate an sk-ony- prefixed API key from the dashboard, and top up your balance in Turkish lira. Then set base_url to https://api.onysoft.com/v1 in your app and set the model to, for example, openai/gpt-6-astra or google/gemini-3.8-flash. You do not need a foreign card or a VPN, and there is no subscription: you pay as you go. Corporate e-invoicing, KVKK compliance, and 24/7 support in Turkish come standard.
Share this article
Related pages
Ready to build?
Access 754+ AI models through a single API. Pay as you go — no subscription.