OpenAI charges $1.20 per million output tokens for GPT-5.6 Luna and fifty dollars for GPT-5.6 Astra. When you cannot point to a published benchmark, that forty-eight-dollar gap becomes the whole conversation. This month's selection leans toward models that make collaboration cheaper, longer, multimodal, or easier to own.
Ten entries made the cut, covering coding, documents, retrieval, music, and visual input.
1. GPT-5.6 Luna

What It Does
GPT-5.6 Luna is a low-cost text model. You access it through an API—an application programming interface that lets your software ask it questions. It handles writing, coding, and instructions.
Its context window holds 1.05 million tokens, which is enough to keep an entire technical library in a single conversation, useful for teams working across massive codebases or document sets.
Why It's Trending
OpenAI introduced Luna as the budget member of its 5.6 family in September 2026. Promotional pricing runs at least through November. No benchmark scores were published, so the case for it rests on arithmetic, not a leaderboard.
Key Capabilities
- General reasoning for text tasks, though no specific score is published.
- Coding assistance via text-in, text-out API calls.
- Long context of 1,050,000 tokens, with separate pricing for that mode.
Best For
- Developers building high-volume coding assistants.
- AI applications that need to keep per-request costs down.
- Enterprises processing large internal text collections.
What Makes It Different
Cost is its main advantage: $0.20 per million input tokens and $1.20 per million output tokens in short context mode. That is a commercial argument, not evidence it matches premium models on difficult tasks.
Model Access
API access through OpenAI's developer documentation.
Things to Consider
- Its weights are closed, with no open licence listed.
- Long-context use costs $0.40 input and $1.80 output per million tokens.
- The captured material does not describe its refusal behaviour or safety limits.
Our Verdict
Luna suits developers and enterprises that need large volumes of ordinary coding or document work, provided their own evaluation shows its cheaper responses are reliable enough for their risk tolerance.
2. GPT-5.6 Astra

What It Does
GPT-5.6 Astra is a high-end text and coding model for API calls. It can hold over a million tokens in its working memory. This lets a coding assistant examine an entire repository or a long project history without constantly clearing old context.
Why It's Trending
OpenAI updated its model page on 10 September 2026, documenting long-context pricing between the 11th and 17th. There is no captured benchmark table. The case for Astra rests entirely on its unusually large memory.
Key Capabilities
- Text responses for software development tasks.
- A 1.05-million-token context window for large repositories.
- API use is confirmed, but the payload does not document built-in tools.
Best For
- Developers analysing entire codebases at once.
- Researchers comparing long documents in a single prompt.
- Enterprises building applications around extended records.
What Makes It Different
Scale is its differentiator, not a published quality score. It offers the huge context window, but you pay for it: standard use costs $10 per million input tokens and $50 per million output tokens.
Model Access
API access through OpenAI's developer documentation.
Things to Consider
- Weights are not open, with no open licence listed.
- Output costs $50 per million tokens at the standard short-context rate.
- The captured results provide no limitations page or refusal details.
Our Verdict
Astra is for teams whose applications genuinely need million-token sessions. Choose it only when the reduced context-management work justifies that steep output price.
3. Kimi K3

What It Does
Kimi K3 is a multimodal reasoning model. It can study more than just text—images and other inputs—and work through a problem before answering. A developer might give it written instructions and visual material, then use its text output for research, coding, or planning.
Why It's Trending
Moonshot released Kimi K3 in July 2026. By August and September, the captured material still presented it as notable. It combines native multimodal input with a context window of 1,048,576 tokens.
Key Capabilities
- Multimodal reasoning that processes problems before answering.
- Accepts multimodal inputs, not just text.
- Long context of 1,048,576 tokens.
- Text output that can support development tasks, though no coding benchmark is captured.
Best For
- Developers building assistants that inspect both text and visuals.
- Researchers analysing long projects with mixed inputs.
- AI applications combining visual context with written answers.
What Makes It Different
Kimi K3 pairs a million-token context with open weights and native multimodal reasoning. The captured pages lack a first-party benchmark table. Its advantage over closed rivals remains a capability claim, not a measured verdict.
Model Access
API and open weights, available through the Kimi app and model distribution.
Things to Consider
- Weights use the Kimi K3 License; detailed terms are not in the payload.
- API pricing is $3 per million input tokens and $15 per million output tokens.
- The captured pages do not detail refusal behaviour.
Our Verdict
Kimi K3 deserves attention from developers building long, mixed-input workflows, especially where open weights matter. Teams should run their own accuracy and safety tests, because comparable first-party benchmarks are missing.
4. Qwen 3.8 Max

What It Does
Qwen 3.8 Max is a vast multimodal model for understanding text and images, though its downloadable open version is text-only. It helps with coding, document questions, and agent-style applications. The hosted version keeps up to a million tokens available.
Why It's Trending
Alibaba's production release landed in July 2026, with reports of both an open model variant and a hosted multimodal product that same month. The hosted service costs $2 per million input tokens and $6 per million output tokens.
Key Capabilities
- The hosted product accepts multimodal input.
- Positioned for coding-agent work.
- 1M tokens hosted; the open version has 262,144 native tokens, extendable to 1,010,000.
- The open model requires thinking steps, according to the captured material.
Best For
- Developers experimenting with coding agents.
- Researchers testing large-model behaviour locally or via API.
- AI applications using hosted multimodal input.
What Makes It Different
Qwen splits the experience. You get a hosted multimodal product and a text-only open model. That makes a direct comparison awkward, but it gives teams a deployment choice.
Model Access
API, open weights, and cloud platform access for the hosted product and open model download.
Things to Consider
- The open model uses a custom qwen3.8-max licence.
- Local deployment is demanding: the model has roughly 2.4 trillion parameters.
- The open model is text-only, unlike the hosted product.
- Hosted usage is $2 input and $6 output per million tokens.
Our Verdict
Qwen 3.8 Max fits research teams and developers willing to manage two product variants, especially for coding-agent experiments. Whether it's practical depends on your hardware and if you can accept the open model's text-only limits.
5. GLM 5.2

What It Does
GLM 5.2 is an open-weight language model for coding and general text. It uses a mixture-of-experts design, where different parts handle different requests. It gives software a million-token space for long tasks.
Why It's Trending
Z.ai released GLM 5.2 on 13 June 2026 under the MIT licence. The captured material describes it as roughly three to seven times cheaper than leading closed models. The available benchmark summary lacks a full score table.
Key Capabilities
- Built for software development and general language tasks.
- 1M token context.
- The benchmark summary places it among frontier models, but gives no complete score breakdown.
- The payload does not document built-in tools.
Best For
- Developers deploying coding assistants with more control.
- Researchers examining an open frontier-scale model.
- Enterprises considering lower-cost API processing.
What Makes It Different
MIT-licensed open weights are the central difference. The official API costs $1.40 input and $4.40 output per million tokens. A missing benchmark breakdown prevents a stronger quality claim.
Model Access
API and open weights, through Z.ai direct and the model distribution.
Things to Consider
- Weights are available under the MIT licence.
- Local use is demanding: the model is about 753 billion parameters.
- Quoted API rates are $1.40 input and $4.40 output per million tokens.
- The captured pages do not provide refusal behaviour or latency figures.
Our Verdict
GLM 5.2 suits researchers and engineering teams that value open weights and deployment control. They'll need serious hardware or the API, and must validate its incomplete benchmark record.
6. GPT-5.4 mini

What It Does
GPT-5.4 mini is a smaller text model for applications that need useful answers without paying frontier-model prices. It powers coding helpers, customer-facing features, and internal automation. Its 400,000-token context holds substantial project material.
Why It's Trending
OpenAI introduced it on 17 March 2026, and it remained in September pricing documents. Its published price is $0.75 per million input tokens and $4.50 per million output tokens.
Key Capabilities
- Handles text-based programming assistance via API.
- 400,000 token context.
- General-purpose text work, with no captured reasoning benchmark.
- The payload confirms API use but not built-in tools.
Best For
- Developers adding coding help to products.
- AI applications serving many moderate-complexity requests.
- Enterprises automating internal text workflows.
What Makes It Different
Mini offers a middle ground between cheap, narrow processing and expensive frontier calls. It has more context than many everyday applications need. The payload supplies no benchmark evidence that it beats older models on quality.
Model Access
API access through OpenAI's announcement and pricing documentation.
Things to Consider
- Weights are closed, with no open licence listed.
- Output costs $4.50 per million tokens, so long answers accumulate charges.
- No limitations page or refusal policy was captured.
Our Verdict
GPT-5.4 mini fits product teams needing a general coding or text assistant at moderate scale. They must test whether the smaller model can handle their application's hardest cases.
7. GPT-4.1 mini

What It Does
GPT-4.1 mini is a compact text model for coding, writing, and simple automation. Its million-token context lets a developer provide a large set of files. Its smaller size keeps routine API calls cheap.
Why It's Trending
This is a slow-burn pick. OpenAI announced it on 14 April 2025, but still listed it in September 2026 model documents. Its documented price remains $0.40 input and $1.60 output per million tokens.
Key Capabilities
- Supports general programming tasks through text.
- 1.05M tokens in the September 2026 model docs.
- Handles general text tasks, with no captured benchmark scores.
- No built-in tool support is documented.
Best For
- Developers building low-cost code helpers.
- AI applications handling routine text requests.
- Enterprises processing long internal documents economically.
What Makes It Different
Its value is persistence and price. It remains a long-context baseline at a fraction of premium-model output cost. This is not a fresh capability story, and the payload gives no current benchmark comparison.
Model Access
API access through OpenAI's product documentation.
Things to Consider
- Weights are not open, with no open licence listed.
- Output is $1.60 per million tokens, but usage depends on context sent.
- No limitations page or refusal details were captured.
Our Verdict
GPT-4.1 mini suits teams seeking a known, inexpensive baseline for routine coding and document work. It's for when long context matters more than access to the newest reasoning features.
8. MiniMax M3

What It Does
MiniMax M3 is a multimodal-input language model. It reads more than plain text and returns written answers. Developers could use it for assistants that combine visual or other media context with coding, research, or document tasks.
Why It's Trending
MiniMax released M3 in June 2026 with a 1-million-token context window, and it remained current in September. DeepInfra pricing in the captured material is $0.28 per million input tokens and $1.10 per million output tokens.
Key Capabilities
- Accepts multimodal input and returns text.
- 1M token context.
- Text output can support coding workflows, though no benchmark is provided.
- The payload identifies it as a language model but gives no reasoning score.
Best For
- Developers prototyping mixed-input assistants.
- AI applications keeping multimodal processing inexpensive.
- Researchers testing million-token multimodal input.
What Makes It Different
M3 pairs multimodal input and a huge context with low quoted hosted pricing. The evidence is thin. The captured page does not specify first-party access, weights, or a benchmark table.
Model Access
Cloud platform access through DeepInfra, according to the captured pricing.
Things to Consider
- The captured page does not specify a licence or whether weights are open.
- Quoted DeepInfra rates are $0.28 input and $1.10 output per million tokens.
- First-party access details were not captured.
- Refusal behaviour and latency are not stated.
Our Verdict
MiniMax M3 suits developers exploring affordable multimodal applications. Treat it as an evaluation candidate, not a settled platform, until access, licence, and safety details are clearer.
9. Lyria 2

What It Does
Lyria 2 is Google DeepMind's music-generation model. You write a prompt, and it produces audio or music. This gives creative teams a way to make draft soundtracks, prototypes, or collaborative musical ideas without starting from scratch.
Why It's Trending
Lyria 2 remains current in Google DeepMind's September 2026 lineup. The captured material supplies no launch date, benchmark scores, or pricing. This is a slow-burn modality pick, not a release-driven one.
Key Capabilities
- Documented interaction is text prompt to audio or music.
- No reasoning benchmark is provided.
- No autonomous editing or external tool use is documented.
- Not a coding model; its relevance is collaborative audio creation.
Best For
- AI applications adding music generation to creative products.
- Researchers studying prompt-to-audio systems.
- Enterprises prototyping soundtrack and audio workflows.
What Makes It Different
Lyria 2 broadens a coding-heavy list into audio. Collaboration here means turning a written idea into a musical draft. There's no captured benchmark or cost figure to establish how it compares.
Model Access
Cloud platform access through Google DeepMind's product surface.
Things to Consider
- Weights are not open, with no open licence listed.
- No pricing figure was captured.
- The captured material provides no safety or rights guidance.
- Access is tied to Google DeepMind's product surface.
Our Verdict
Lyria 2 suits creative-product teams and researchers testing prompt-to-music collaboration. They must accept closed access and establish their own checks for quality, rights, and production consistency.
10. OpenAI text-embedding-3-large

What It Does
This model converts text into a list of numbers called an embedding, which represents meaning in a form software can compare. It lets a team search an internal knowledge base, group similar documents, or retrieve the right code note before another model writes an answer.
Why It's Trending
This is a slow-burn infrastructure pick. OpenAI still lists it in September 2026. The payload identifies its continuing use in search, clustering, and retrieval pipelines, not a new launch.
Key Capabilities
- Supplies the search layer that lets an agent find relevant documents.
- It does not reason or write answers; it represents text for comparison.
- Can support code and documentation retrieval, though no coding score is captured.
- No context figure is provided.
Best For
- Developers building document and code search.
- AI applications grounding assistants in private information.
- Enterprises organising large text collections.
What Makes It Different
It doesn't try to answer the user directly. Its job is to make retrieval work, which can improve a collaborative assistant without replacing the model that writes the final response.
Model Access
API access through OpenAI's embeddings guide.
Things to Consider
- Weights are not open, with no open licence listed.
- No pricing figure was captured.
- Retrieval pipelines depend on the OpenAI API.
- No limitations page was captured.
Our Verdict
This embedding model suits developers building searchable knowledge and coding assistants. They'll need to budget for API dependence and test whether retrieved documents are relevant enough for their final answers.
The strongest case for switching models isn't a leaderboard claim. It's fit. Use Luna for volume, Astra for huge working sets, Kimi and Qwen for multimodal work, GLM for open control.
And remember the retrieval layer—embeddings—that collaboration systems quietly depend on.