OpenAI charges $1.20 per million output tokens for GPT-5.6 Luna and fifty dollars for GPT-5.6 Astra. When you cannot point to a published benchmark, that forty-eight-dollar gap becomes the whole conversation. This month's selection leans toward models that make collaboration cheaper, longer, multimodal, or easier to own.

Ten entries made the cut, covering coding, documents, retrieval, music, and visual input.

1. GPT-5.6 Luna

GPT-5.6 Luna

What It Does

GPT-5.6 Luna is a low-cost text model. You access it through an API—an application programming interface that lets your software ask it questions. It handles writing, coding, and instructions.

Its context window holds 1.05 million tokens, which is enough to keep an entire technical library in a single conversation, useful for teams working across massive codebases or document sets.

OpenAI introduced Luna as the budget member of its 5.6 family in September 2026. Promotional pricing runs at least through November. No benchmark scores were published, so the case for it rests on arithmetic, not a leaderboard.

Key Capabilities

  • General reasoning for text tasks, though no specific score is published.
  • Coding assistance via text-in, text-out API calls.
  • Long context of 1,050,000 tokens, with separate pricing for that mode.

Best For

  • Developers building high-volume coding assistants.
  • AI applications that need to keep per-request costs down.
  • Enterprises processing large internal text collections.

What Makes It Different

Cost is its main advantage: $0.20 per million input tokens and $1.20 per million output tokens in short context mode. That is a commercial argument, not evidence it matches premium models on difficult tasks.

Model Access

API access through OpenAI's developer documentation.

Things to Consider

  • Its weights are closed, with no open licence listed.
  • Long-context use costs $0.40 input and $1.80 output per million tokens.
  • The captured material does not describe its refusal behaviour or safety limits.

Our Verdict

Luna suits developers and enterprises that need large volumes of ordinary coding or document work, provided their own evaluation shows its cheaper responses are reliable enough for their risk tolerance.

2. GPT-5.6 Astra

GPT-5.6 Astra

What It Does

GPT-5.6 Astra is a high-end text and coding model for API calls. It can hold over a million tokens in its working memory. This lets a coding assistant examine an entire repository or a long project history without constantly clearing old context.

OpenAI updated its model page on 10 September 2026, documenting long-context pricing between the 11th and 17th. There is no captured benchmark table. The case for Astra rests entirely on its unusually large memory.

Key Capabilities

  • Text responses for software development tasks.
  • A 1.05-million-token context window for large repositories.
  • API use is confirmed, but the payload does not document built-in tools.

Best For

  • Developers analysing entire codebases at once.
  • Researchers comparing long documents in a single prompt.
  • Enterprises building applications around extended records.

What Makes It Different

Scale is its differentiator, not a published quality score. It offers the huge context window, but you pay for it: standard use costs $10 per million input tokens and $50 per million output tokens.

Model Access

API access through OpenAI's developer documentation.

Things to Consider

  • Weights are not open, with no open licence listed.
  • Output costs $50 per million tokens at the standard short-context rate.
  • The captured results provide no limitations page or refusal details.

Our Verdict

Astra is for teams whose applications genuinely need million-token sessions. Choose it only when the reduced context-management work justifies that steep output price.

3. Kimi K3

Kimi K3

What It Does

Kimi K3 is a multimodal reasoning model. It can study more than just text—images and other inputs—and work through a problem before answering. A developer might give it written instructions and visual material, then use its text output for research, coding, or planning.

Moonshot released Kimi K3 in July 2026. By August and September, the captured material still presented it as notable. It combines native multimodal input with a context window of 1,048,576 tokens.

Key Capabilities

  • Multimodal reasoning that processes problems before answering.
  • Accepts multimodal inputs, not just text.
  • Long context of 1,048,576 tokens.
  • Text output that can support development tasks, though no coding benchmark is captured.

Best For

  • Developers building assistants that inspect both text and visuals.
  • Researchers analysing long projects with mixed inputs.
  • AI applications combining visual context with written answers.

What Makes It Different

Kimi K3 pairs a million-token context with open weights and native multimodal reasoning. The captured pages lack a first-party benchmark table. Its advantage over closed rivals remains a capability claim, not a measured verdict.

Model Access

API and open weights, available through the Kimi app and model distribution.

Things to Consider

  • Weights use the Kimi K3 License; detailed terms are not in the payload.
  • API pricing is $3 per million input tokens and $15 per million output tokens.
  • The captured pages do not detail refusal behaviour.

Our Verdict

Kimi K3 deserves attention from developers building long, mixed-input workflows, especially where open weights matter. Teams should run their own accuracy and safety tests, because comparable first-party benchmarks are missing.

4. Qwen 3.8 Max

Qwen 3.8 Max

What It Does

Qwen 3.8 Max is a vast multimodal model for understanding text and images, though its downloadable open version is text-only. It helps with coding, document questions, and agent-style applications. The hosted version keeps up to a million tokens available.

Alibaba's production release landed in July 2026, with reports of both an open model variant and a hosted multimodal product that same month. The hosted service costs $2 per million input tokens and $6 per million output tokens.

Key Capabilities

  • The hosted product accepts multimodal input.
  • Positioned for coding-agent work.
  • 1M tokens hosted; the open version has 262,144 native tokens, extendable to 1,010,000.
  • The open model requires thinking steps, according to the captured material.

Best For

  • Developers experimenting with coding agents.
  • Researchers testing large-model behaviour locally or via API.
  • AI applications using hosted multimodal input.

What Makes It Different

Qwen splits the experience. You get a hosted multimodal product and a text-only open model. That makes a direct comparison awkward, but it gives teams a deployment choice.

Model Access

API, open weights, and cloud platform access for the hosted product and open model download.

Things to Consider

  • The open model uses a custom qwen3.8-max licence.
  • Local deployment is demanding: the model has roughly 2.4 trillion parameters.
  • The open model is text-only, unlike the hosted product.
  • Hosted usage is $2 input and $6 output per million tokens.

Our Verdict

Qwen 3.8 Max fits research teams and developers willing to manage two product variants, especially for coding-agent experiments. Whether it's practical depends on your hardware and if you can accept the open model's text-only limits.

5. GLM 5.2

GLM 5.2

What It Does

GLM 5.2 is an open-weight language model for coding and general text. It uses a mixture-of-experts design, where different parts handle different requests. It gives software a million-token space for long tasks.

Z.ai released GLM 5.2 on 13 June 2026 under the MIT licence. The captured material describes it as roughly three to seven times cheaper than leading closed models. The available benchmark summary lacks a full score table.

Key Capabilities

  • Built for software development and general language tasks.
  • 1M token context.
  • The benchmark summary places it among frontier models, but gives no complete score breakdown.
  • The payload does not document built-in tools.

Best For

  • Developers deploying coding assistants with more control.
  • Researchers examining an open frontier-scale model.
  • Enterprises considering lower-cost API processing.

What Makes It Different

MIT-licensed open weights are the central difference. The official API costs $1.40 input and $4.40 output per million tokens. A missing benchmark breakdown prevents a stronger quality claim.

Model Access

API and open weights, through Z.ai direct and the model distribution.

Things to Consider

  • Weights are available under the MIT licence.
  • Local use is demanding: the model is about 753 billion parameters.
  • Quoted API rates are $1.40 input and $4.40 output per million tokens.
  • The captured pages do not provide refusal behaviour or latency figures.

Our Verdict

GLM 5.2 suits researchers and engineering teams that value open weights and deployment control. They'll need serious hardware or the API, and must validate its incomplete benchmark record.

6. GPT-5.4 mini

GPT-5.4 mini

What It Does

GPT-5.4 mini is a smaller text model for applications that need useful answers without paying frontier-model prices. It powers coding helpers, customer-facing features, and internal automation. Its 400,000-token context holds substantial project material.

OpenAI introduced it on 17 March 2026, and it remained in September pricing documents. Its published price is $0.75 per million input tokens and $4.50 per million output tokens.

Key Capabilities

  • Handles text-based programming assistance via API.
  • 400,000 token context.
  • General-purpose text work, with no captured reasoning benchmark.
  • The payload confirms API use but not built-in tools.

Best For

  • Developers adding coding help to products.
  • AI applications serving many moderate-complexity requests.
  • Enterprises automating internal text workflows.

What Makes It Different

Mini offers a middle ground between cheap, narrow processing and expensive frontier calls. It has more context than many everyday applications need. The payload supplies no benchmark evidence that it beats older models on quality.

Model Access

API access through OpenAI's announcement and pricing documentation.

Things to Consider

  • Weights are closed, with no open licence listed.
  • Output costs $4.50 per million tokens, so long answers accumulate charges.
  • No limitations page or refusal policy was captured.

Our Verdict

GPT-5.4 mini fits product teams needing a general coding or text assistant at moderate scale. They must test whether the smaller model can handle their application's hardest cases.

7. GPT-4.1 mini

GPT-4.1 mini

What It Does

GPT-4.1 mini is a compact text model for coding, writing, and simple automation. Its million-token context lets a developer provide a large set of files. Its smaller size keeps routine API calls cheap.

This is a slow-burn pick. OpenAI announced it on 14 April 2025, but still listed it in September 2026 model documents. Its documented price remains $0.40 input and $1.60 output per million tokens.

Key Capabilities

  • Supports general programming tasks through text.
  • 1.05M tokens in the September 2026 model docs.
  • Handles general text tasks, with no captured benchmark scores.
  • No built-in tool support is documented.

Best For

  • Developers building low-cost code helpers.
  • AI applications handling routine text requests.
  • Enterprises processing long internal documents economically.

What Makes It Different

Its value is persistence and price. It remains a long-context baseline at a fraction of premium-model output cost. This is not a fresh capability story, and the payload gives no current benchmark comparison.

Model Access

API access through OpenAI's product documentation.

Things to Consider

  • Weights are not open, with no open licence listed.
  • Output is $1.60 per million tokens, but usage depends on context sent.
  • No limitations page or refusal details were captured.

Our Verdict

GPT-4.1 mini suits teams seeking a known, inexpensive baseline for routine coding and document work. It's for when long context matters more than access to the newest reasoning features.

8. MiniMax M3

MiniMax M3

What It Does

MiniMax M3 is a multimodal-input language model. It reads more than plain text and returns written answers. Developers could use it for assistants that combine visual or other media context with coding, research, or document tasks.

MiniMax released M3 in June 2026 with a 1-million-token context window, and it remained current in September. DeepInfra pricing in the captured material is $0.28 per million input tokens and $1.10 per million output tokens.

Key Capabilities

  • Accepts multimodal input and returns text.
  • 1M token context.
  • Text output can support coding workflows, though no benchmark is provided.
  • The payload identifies it as a language model but gives no reasoning score.

Best For

  • Developers prototyping mixed-input assistants.
  • AI applications keeping multimodal processing inexpensive.
  • Researchers testing million-token multimodal input.

What Makes It Different

M3 pairs multimodal input and a huge context with low quoted hosted pricing. The evidence is thin. The captured page does not specify first-party access, weights, or a benchmark table.

Model Access

Cloud platform access through DeepInfra, according to the captured pricing.

Things to Consider

  • The captured page does not specify a licence or whether weights are open.
  • Quoted DeepInfra rates are $0.28 input and $1.10 output per million tokens.
  • First-party access details were not captured.
  • Refusal behaviour and latency are not stated.

Our Verdict

MiniMax M3 suits developers exploring affordable multimodal applications. Treat it as an evaluation candidate, not a settled platform, until access, licence, and safety details are clearer.

9. Lyria 2

Lyria 2

What It Does

Lyria 2 is Google DeepMind's music-generation model. You write a prompt, and it produces audio or music. This gives creative teams a way to make draft soundtracks, prototypes, or collaborative musical ideas without starting from scratch.

Lyria 2 remains current in Google DeepMind's September 2026 lineup. The captured material supplies no launch date, benchmark scores, or pricing. This is a slow-burn modality pick, not a release-driven one.

Key Capabilities

  • Documented interaction is text prompt to audio or music.
  • No reasoning benchmark is provided.
  • No autonomous editing or external tool use is documented.
  • Not a coding model; its relevance is collaborative audio creation.

Best For

  • AI applications adding music generation to creative products.
  • Researchers studying prompt-to-audio systems.
  • Enterprises prototyping soundtrack and audio workflows.

What Makes It Different

Lyria 2 broadens a coding-heavy list into audio. Collaboration here means turning a written idea into a musical draft. There's no captured benchmark or cost figure to establish how it compares.

Model Access

Cloud platform access through Google DeepMind's product surface.

Things to Consider

  • Weights are not open, with no open licence listed.
  • No pricing figure was captured.
  • The captured material provides no safety or rights guidance.
  • Access is tied to Google DeepMind's product surface.

Our Verdict

Lyria 2 suits creative-product teams and researchers testing prompt-to-music collaboration. They must accept closed access and establish their own checks for quality, rights, and production consistency.

10. OpenAI text-embedding-3-large

OpenAI text-embedding-3-large

What It Does

This model converts text into a list of numbers called an embedding, which represents meaning in a form software can compare. It lets a team search an internal knowledge base, group similar documents, or retrieve the right code note before another model writes an answer.

This is a slow-burn infrastructure pick. OpenAI still lists it in September 2026. The payload identifies its continuing use in search, clustering, and retrieval pipelines, not a new launch.

Key Capabilities

  • Supplies the search layer that lets an agent find relevant documents.
  • It does not reason or write answers; it represents text for comparison.
  • Can support code and documentation retrieval, though no coding score is captured.
  • No context figure is provided.

Best For

  • Developers building document and code search.
  • AI applications grounding assistants in private information.
  • Enterprises organising large text collections.

What Makes It Different

It doesn't try to answer the user directly. Its job is to make retrieval work, which can improve a collaborative assistant without replacing the model that writes the final response.

Model Access

API access through OpenAI's embeddings guide.

Things to Consider

  • Weights are not open, with no open licence listed.
  • No pricing figure was captured.
  • Retrieval pipelines depend on the OpenAI API.
  • No limitations page was captured.

Our Verdict

This embedding model suits developers building searchable knowledge and coding assistants. They'll need to budget for API dependence and test whether retrieved documents are relevant enough for their final answers.

The strongest case for switching models isn't a leaderboard claim. It's fit. Use Luna for volume, Astra for huge working sets, Kimi and Qwen for multimodal work, GLM for open control.

And remember the retrieval layer—embeddings—that collaboration systems quietly depend on.