Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Qwen is a family of models developed by the Qwen Team at Alibaba Group, not a single chatbot. Its public release history runs from early Qwen checkpoints in 2023 through Qwen2, Qwen2.5 and Qwen3, then a later Qwen3.5–Qwen3.8 sequence recorded in the team’s repository through August 2026. The shift toward “agentic” AI is not just a change in model capability: useful agents also need software to coordinate tools, permissions, memory and evaluation.

What is Qwen?

Qwen is a model family whose releases include different generations, sizes and capabilities. A model checkpoint is the downloadable set of learned weights; a chatbot or agent is an application that uses a model, often with additional software and services around it. The distinction matters: the name Qwen alone does not tell you which version, modality, license or deployment route you are getting.

The project’s public release trail begins in 2023. Its early repository records Qwen-7B and Qwen-7B-Chat releases in August, followed by Qwen-14B and Qwen-14B-Chat in September. It also records an Int4 Qwen-7B-Chat release and finetuning support that year. This is the project’s own milestone history, not necessarily a complete record of every model developed internally.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How did Qwen evolve?

The chronology shows a move from early general-purpose checkpoints to successive model generations, mixture-of-experts releases, and later documentation focused on reasoning modes and tool use. Dates below reflect milestones recorded by the Qwen Team, rather than a claim that every model was available in every region or through every service on that date.

Period Milestones recorded by the Qwen Team What the change indicates
2023 Qwen-7B and Qwen-7B-Chat in August; Qwen-14B and Qwen-14B-Chat in September; an Int4 Qwen-7B-Chat release and finetuning support. The first public checkpoints and early options for chat use, quantization and adaptation.
2024 Qwen1.5 in February; Qwen1.5-MoE-A2.7B in March, described by the project as its first MoE release; Qwen2 in June; Qwen2.5 in September. A growing model family, including a mixture-of-experts design as well as new generations.
2025 Qwen3 announced in April; refreshed Qwen3-2507 releases in July and August. The team documented thinking and non-thinking modes, tool integration and a range of model sizes.
2026 The Qwen3.8 repository records Qwen3.5 releases beginning February 16, additional sizes in February and March, Qwen3.6 releases in April, and Qwen3.8 releases in August. The dated sequence extends beyond Qwen3. The repository’s August 2026 entries are the latest milestones covered here.

What changed with Qwen3?

Qwen3 made reasoning mode and model architecture more explicit parts of the family. The Qwen Team documents both thinking and non-thinking modes, along with tool integration in both. In broad terms, a thinking mode is intended for tasks that benefit from more deliberate reasoning; a non-thinking mode is a different operating mode for other uses. Which modes are available and how they behave depends on the specific checkpoint and software used to serve it.

The Qwen Team lists these Qwen3 sizes in its documentation:

  • Dense models: 0.6B, 1.7B, 4B, 8B, 14B and 32B.
  • Mixture-of-experts (MoE) models: 30B-A3B and 235B-A22B. In these names, the A figures denote the active parameter count indicated by the model naming; they should not be mistaken for the total parameter count.

Dense and MoE models make different trade-offs. A dense model uses its parameters throughout inference, while an MoE model routes computation through a subset of its experts for a given input. Parameter counts alone do not establish which model will be faster, more capable or less costly for a particular workload; serving setup, hardware and evaluation task matter too.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Qwen Team also describes Qwen3 as supporting more than 100 languages and dialects. Treat that as the team’s stated coverage, not a guarantee of equal quality across every language or task. Similarly, descriptions of Qwen3 performance on agent tasks are the team’s claims; the material cited here does not establish an independent overall ranking.

What makes an AI system “agentic”?

A model can produce a plan or a tool-call request without being a complete, reliable agent. An operating agent needs an application to decide which tools are available, execute approved calls, handle results, maintain any needed state and respond to errors. Its permissions determine what it is allowed to do; evaluation helps determine whether it does those things correctly and safely.

Qwen-Agent is the Qwen Team’s framework for building such applications around Qwen instruction-following models. The project describes support for tool use, planning and memory, and includes examples such as a browser assistant, a code interpreter and custom assistants. Its documented development includes Qwen3 tool-call demonstrations, MCP cookbooks, Qwen3-Coder and Qwen3-VL tool-call demonstrations, and a Qwen3.5 agent example.

This development supports a concrete distinction: Qwen models provide capabilities that an agent can use, while Qwen-Agent and the application built with it provide some of the surrounding orchestration. It does not establish that every tool call or multi-step task will be dependable. For consequential use, developers still need to test the particular workflow, constrain permissions, inspect tool inputs and outputs, and provide a way to recover from errors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does “open” mean every Qwen model has the same license?

No. “Open-weight” is the more precise term for models whose weights are released for download, and license terms vary across generations. The Qwen3 repository states: “All our open-weight models are licensed under Apache 2.0.” That statement applies to the open-weight models covered by that repository; it should not be extended to every historical Qwen checkpoint.

The older Qwen repository documents separate Tongyi Qianwen license agreements for early Qwen-72B, Qwen-14B and Qwen-7B checkpoints, and describes different terms for Qwen-1.8B. Before using any checkpoint—especially in a commercial product—read the license attached to that exact model and check whether its terms fit the intended use.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How can you run Qwen?

Local inference is an option, not a requirement. Qwen documentation describes multiple local runtimes and serving frameworks, while hosted inference is another possible route. The right choice depends on the particular checkpoint, workload, privacy needs, latency expectations, operational capacity and cost model; the documented material does not provide a like-for-like comparison across every current deployment option.

One Qwen Team example illustrates why deployment figures need context: its Qwen3.8-27B serving command specifies a 262,144-token context length and tensor-parallel size 4. These are settings in that example, not a guarantee that every runtime or request will operate equally at that context length, nor a universal minimum hardware requirement. Tensor parallelism is a way to divide model computation across devices; the example does not establish that every user needs four GPUs or any particular workstation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should you choose a Qwen model or agent setup?

Start with the workload rather than the newest model name. Check the model card and deployment documentation for the exact checkpoint under consideration, then compare the factors that affect your use case:

  • Architecture and size: Determine whether a dense or MoE checkpoint fits your available memory, throughput needs and quality requirements. A model’s headline parameter count is not a substitute for measuring it in your serving setup.
  • Task mode: Check whether the checkpoint supports the reasoning or general-use mode your task calls for, and whether its latency and output behavior fit the application.
  • Modality: Confirm whether the particular release accepts only text or also supports the image, audio or other inputs your application needs. Do not assume that support in one Qwen model applies to the whole family.
  • Context and workload: Match the documented context configuration to the prompt and generation lengths you expect. A published maximum or sample command does not by itself show the performance of your own workload.
  • Deployment route: Compare hosted and local options against privacy requirements, latency, recurring costs, hardware capacity and the effort of operating the system.
  • License and governance: Verify the exact checkpoint’s license and consider how the application handles data and controls model access to tools.
  • Agent framework: For tool-using applications, assess orchestration, available tools, memory behavior, observability, permission boundaries and evaluation—not just the model name.

For performance comparisons, look for results tied to a named checkpoint, dataset, metric, evaluation setup and publishing organization. The Qwen Team’s capability descriptions are useful for understanding what the project says it supports, but they are not an independent head-to-head evaluation of all current models.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.