Last updated: August 2026
A 27-billion-parameter model that watches hour-long videos, reads screenshots and documents in a single one-million-token context, ships under a genuinely permissive Apache-2.0 license, and costs fifty cents per million input tokens on its official hosted API. Any one of those facts would be a normal August headline in 2026. All five together is why the Qwen3.8 open-weights drop, which quietly finished landing on Hugging Face between August 5 and 14, deserves a closer look from anyone who produces content for a living.
The short answer: Qwen3.8-27B is a dense vision-language model (27.78B parameters) released as open weights by Alibaba's Qwen team. It accepts text, images, and video as input, outputs text, extends to a 1M-token context window, and is available three ways: free download under Apache-2.0, community quantizations that run on a 24GB GPU, or a hosted API on Qwen Cloud at $0.50/$3.00 per million tokens.

What actually shipped, and when
Open-source releases usually arrive as a single announcement followed by silence. Qwen did the opposite this month, and the staged rollout matters if you are dating your coverage or your research:
- August 3, 2026 — the official Qwen blog published "Qwen3.8-Max: A New Bar for Coding and Cowork", announcing the flagship Qwen3.8-Max (2.4T parameters, 95B active) on Qwen Cloud and explicitly promising that "the open weights will be released next week."
- August 5, 2026 — the Qwen3.8-27B repository was created on Hugging Face (per the platform's API, which exposes creation timestamps for every repo).
- August 8, 2026 — the giant of the family, Qwen3.8-2.4T-A95B, landed: 2.46T total parameters, 95B active, the first Qwen-Max-class model whose weights were ever published.
- August 13–14, 2026 — the official FP8 quantization of the 27B arrived and the model cards received their final benchmark updates.
- August 17, 2026 — the hosted qwen3.8-27b page on Qwen Cloud was updated with final pricing, rate limits, and a one-million-token free quota for new accounts.
So if you see posts claiming this launched "today", check the date against that timeline. We covered the hosted Qwen3.8-Max launch earlier this month; this piece is about the weights themselves and the smaller sibling, which is where content creators actually live: Qwen 3.8-Max for Content Teams.
The 27B in numbers that matter to writers and marketers
Spec sheets are cheap. Here is what each headline number means for a working content operation:
| Specification | Official value | Why a creator cares |
|---|---|---|
| Parameters | 27.78B dense | Small enough to self-host quantized; large enough for serious drafting |
| Inputs | Text, images, video (hour-scale supported) | Analyze source material, not just prompts |
| Output | Text only | It writes about your media; it does not generate media |
| Context | 262,144 native → 1M via YaRN | Whole books, codebases, or transcripts in one request |
| Thinking mode | On by default, adjustable (low/medium/xhigh), can be disabled | Control the cost/speed/quality tradeoff per task |
| License | Apache-2.0 | Commercial use, redistribution, fine-tuning — all allowed |
| Card benchmarks | IFBench 79.5, CoWorkBench 70.7 | Strong instruction-following for structured briefs and templates |
The instruction-following figure deserves a note, because it is the spec creators actually feel in daily use. IFBench 79.5 measures how reliably a model honors complex formatting and constraint-laden instructions — exactly the skill you depend on when you ask for "a 60-word Instagram caption in a casual tone with one emoji and no hashtags." CoWorkBench 70.7 measures long-horizon office productivity tasks, which is the official card's way of saying the model is tuned for deliverables, not just chat.
Adoption has been unusually fast even by open-weights standards. As of August 19, the official Hugging Face organization page shows more than one million downloads of Qwen3.8-27B in the past month, 11,000+ likes, 608 community quantizations, and 135 fine-tunes — a mature ecosystem two weeks after the repo appeared.
The three ways to run it, priced honestly
Option 1: The hosted API (zero setup)
The official Qwen Cloud page for qwen3.8-27b — last updated August 17 — lists the full commercial terms:
- Input: $0.50 per 1M tokens
- Output: $3.00 per 1M tokens
- Implicit cache reads: $0.10 per 1M; explicit cache creation $0.625, cache reads $0.05
- Full 1M context: up to 991K input tokens, up to 131K output tokens
- Throughput: 5M tokens per minute, 5,000 requests per minute
- A free quota of 1M tokens for new accounts to test before paying
The endpoint is OpenAI-API-compatible (it rides on the DashScope international endpoint), which means existing tooling built for GPT models can switch by changing a base URL and a key. Hosted features include function calling, JSON structured outputs, web search, batch processing, and built-in tools through the Responses API — code interpreter, web extractor, and image search among them — plus fine-tuning on your own data.
Option 2: Self-hosted on consumer hardware
Arithmetic on the published weights gives realistic hardware targets (label these as estimates, because your mileage varies with tooling):
| Build | Approximate size | Practical hardware |
|---|---|---|
| Full BF16 | ~55 GB | Workstation-class GPU or 128 GB unified memory |
| Official FP8 | ~28 GB | 32 GB GPU or 48 GB unified memory |
| Community 4-bit | ~16–17 GB | RTX 4090/5090 (24 GB) or a 32 GB MacBook |
Ollama and LM Studio users can start with the 4-bit builds today; production deployers have an official vLLM recipe published for this exact model. The license permits all of it, including client work.
Option 3: The free quota first, decisions later
Given that a million free tokens are enough to transcribe-analyze several long videos or draft a week of social copy, the rational first move is to burn the free quota on your real workload before choosing between the other two options.

How the price stacks up against the models you already use
Context is everything in pricing conversations, so here are the current official rates for the closest competitors, all verified against their primary sources in our recent coverage:
| Model | Input / 1M | Output / 1M | Vision input | Open weights |
|---|---|---|---|---|
| Qwen3.8-27B | $0.50 | $3.00 | Images + video | Yes (Apache-2.0) |
| Gemini 3.7 Flash | $0.75 intro, then $1.50 | $3.75 intro, then $7.50 | Yes | No |
| Grok 4.6 | $2.00 | $6.00 | Images | No |
| GPT-5.6 Sol | $5.00 | $30.00 | Limited | No |
Read that table carefully before declaring a winner. Qwen3.8-27B is not claiming to out-reason GPT-5.6 Sol on the hardest analytical tasks — the flagship Qwen3.8-Max and GPT's top tier still lead on deep reasoning. The pitch here is unit economics for volume work: transcription analysis, first drafts, content repurposing, bulk classification. At fifty cents per million input tokens, a workload that would cost $500 on one model costs $50 here, and $0 on your own GPU. If you followed our coverage of GPT-5.6 Sol's price cut on OpenRouter, you already know the market is racing downward; this release is the first time the floor has dropped this low with open weights attached.
What this means for your content workflow
- Media analysis becomes a commodity. Feed a webinar recording, a competitor's ad gallery, or a folder of product photos and ask for structured takeaways. The model reads the media directly — no separate transcription or captioning step for the understanding layer.
- Long-source work stops being painful. Affiliate disclosures, platform policy PDFs, partnership agreements: a million-token context swallows them whole, and the adjustable thinking mode lets you spend compute only on the hard questions.
- Privacy-sensitive work has an exit ramp. Agencies handling unreleased campaign material for regional brands can run everything on-premise under Apache-2.0 with no data leaving the building.
- Tool lock-in weakens. OpenAI-compatible endpoints mean your internal tools, scripts, and automation chains can treat Qwen as a drop-in backend and renegotiate the relationship any time.
- The Arabic question stays open — deliberately. The card does not officially claim Arabic support. The 248,320-entry vocabulary implies broad multilingual coverage, but that is inference, not a promise. Test on your own text before committing a workflow.
The 2.4T flagship sibling: read the license before you celebrate
Headlines elsewhere will tell you "Qwen open-sourced its flagship." That is true with an asterisk worth two minutes of your attention.
The Qwen/Qwen3.8-2.4T-A95B repository — live since August 8 — really does contain the Max-class model: 2.46T total parameters, 95B active, text-only, thinking always on, with official card benchmarks of 92.6 on GPQA Diamond and 73.5 on FrontierSWE, explicitly benchmarked against Claude Opus 4.8 and GPT-5.6 Sol on agentic coding. But it ships under a custom "qwen3.8-max" license, not Apache-2.0, and its BF16 weights span roughly 4.9 terabytes — datacenter territory, not desktop hardware.
For content teams, the practical takeaway is simpler: the permissive horse in this race is the 27B. Treat the giant repo as proof of the family's ceiling and a research artifact, and route production work through the Apache-licensed model or the hosted API.
Worth noting for scale-of-capability context: the same official blog documents the flagship autonomously running a software project for sixteen continuous days — 265 commits and 127 merged pull requests with no human intervention. That is developer-flavored news, but it explains why the "Cowork" in the announcement was not marketing puffery: this generation of the family is explicitly tuned to finish deliverables end-to-end.
Honest limits, because every launch has them
- No image or video generation. This model perceives; it does not paint. For generation you still need a dedicated image model — we reviewed Qwen's own Qwen Image 3.0 with Arabic text rendering recently.
- "Free" assumes expensive hardware. The self-hosted path is free of license fees, not of a GPU purchase, electricity, and maintenance time.
- Multilingual quality is unverified beyond the card. Wide vocabulary coverage is not the same as native-quality output in your language; benchmark scores do not translate style.
- The ecosystem is fourteen days old. Community quantizations vary in quality; official FP8 only landed August 13. Expect tooling churn.
- Vision ≠ omniscience. Hour-scale video input is supported at the architecture level, but attention across very long media degrades the way it does for every current model. Verify details that matter.
- Rate limits bind at scale. 5,000 requests per minute is generous for a team and insufficient for a platform; batch mode exists partly for this reason.
A five-step evaluation plan for a content team
- Claim the free million tokens on Qwen Cloud and spend them on your single most repetitive task — the one you would automate first if cost disappeared.
- Run a blind quality test. Same brief, three models: your current default, the 27B hosted, and (if you have the hardware) the local 4-bit build. Score outputs on your own rubric, not public leaderboards.
- Price your real volume. One week of token analytics tells you whether self-hosting pays for itself inside a quarter or never.
- Stress the media pipeline. Upload the longest video and the messiest scanned document you actually own, and see where comprehension breaks — every workflow has a cliff.
- Keep a switching lane open. Because the API mirrors OpenAI's, wire it behind a small abstraction in your scripts so the next price drop — this market moves monthly — costs you a configuration change, not a rebuild.
And a candid note for readers whose priority is polished multilingual output rather than infrastructure tinkering: a dedicated writing platform with a native-language interface will beat a raw model for day-one productivity. Tools like ArWriter exist precisely for that lane — and the two approaches stack cleanly, with the platform for daily drafting and the open model for heavy analysis underneath.
Frequently asked questions
Is Qwen3.8-27B really free?
The weights are free to download, run, fine-tune, and use commercially under Apache-2.0. Running them yourself requires capable hardware (roughly 16–55 GB depending on the build), and the official hosted API is paid: $0.50 per million input tokens and $3.00 per million output tokens, with a 1M-token free trial quota.
What GPU do I need to run Qwen3.8-27B locally?
Estimate about 16–17 GB of memory for a 4-bit quantization, 28 GB for the official FP8 build, and 55 GB for full BF16. In practice, a 24 GB card (RTX 4090 or 5090) handles the quantized version comfortably, as do 32 GB unified-memory Macs.
Can Qwen3.8-27B generate images or video?
No. It is a vision-language model: it reads and analyzes images and video and writes text about them. Generating visuals requires a separate image or video model.
How is Qwen3.8-27B different from Qwen3.8-Max?
Max is the 2.4-trillion-parameter flagship served on Qwen Cloud (with open weights under a custom license). The 27B is the dense, Apache-2.0-licensed sibling with image and video input in the open release itself, designed to be self-hostable.
Does Qwen3.8-27B support a 1M token context out of the box?
The model card lists 262,144 tokens natively, extensible to one million via YaRN configuration. On the official hosted API, the full 1M context is already enabled, with up to 991K input and 131K output tokens per request.
Sources
- Official Qwen3.8-27B model card — Hugging Face
- Qwen3.8-Max announcement — official Qwen blog, August 3, 2026
- Official qwen3.8-27b pricing page — Qwen Cloud
- Qwen3.8-2.4T-A95B repository — Hugging Face
- Official vLLM serving recipe for Qwen3.8-27B
The open-weights floor just dropped to fifty cents per million tokens with a permissive license attached — and for once, the constraint is not the price but your willingness to test whether a 27B model is enough for your workload. Claim the free quota, run your hardest real task through it, and let your own rubric decide. If polished, interface-first writing tools fit your pace better than model plumbing, ArWriter starts free and scales from $4.99 a month.