Qwen logo

Qwen composer: a quiet bar for a multimodal studio

Updated August 25, 2026

Qwen bets on a calm default for a multimodal studio. First send should feel like messaging. Image, video, search, research, and slides wait until intent is clear. Effort is a separate lever from the chat model. Dictation stays in the text loop. Live voice and video leave for their own session. Progressive disclosure is the spine. Clarity frays when multiple models and voice paths compete once a mode is armed.

Calm default

Open chat.qwen.ai on a new chat with the + menu closed.
Open chat.qwen.ai on a new chat with the + menu closed.

What works

  • The default bar reads as messaging, not a mode catalog. Capability waits for intent.
  • The header names the active chat model without forcing a picker on first paint.
  • Sidebar structure (projects, history) stays out of the composer, so the send path stays short.

What we would push on

  • Generic help copy does not preview that this product does images, video, or slides. Discovery depends on opening +.
  • Two voice controls on the same bar without labels. Dictation and live session look like twins until you tap wrong.

Product bet

Alibaba is betting the default loop stays familiar chat. Multimodal and agent jobs stay nested so casual Q&A does not look like a studio on day one.

Tradeoff

DecisionBenefitCost
Calm bar with multimodal modes only in +Low cognitive load on first sendImage, video, and dev work are invisible until +

Takeaway

If + holds your real product surface, keep the default bar typeable and uncluttered. Label voice dictation vs voice chat before users tap the wrong control.

Model choice in the header

Click Qwen3.7-Plus in the header to open the model menu.
Click Qwen3.7-Plus in the header to open the model menu.

What works

  • Each model row gets a job sentence, not only a version ID. People can pick by outcome.
  • Current selection is marked. Collapsed long tail keeps the first open short.
  • Comparison lives as an opt-in in the menu, not a permanent header chip.

What we would push on

  • Version brand names still need a plain reason to pay up. Casual users do not decode Plus vs Max.
  • No speed, context, or price hint on the rows. Cost stays opaque until after send.

Product bet

Qwen ships version numbers as the brand. The picker launches the flagship while a safer default stays selected.

Tradeoff

DecisionBenefitCost
Header model menu with comparison toggle and collapsed long tailFlagship visible; power users can comparePlus/Max jargon; no cost or speed preview

Takeaway

Keep version IDs if that is your brand, but add one plain line per row about speed, cost, or what unlocks on the higher tier.

Pattern: Model Selection UIRows mix version IDs with one-line job copy and a Model Comparison toggle at the top of the menu.

Pattern: Cost Transparency

+ menu: the real product

Click + in the composer to open the attach and modes menu.
Click + in the composer to open the attach and modes menu.

What works

  • Upload spells formats upfront. Multimodal scope is explicit, not implied.
  • Generative modes sit beside search and research as peers. Media and knowledge are one catalog.
  • Deliverable names (sites, slides) beat abstract model feature names.

What we would push on

  • A long single flyout is a lot for first open. Home-row pills advertise jobs without a hunt.
  • If a mode rewrites the bar vs only adds a chip, the menu should say so. Ambiguous arming burns trust.

Product bet

Qwen is selling a multimodal studio inside one chat shell. The + menu is the SKU list.

Tradeoff

DecisionBenefitCost
Single + flyout for attach, generative modes, search, and dev jobsOne attach surface; calm default barHeavy first open; no home-row advertising like Kimi

Takeaway

When + is your mode catalog, spell formats on Upload and group generative vs research vs dev so the list scans fast.

Inline attachment preview

Upload an image via +; keep the composer open with the thumbnail visible.
Upload an image via +; keep the composer open with the thumbnail visible.

What works

  • Attachments live in the bar with remove. Confirm before send without leaving chat.
  • The composer still reads as typeable. Attach does not turn the bar into an upload form.
  • Send state changes when input is ready. Idle and ready feel different.

What we would push on

  • Thumbnail alone is ambiguous across file types. Name or type belongs on the chip.
  • Effort stays visible with no signal that files change which brain or mode runs.

Tradeoff

DecisionBenefitCost
In-bar thumbnail with dismiss, generic placeholder unchangedVisual confirm before send; bar stays familiarNo metadata on the chip; effort picker unexplained with files

Takeaway

Show attachments inside the composer with a remove control. Add file type or filename when the thumbnail alone is ambiguous.

Effort separate from model

Click Auto on the right side of the composer.
Click Auto on the right side of the composer.

What works

  • Effort uses plain words, not version IDs. People tune depth without retooling the model brand.
  • Default is marked. The control sits in the composer where send decisions happen.

What we would push on

  • No outcome copy for time or tokens. Fast vs Thinking is a guess.
  • Header model and effort are two levers with no shared sentence. Users need a map.

Product bet

Auto keeps everyday sends cheap. Thinking is the opt-in for hard prompts without forcing reasoning on every message.

Tradeoff

DecisionBenefitCost
Effort menu separate from header model pickerReasoning depth without retooling the whole barNo outcome copy or latency preview on Thinking

Takeaway

Put thinking effort in the composer, not the model menu. One line per option beats bare labels.

Split voice products

Click the black waveform button on the right of the composer.
Click the black waveform button on the right of the composer.

What works

  • Voice Chat and Video Chat are labeled products, not mic settings.
  • Live session launches from a different control than dictation. Two jobs, two entries.

What we would push on

  • Session limits and which omni model runs should show before connect, not after.
  • Unlabeled twin icons on the bar still invite mis-taps even when the menu is clear.

Product bet

Realtime audio and video are products, not composer modes. A dedicated launcher keeps them out of the text state machine.

Tradeoff

DecisionBenefitCost
Separate launcher for Voice vs Video; mic stays for dictationRealtime sessions do not hijack the text barTwo voice icons; limits hidden until connect

Takeaway

Split dictation from voice chat at the control level. Name both paths on the bar or in a tooltip.

Immersive voice session

Start Voice Chat from the waveform menu.
Start Voice Chat from the waveform menu.

What works

  • Full-screen session signals a mode switch. It feels like a call, not a chat bubble with audio.
  • A visible time cap sets expectation before people settle in.
  • Close, settings, and mute land in predictable corners.

What we would push on

  • Persona choice buried in settings on first visit hides a core product decision.
  • A short cap is honest. It may still feel stingy for real work conversations.

Tradeoff

DecisionBenefitCost
Immersive voice UI with visible session capClear mode switch; time limit upfrontPersona choice deferred to settings; short caps can feel stingy

Takeaway

Voice sessions deserve their own screen and a visible timer. Surface voice choice before connect if personas matter.

Pattern: Voice Input

Voice personas

In Voice Chat, open settings to reach Select voice.
In Voice Chat, open settings to reach Select voice.

What works

  • Search and filter scale when the catalog is long.
  • Personas get personality blurbs and language detail, not only a gender tag.
  • Committing a voice to a fresh thread makes the choice durable.

What we would push on

  • Poetic catalog copy is hard to scan for a meeting. Use-case filters beat vibe alone.
  • Role-play tone as the default catalog skew can fight a general-assistant brand.

Product bet

Voice Chat is partly entertainment. Persona depth and multilingual support sell omni models to people who want character, not just transcription.

Tradeoff

DecisionBenefitCost
Searchable persona catalog with playful copy and language matrixDifferentiation vs flat TTS; global language list visibleHard to pick quickly; tone skews role-play

Takeaway

If you ship personas, add a neutral default and filter by use case, not only search and vibe copy.

Inline dictation

Click the mic in the composer and speak.
Click the mic in the composer and speak.

What works

  • Dictation expands in place with live transcript and confirm/cancel. Edit before send.
  • You stay on the home composer. No mode page for speech-to-text.

What we would push on

  • The live-session control still sits beside the mic. Split jobs need split labels, or people will keep colliding.

Tradeoff

DecisionBenefitCost
In-place dictation card with transcript and confirm/cancelEdit before send; no mode pageCompetes with Voice Chat on the same bar

Takeaway

Dictation should expand the composer with transcript and explicit confirm. Keep it off the same button as realtime voice chat.

Pattern: Voice Input

Create Image mode

From +, choose Create Image.
From +, choose Create Image.

What works

  • Armed mode shows as a removable chip. Scope is visible before spend.
  • Image model and aspect sit with the mode. Format decisions happen pre-prompt.
  • Example gallery teaches the job without leaving the shell.

What we would push on

  • Placeholder should rewrite to the armed job. Generic chat copy fights mode clarity.
  • Chat model in the header plus image model in the bar is dual branding with no bridge.

Product bet

Image gen is a first-class mode, not a plugin. In-bar chip plus gallery competes with browse-first tools while chat stays the shell.

Tradeoff

DecisionBenefitCost
Create Image chip with image-model and aspect controls plus example galleryScope set before generate; browse plus typeDual model pickers; generic placeholder

Takeaway

When image mode is on, show the chip, model, aspect, and examples together. Change placeholder copy to a create job.

Teach prompts from examples

In Create Image mode, hover a gallery card and click Use Prompt.
In Create Image mode, hover a gallery card and click Use Prompt.

What works

  • Example cards reveal real prompt text. Users learn structure, not only aesthetics.
  • One-click inject beats copy-paste from alt text.

What we would push on

  • Hover-only teaching fails on touch. Always-visible beats hover theater.
  • If only some cards teach, the affordance is easy to miss.

Tradeoff

DecisionBenefitCost
Example cards with full prompt text and Use Prompt CTATeaches prompt shape; lowers blank-page anxietyHover dependent; affordance easy to miss

Takeaway

Show the prompt on example cards and let users inject it. Always visible beats hover-only.

Pattern: Prompt StartersFeatured images expose the full generation prompt on hover with a Use Prompt button, not only a silent thumbnail.

Image engine vs chat engine

In Create Image mode, open the Qwen-Image dropdown in the bar.
In Create Image mode, open the Qwen-Image dropdown in the bar.

What works

  • Image models live with the image mode, not the chat header. Engines stay decoupled.
  • Active image engine is marked without leaving the mode.

What we would push on

  • Version-only rows need job copy. Quality and speed should not be a guess.
  • Three model names on one screen without a map is overload.

Tradeoff

DecisionBenefitCost
Nested image-model picker inside Create Image modeChat and image engines stay decoupledNo job copy; model name overload

Takeaway

Split chat and image model pickers, but say what the newer image engine improves in one line inside the menu.

Aspect ratio with the mode

In Create Image mode, open the aspect ratio dropdown.
In Create Image mode, open the aspect ratio dropdown.

What works

  • Format is chosen before the prompt, same pattern as other image studios.
  • Common feed, portrait, and slide shapes are covered.

What we would push on

  • Text-only ratios scan worse than shape icons when the list is mostly numbers.

Takeaway

Put aspect ratio in the mode bar for image gen. Icons help when the list is mostly numbers.

Web search as a mode

From +, choose Web search.
From +, choose Web search.

What works

  • Search is an armed, removable chip, not a hidden always-on globe.
  • Starters under the bar lower the blank-page cost of first search.
  • Effort stays available. Search mode does not rebuild the whole composer.

What we would push on

  • Generic head-term starters miss a trust story. Recency and citations belong in the nudge.
  • Search mode, thinking effort, and Deep Research need a clear relationship or they collide.

Product bet

Web search is a mode chip, not always-on browse. Starters nudge first search without opening a research product.

Tradeoff

DecisionBenefitCost
Removable Web search chip plus query starters under the barMode obvious; low-friction first queryGeneric starters; overlap with Deep Research in +

Takeaway

When search is a mode, show a removable chip and starters that match your trust story, not generic head terms.

Copy this

  • Calm default bar with the multimodal catalog behind +
  • Effort menu separate from header model versions
  • Separate entry for realtime Voice/Video vs in-bar dictation
  • Voice session with its own screen and a visible time cap
  • Create Image as an armed mode with model, aspect, and example gallery
  • Example cards that inject the full generation prompt
  • Web search as a removable chip with starters under the bar

Skip this

  • Unlabeled dictation and live-voice controls on the same bar
  • Model rows without cost, speed, or plain-language why
  • A heavy + catalog with no home advertising for key jobs
  • Poetic voice personas as the only catalog tone
  • Generic placeholder while a generative mode is armed
  • Chat model, image engine, and effort visible with no map between them

How others design the composer

How other products handle the same job, and what each tradeoff reveals.

Compare composer UX across products

ChatGPT, Claude, Perplexity, and Gemini side by side: default bar, tools, cost, and patterns to copy.

Full comparison

Useful for a critique or spec? Share it.

Original gallery pages: Tool Switching in Composer