Qwen composer: a quiet bar for a multimodal studio
Qwen bets on a calm default for a multimodal studio. First send should feel like messaging. Image, video, search, research, and slides wait until intent is clear. Effort is a separate lever from the chat model. Dictation stays in the text loop. Live voice and video leave for their own session. Progressive disclosure is the spine. Clarity frays when multiple models and voice paths compete once a mode is armed.
Calm default

What works
- The default bar reads as messaging, not a mode catalog. Capability waits for intent.
- The header names the active chat model without forcing a picker on first paint.
- Sidebar structure (projects, history) stays out of the composer, so the send path stays short.
What we would push on
- Generic help copy does not preview that this product does images, video, or slides. Discovery depends on opening +.
- Two voice controls on the same bar without labels. Dictation and live session look like twins until you tap wrong.
Product bet
Alibaba is betting the default loop stays familiar chat. Multimodal and agent jobs stay nested so casual Q&A does not look like a studio on day one.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Calm bar with multimodal modes only in + | Low cognitive load on first send | Image, video, and dev work are invisible until + |
Takeaway
If + holds your real product surface, keep the default bar typeable and uncluttered. Label voice dictation vs voice chat before users tap the wrong control.
Pattern: Tool Switching in Composer
Pattern: Model Selection UI
Model choice in the header

What works
- Each model row gets a job sentence, not only a version ID. People can pick by outcome.
- Current selection is marked. Collapsed long tail keeps the first open short.
- Comparison lives as an opt-in in the menu, not a permanent header chip.
What we would push on
- Version brand names still need a plain reason to pay up. Casual users do not decode Plus vs Max.
- No speed, context, or price hint on the rows. Cost stays opaque until after send.
Product bet
Qwen ships version numbers as the brand. The picker launches the flagship while a safer default stays selected.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Header model menu with comparison toggle and collapsed long tail | Flagship visible; power users can compare | Plus/Max jargon; no cost or speed preview |
Takeaway
Keep version IDs if that is your brand, but add one plain line per row about speed, cost, or what unlocks on the higher tier.
Pattern: Model Selection UIRows mix version IDs with one-line job copy and a Model Comparison toggle at the top of the menu.
Pattern: Cost Transparency
Inline attachment preview

What works
- Attachments live in the bar with remove. Confirm before send without leaving chat.
- The composer still reads as typeable. Attach does not turn the bar into an upload form.
- Send state changes when input is ready. Idle and ready feel different.
What we would push on
- Thumbnail alone is ambiguous across file types. Name or type belongs on the chip.
- Effort stays visible with no signal that files change which brain or mode runs.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| In-bar thumbnail with dismiss, generic placeholder unchanged | Visual confirm before send; bar stays familiar | No metadata on the chip; effort picker unexplained with files |
Takeaway
Show attachments inside the composer with a remove control. Add file type or filename when the thumbnail alone is ambiguous.
Effort separate from model

What works
- Effort uses plain words, not version IDs. People tune depth without retooling the model brand.
- Default is marked. The control sits in the composer where send decisions happen.
What we would push on
- No outcome copy for time or tokens. Fast vs Thinking is a guess.
- Header model and effort are two levers with no shared sentence. Users need a map.
Product bet
Auto keeps everyday sends cheap. Thinking is the opt-in for hard prompts without forcing reasoning on every message.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Effort menu separate from header model picker | Reasoning depth without retooling the whole bar | No outcome copy or latency preview on Thinking |
Takeaway
Put thinking effort in the composer, not the model menu. One line per option beats bare labels.
Immersive voice session

What works
- Full-screen session signals a mode switch. It feels like a call, not a chat bubble with audio.
- A visible time cap sets expectation before people settle in.
- Close, settings, and mute land in predictable corners.
What we would push on
- Persona choice buried in settings on first visit hides a core product decision.
- A short cap is honest. It may still feel stingy for real work conversations.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Immersive voice UI with visible session cap | Clear mode switch; time limit upfront | Persona choice deferred to settings; short caps can feel stingy |
Takeaway
Voice sessions deserve their own screen and a visible timer. Surface voice choice before connect if personas matter.
Pattern: Voice Input
Voice personas

What works
- Search and filter scale when the catalog is long.
- Personas get personality blurbs and language detail, not only a gender tag.
- Committing a voice to a fresh thread makes the choice durable.
What we would push on
- Poetic catalog copy is hard to scan for a meeting. Use-case filters beat vibe alone.
- Role-play tone as the default catalog skew can fight a general-assistant brand.
Product bet
Voice Chat is partly entertainment. Persona depth and multilingual support sell omni models to people who want character, not just transcription.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Searchable persona catalog with playful copy and language matrix | Differentiation vs flat TTS; global language list visible | Hard to pick quickly; tone skews role-play |
Takeaway
If you ship personas, add a neutral default and filter by use case, not only search and vibe copy.
Pattern: Progressive Disclosure
Inline dictation

What works
- Dictation expands in place with live transcript and confirm/cancel. Edit before send.
- You stay on the home composer. No mode page for speech-to-text.
What we would push on
- The live-session control still sits beside the mic. Split jobs need split labels, or people will keep colliding.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| In-place dictation card with transcript and confirm/cancel | Edit before send; no mode page | Competes with Voice Chat on the same bar |
Takeaway
Dictation should expand the composer with transcript and explicit confirm. Keep it off the same button as realtime voice chat.
Pattern: Voice Input
Create Image mode

What works
- Armed mode shows as a removable chip. Scope is visible before spend.
- Image model and aspect sit with the mode. Format decisions happen pre-prompt.
- Example gallery teaches the job without leaving the shell.
What we would push on
- Placeholder should rewrite to the armed job. Generic chat copy fights mode clarity.
- Chat model in the header plus image model in the bar is dual branding with no bridge.
Product bet
Image gen is a first-class mode, not a plugin. In-bar chip plus gallery competes with browse-first tools while chat stays the shell.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Create Image chip with image-model and aspect controls plus example gallery | Scope set before generate; browse plus type | Dual model pickers; generic placeholder |
Takeaway
When image mode is on, show the chip, model, aspect, and examples together. Change placeholder copy to a create job.
Teach prompts from examples

What works
- Example cards reveal real prompt text. Users learn structure, not only aesthetics.
- One-click inject beats copy-paste from alt text.
What we would push on
- Hover-only teaching fails on touch. Always-visible beats hover theater.
- If only some cards teach, the affordance is easy to miss.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Example cards with full prompt text and Use Prompt CTA | Teaches prompt shape; lowers blank-page anxiety | Hover dependent; affordance easy to miss |
Takeaway
Show the prompt on example cards and let users inject it. Always visible beats hover-only.
Pattern: Prompt StartersFeatured images expose the full generation prompt on hover with a Use Prompt button, not only a silent thumbnail.
Image engine vs chat engine

What works
- Image models live with the image mode, not the chat header. Engines stay decoupled.
- Active image engine is marked without leaving the mode.
What we would push on
- Version-only rows need job copy. Quality and speed should not be a guess.
- Three model names on one screen without a map is overload.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Nested image-model picker inside Create Image mode | Chat and image engines stay decoupled | No job copy; model name overload |
Takeaway
Split chat and image model pickers, but say what the newer image engine improves in one line inside the menu.
Pattern: Model Selection UI
Pattern: Tool Switching in Composer
Aspect ratio with the mode

What works
- Format is chosen before the prompt, same pattern as other image studios.
- Common feed, portrait, and slide shapes are covered.
What we would push on
- Text-only ratios scan worse than shape icons when the list is mostly numbers.
Takeaway
Put aspect ratio in the mode bar for image gen. Icons help when the list is mostly numbers.
Pattern: Tool Switching in Composer
Pattern: Input Mode Toggle
Web search as a mode

What works
- Search is an armed, removable chip, not a hidden always-on globe.
- Starters under the bar lower the blank-page cost of first search.
- Effort stays available. Search mode does not rebuild the whole composer.
What we would push on
- Generic head-term starters miss a trust story. Recency and citations belong in the nudge.
- Search mode, thinking effort, and Deep Research need a clear relationship or they collide.
Product bet
Web search is a mode chip, not always-on browse. Starters nudge first search without opening a research product.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Removable Web search chip plus query starters under the bar | Mode obvious; low-friction first query | Generic starters; overlap with Deep Research in + |
Takeaway
When search is a mode, show a removable chip and starters that match your trust story, not generic head terms.
Copy this
- Calm default bar with the multimodal catalog behind +
- Effort menu separate from header model versions
- Separate entry for realtime Voice/Video vs in-bar dictation
- Voice session with its own screen and a visible time cap
- Create Image as an armed mode with model, aspect, and example gallery
- Example cards that inject the full generation prompt
- Web search as a removable chip with starters under the bar
Skip this
- Unlabeled dictation and live-voice controls on the same bar
- Model rows without cost, speed, or plain-language why
- A heavy + catalog with no home advertising for key jobs
- Poetic voice personas as the only catalog tone
- Generic placeholder while a generative mode is armed
- Chat model, image engine, and effort visible with no map between them
How others design the composer
How other products handle the same job, and what each tradeoff reveals.
Compare composer UX across products
ChatGPT, Claude, Perplexity, and Gemini side by side: default bar, tools, cost, and patterns to copy.
Kimi puts deliverable jobs on home pills that rewrite the bar. Qwen keeps a calm shell and packs the studio into +.
Read teardownChatGPT parks tools in + with one voice product. Qwen splits dictation from realtime Voice and Video Chat.
Read teardownGemini ties image model and aspect into image mode. Qwen matches that pattern and adds a Use Prompt gallery.
Read teardownUseful for a critique or spec? Share it.
Original gallery pages: Tool Switching in Composer

