All work
[ Concept study ] [ Conversation design ] [ Multimodal ] [ 2026 ]

Cue

A voice-first music assistant that knows when to talk and when to show. A self-initiated study in conversation design across a voice-only speaker, a smart display, a phone and the car.

Role
Conversation & multimodal designer
Type
Self-initiated concept
Surfaces
Voice-only speaker · Smart display · Phone · Car
Deliverables
Principles, persona, sample dialogs, dialog flow, repair strategy, multimodal spec, working prototype
Talk to the prototype

Independent concept, not affiliated with any assistant platform. Catalogue is illustrative.

01 · [ The problem ]

Music is the easy request that's hard to get right

"Play something" is the most common thing people say to a voice assistant, and the one with the most ways to disappoint.

A music request is short, casual and full of assumptions. "Play Blonde" could be an album, a playlist or a song. "Play the new Tyler album" needs the assistant to know what "new" means today. "Add this to my gym playlist" only works if the assistant remembers what "this" is. And the right answer sounds different on a kitchen speaker than it looks on a screen in a car.

I picked four problems to design against, and built a prototype to test whether the rules hold up when someone actually types or says things to it.

  1. A

    Ambiguity

    One name, several entities. Resolve with the least friction the surface allows.

  2. B

    Context

    "This", "that", "again", "the next one": anaphora across turns and across surfaces.

  3. C

    Repair

    Noise, mishearing, out-of-catalogue. Recover in one turn, escalate gracefully, always offer an exit.

  4. D

    Modality

    Same intent, different surface. Decide what is said, what is shown, and what is both.

02 · [ Principles ]

Six rules Cue never breaks

  1. 01

    Be cooperative, then be quiet.

    Answer the question that was asked, at the level of detail it was asked at. Confirmation of a low-risk action is the action itself. Music starting is the confirmation.

  2. 02

    Confirm in proportion to the risk.

    Implicit for play/skip. Explicit for anything that changes the user's library, spends money, or is heard at low confidence.

  3. 03

    Never read a list a screen could show.

    Voice-only reads at most three options. With a screen, voice asks the question and the screen carries the choices.

  4. 04

    One breath.

    Every spoken response fits in a single breath, roughly twelve words. If it needs more, it needs a screen or a follow-up turn.

  5. 05

    Always a way out.

    Two failed turns and Cue stops guessing: it offers concrete things to say and a way to stop. It never loops.

  6. 06

    One persona, every surface.

    Shorter on a speaker, quieter in a car, richer on a display, but the same voice. Consistency is what makes the assistant feel like one thing.

03 · [ Persona & voice ]

The friend behind the record-shop counter

Cue is knowledgeable without showing off, warm without being chirpy, and comfortable with silence. It speaks in short, complete sentences and stops. It never apologises twice, never uses exclamation marks, and never says "great choice".

TraitSounds likeNever
Brief"Blonde, by Frank Ocean.""Sure! Now playing the album Blonde by the artist Frank Ocean."
Warm, dry"Here's Kitchen Radio. No skips needed.""Awesome!! Let's get cooking! 🎵"
Honest about limits"I don't have that one. Want SZA radio instead?""Sorry, I didn't understand. Please try again."
Proportionate"Added Snooze to Gym.""Are you sure you want to add this song to the playlist?"
04 · [ Sample dialogs ]

Four conversations

Each script is written for a voice-only speaker first, then annotated with what changes when a screen is present. Voice is what Cue says. Screen is what a display shows instead of, or alongside, speech.

D1 · Happy path

Activity request

User

Play something for cooking.

Cue

Voice Here's Kitchen Radio. Upbeat, no skips needed.

Screen Now-playing card with the playlist name and a one-line reason: "Because you cook on Sundays." Voice shortens to "Kitchen Radio."

Why: the reason line builds trust in a recommendation without asking the user to evaluate it. On voice it's one clause; on screen it's a caption.

D2 · Disambiguation

One name, two things

User

Play Blonde.

Cue

Voice I've got the album Blonde by Frank Ocean, and your playlist Blonde Hour. Which one?

Screen Two tappable options with artwork. Voice says only: "Which Blonde?"

User

The album.

Cue

Voice Blonde, by Frank Ocean.

Why: options are ordered by likelihood (library item first when the name is an exact match), capped at three, and each is described by the attribute that separates it: type, not title.

D3 · Repair

Noise, then low confidence

User

(muffled, blender running)

Cue

Voice Sorry, I missed that. What should I play?

User

Play the new Tyler album.

ASR confidence 0.62
Cue

Voice Chromakopia, by Tyler, the Creator?

Screen Album card with a single "Play" button; a wrong guess is one tap from fixed, so voice doesn't wait for a yes.

User

Yeah.

Cue

Voice Playing.

Why: the first repair re-asks the open question rather than saying "I didn't understand". The second is a confirmation phrased as the answer. If right, the user just says yes; if wrong, they've already heard what Cue thought it heard.

D4 · Context

Carrying "this" across turns

User

Play SZA.

Cue

Voice SZA radio, starting with Kill Bill.

User

Skip.

Cue

Earcon Short tick, no speech. Voice-only adds "Next: Snooze" only if the user has skipped twice in a row.

User

Add this to my gym playlist.

Cue

Voice Added Snooze to Gym.

User

What's this?

Cue

Voice Snooze, by SZA, from SOS.

Screen Already visible. Voice answers anyway, because the user asked.

Why: "this" always resolves to the currently playing track, and the confirmation names it, so a wrong resolution is caught immediately. Skips don't get speech because the music changing is the feedback.

05 · [ Multimodal response spec ]

Same intent, four surfaces

The spec is the contract between the dialog and every surface. Voice is designed to be sufficient on its own; screens subtract speech and add choice.

IntentVoice-only speakerSmart displayPhoneCar
Play (resolved) "Blonde, by Frank Ocean." Music starts. "Blonde." Now-playing card: art, title, queue. No speech. Now-playing sheet + haptic. "Blonde, by Frank Ocean." Large title only; no lists.
Disambiguate Reads ≤3 options by distinguishing attribute. "Which one?" "Which Blonde?" Tappable options with art. Bottom sheet of options; voice optional. Reads 2 options max, defaults to the likeliest after 4 s of silence.
Skip Earcon. Speech only on 2nd consecutive skip. Earcon. Card crossfades. Haptic. Card crossfades. Earcon. New title announced.
Low confidence Confirm as the answer: "Chromakopia?" Card + single Play button; plays on yes or tap. Inline "Did you mean…" chip. Confirm as the answer; never a list.
No match ×2 Three example phrases, then "or say stop." "Try one of these" + chips: artist, album, mood. Search field pre-focused with the heard text. "I couldn't find that. Try again when you're parked, or say a mood."
Add to playlist "Added Snooze to Gym." Explicit, because it's a library change. Same speech + toast with undo. Toast with undo; no speech. Same speech; no toast.
06 · [ Dialog flow ]

How a request travels

Cue dialog flow An utterance goes to NLU for intent and slots, then a confidence gate: above 0.8 act, between 0.5 and 0.8 confirm, below 0.5 repair. Acting resolves entities: one match plays, several disambiguate, none triggers a no-match repair. After two failed repairs, Cue offers examples and an exit. Playing stores context for follow-ups like skip, add and what's this. UTTERANCE"play blonde" NLUintent + slots CONF? ≥ .80 · ACT .50–.80 · CONFIRM < .50 · REPAIR RESOLVE ENTITY 1 → PLAY >1 → DISAMBIGUATE 0 → NO MATCH RE-ASK (max 2) EXAMPLES + "OR SAY STOP" CONTEXT STORE skip · add · what's this highmidlow resolved 2nd fail
Confidence gate → entity resolution → repair with a hard cap. Every branch ends in music or an exit.
07 · [ Repair strategy ]

When it goes wrong

ConditionFirst attemptSecond attemptThen
No input"What should I play?""You can say an artist, an album, or a mood."Close the session silently. Never a third prompt.
No match"I don't have that one. Want [nearest artist] instead?"Three example phrases + "or say stop."Screen: search pre-filled with the heard text.
Low confidenceConfirm as the answer: "Chromakopia?"Offer the top two alternatives.Fall through to no-match.
Ambiguous≤3 options by distinguishing attribute.Default to the likeliest and say so: "Going with the album. Say 'the playlist' to switch."Play. A wrong guess is one turn from fixed.
Out of scope"I can't do that here, but I can play, skip, or save music."NoneNever pretend to have done it.
Playback error"That one won't play right now. Try the next track?"Skip automatically on yes.Log for content team; don't blame the user.
08 · [ Prototype ]

Talk to Cue

A working simulator of the rules above. Type or say a request, switch the surface, and watch the same intent produce different responses. The trace shows what the assistant heard, what it decided, and how confident it was.

Idle

NLU trace
...

What the prototype proves

The same intent yields different responses per surface: on the speaker, disambiguation is read aloud (max three); on the display, voice asks the question and the screen carries the options. Typos trigger a confirm-as-the-answer at mid confidence; unknown artists trigger the two-step repair with a hard exit; "this" and "skip" resolve against a context store.

It's rule-based on purpose. The point is the dialog policy, not the model. Swap the NLU for a real one and the policy holds.

09 · [ Accessibility & edge cases ]

The unglamorous parts

10 · [ How I'd validate it ]

Prove it with people

This is a concept, so the honest next step is testing, not polish. I'd run Wizard-of-Oz sessions first: a facilitator plays Cue from the script while participants cook, drive (simulated) and sit with a display, then move the same tasks onto the prototype.

Tasks: play a specific album with an ambiguous name; recover from a mishearing; save the current track; find music for an activity. Each measured on the same five numbers.

Task success
Did music the user wanted start? Target ≥ 90% within 2 turns.
Turns to success
Median turns per task. Target ≤ 2 on voice-only, 1 on display.
Repair exit rate
Share of sessions that hit the two-strike exit. Target < 5%.
Confirm accuracy
How often the mid-confidence guess was right. Tunes the gate thresholds.
Perceived effort
Single Ease Question after each task, compared across surfaces.
11 · [ Reflection ]

Why I built this

I DJ. Reading a room and deciding the next track is the same job as designing a good turn: know what they meant, don't over-explain, keep the energy moving.

Designing for a screen taught me to show everything; designing for voice taught me to say almost nothing. The hardest part of this study was the multimodal spec: deciding, intent by intent, what a screen should take away from speech rather than add to it.

What I'd do next: replace the rule-based NLU with a real model and see which principles survive contact with actual confidence scores; then test the car surface properly, because the four-second default is a guess that needs a driving simulator to check.