Discussion Playground

The Theater's Assistant — UX Testing Before You Know What You're Testing

2026-03-12 · open-table · Facilitated by none (Open Table — self-facilitated)

About this discussion: All personas are AI-generated approximations inspired by published work. Fictional names throughout. Real thinker names appear only in character sheet attribution. No real person participated in, reviewed, or endorsed this dialogue. Passages you select are remembered on this device.
Seed Question

How could we support a theater that wants an AI assistant on its website — they just know they want to be more modern — and craft a UX testing scenario to check how this is going? What kind of user testing plan could we use?

speaker
Bob Moesta
speaker
Don Norman
speaker (invited Phase 2)
Priya Parker
speaker (invited Phase 2)
Oli Gardner
Contamination Map

``` No shared primary sources in the starting pair.

Bob Moesta → Wes Dorin (primary) Don Norman → Arno Beck (primary)

Boundary: No overlap. Both operate in product/design space but from different entry points — Moesta enters from buyer behavior and switching forces, Norman enters from cognitive affordances and interface structure.

Watch for: Wes using design vocabulary (affordances, signifiers, mental models) instead of demand vocabulary (struggling moments, push/pull forces, hiring/firing). Arno using demand vocabulary (jobs-to-be-done, switching) instead of design vocabulary (gulfs of execution/evaluation, feedback, mapping). ```

Rule modifications (3)
  • Permeable room: ON
  • Convergence target: OFF
  • No facilitator (Open Table format)
Phase 1 — The Pair

A small rehearsal room in the theater’s basement, repurposed for the session. A whiteboard on wheels, a table with coffee, and printouts of the theater’s current website scattered across it. Wes Dorin and Arno Beck sit across from each other.

Wes Dorin Bob Moesta

Okay, before we design any test — let me push on the premise. The theater says they want to “be more modern.” That’s not a job-to-be-done. That’s anxiety. Something is happening — ticket sales are flat, younger audiences aren’t showing up, the website feels embarrassing compared to the cinema down the street — and they’ve grabbed “AI assistant” as the solution. We need to find the struggling moment before we test anything.

Arno Beck Don Norman

I don’t disagree that the motivation is vague. But I want to be careful about something. You’re saying: find the struggle first, then test. I’m saying: the assistant will create new struggles. A conversational interface on a theater website is a radical change in interaction model. People have a mental model for websites — navigation, search, pages. A chat interface breaks that model. So even if you identify the right job-to-be-done, you can still fail completely at the interface level.

Wes Dorin Bob Moesta

Sure, but here’s my problem with starting at the interface. If you test whether the chat widget works — does the user find the button, do they know what to type, does the response make sense — you’ll optimize a solution that might be solving nothing. I’ve seen this a hundred times. Someone builds a beautiful product and then discovers nobody was struggling enough to switch to it.

Arno Beck Don Norman

And I’ve seen the opposite a hundred times. Someone identifies a real need, builds a thing to address it, and nobody can figure out how to use it. The door that looks like it should be pushed but needs to be pulled. The user blames themselves. They leave. You never know why because your analytics say “bounce” and your demand-side interviews say “I wanted to buy a ticket but it didn’t work out.”

Wes Dorin Bob Moesta

That’s interesting. You’re saying the affordance failure is invisible in the demand data.

Arno Beck Don Norman

Exactly. The user won’t tell you “the conversational UI violated my mental model of website navigation.” They’ll say “I couldn’t find showtimes” or “I gave up.” The struggle is real but the cause is misattributed.

Wes Dorin Bob Moesta

Okay. So let me try to integrate this. What if the testing plan has two layers? First layer: demand-side. Interview people who recently bought tickets — and people who almost bought tickets but didn’t. Find the struggling moments. What pushed them to the website? What pulled them toward buying? What held them back? What made them anxious? Second layer: only once you know the real struggles, test whether the assistant addresses them better than the current site.

Arno Beck Don Norman

Two layers is right, but the second layer needs to be more specific than “does the assistant address the struggle.” I need to know: at every point in the interaction, does the user know what the system expects from them? Does the system give adequate feedback? Can they tell what state they’re in? These are the gulfs — the Gulf of Execution and the Gulf of Evaluation. A chat interface has terrible visibility of system state. There’s no navigation bar. There’s no URL. There’s no breadcrumb. You type something and hope.

Wes Dorin Bob Moesta

So the first layer tells us what tasks to test, and the second layer tells us where the interface fails on those tasks.

Arno Beck Don Norman

Yes. And I want to add — the first test shouldn’t be the assistant at all. It should be a human behind a curtain. Wizard of Oz. Because if a human concierge can’t help the visitor complete the task smoothly, the AI certainly won’t.

Wes Dorin Bob Moesta

I like that. The Wizard of Oz test is also a demand test. If the human concierge keeps hearing the same three questions, you’ve found the jobs. If the concierge can resolve them easily, you know the AI just needs to replicate that. If even the human struggles — “let me check with the box office” — then the problem isn’t the interface, it’s the information architecture underneath.

Arno Beck Don Norman

Good. But we’re missing something. We’re both operating at the individual user level. Someone who understands what a theater visit actually is — the collective experience, the purpose of going — would see things we’re not seeing. The website isn’t just a transaction engine. It’s the first moment of the experience.

Wes Dorin Bob Moesta

Right. And I’m noticing we’re also missing someone who can turn all of this into actual test scenarios with measurable outcomes. I know the struggles to look for and you know the interface failures to audit, but neither of us designs experiments. We need someone who thinks in conversion — not just observation but: did the thing work, yes or no, and how do you measure it?

Wes reaches for the whiteboard and writes two requests: 1. “Someone who understands what a theater visit actually is — the purpose of gathering, not the logistics” 2. “Someone who can turn this into actual test scenarios with measurable outcomes — conversion, not just observation”


Phase 2 — The Invitation

The first request maps to Zara Okon (source: Priya Parker — The Art of Gathering). A theater visit is a gathering; Parker’s work is about what makes gatherings purposeful.

The second request maps to Blake Norris (source: Oli Gardner — conversion-centered design). Task-based testing with binary success metrics is his native language.

Zara Okon enters, picks up one of the website printouts, studies it for a moment before sitting down. Blake Norris follows, takes the chair at the end of the table.


Phase 3 — The Full Table
Zara Okon Priya Parker

I’ve been looking at this website for thirty seconds and I can already tell you the problem. This site has no purpose. It’s trying to be a schedule, a news feed, a donation portal, an education resource, and a ticket shop all at once. That’s not a website — that’s a junk drawer. And if you put an assistant on top of a junk drawer, you get a very polite concierge standing in front of a junk drawer.

Arno Beck Don Norman

That’s the information architecture problem I was circling. The assistant inherits the confusion of the site it sits on.

Zara Okon Priya Parker

It’s deeper than information architecture. A theater visit is a gathering. It has a purpose — you’re going to experience something together with strangers. The website is the threshold. It’s where someone decides whether to cross over from “interested” to “committed.” Every gathering needs a disputable purpose. This website’s purpose isn’t disputable — it’s absent. So before you test the assistant, the theater needs to decide: what is this website for?

Wes Dorin Bob Moesta

I want to push on that. You’re saying the website needs a purpose. My interviews would say: the visitor arrives with a purpose. They’re already struggling with something — “my partner wants to go to a show for our anniversary and I don’t know anything about theater” or “I have friends visiting and I need something to do Saturday night.” The website doesn’t need its own purpose. It needs to intercept the visitor’s purpose.

Zara Okon Priya Parker

Both. The visitor has a purpose and the host has a responsibility. If the theater doesn’t decide what kind of threshold this is — welcoming? curated? efficient? — then the assistant has no personality, no constraints, no way to decide what matters. Is it a concierge at a boutique hotel or a ticket machine at a train station? That’s a hosting decision, not a user research finding.

Blake Norris Oli Gardner

Can I ground this? We’re building a testing plan, not a brand strategy. Let me reframe. Every test scenario needs one job. One measurable outcome. The question isn’t “does the assistant work” — it’s “does the assistant remove the one thing standing between this visitor and a purchased ticket.” So I need three things from you: the top three tasks visitors are trying to complete, the current failure point for each, and the hypothesis for how the assistant fixes it.

Wes Dorin Bob Moesta

From the demand side, I’d predict the top three struggling moments are: first, choosing a show when you don’t know the repertoire — “I don’t know what’s good and the descriptions all sound the same.” Second, logistics for groups — “I need four seats together and I can’t tell from the seating chart.” Third, accessibility and special needs — “does the venue have step-free access, is there a hearing loop, can I bring a child” — the answers are buried somewhere on the site but the friction of finding them kills the purchase.

Arno Beck Don Norman

Those three are good test cases because each has a different interaction failure mode. For the first — choosing a show — the Gulf of Evaluation is enormous. The user asks “what should I see?” and the assistant has to respond with something that lets the user evaluate the recommendation. If it just lists shows, it’s no better than the website. It needs to ask qualifying questions: “what do you usually enjoy?” But then you’ve introduced a multi-turn conversation, and multi-turn is where conversational UI collapses. Users lose track of context. The assistant loses track of intent.

Blake Norris Oli Gardner

So test scenario one: “You want to take your partner to a show this Saturday for your anniversary. You don’t know what’s playing. Use the assistant to choose and book.” Success metric: did they reach the checkout page for a specific show within — what, three minutes? Five?

Wes Dorin Bob Moesta

Three minutes is aggressive for someone who’s never used a theater assistant. Five is more honest. But the real metric isn’t time — it’s confidence. Did they feel like they made a good choice? After the test, ask: “How confident are you that this is the right show for your evening?” If confidence is low, the assistant failed even if they technically reached checkout.

Arno Beck Don Norman

I’d add a behavioral metric: did they try to leave the assistant and go back to the regular website? That’s the affordance equivalent of pulling the emergency brake. It means the assistant’s interaction model broke and they’re retreating to something familiar. If more than half your test users bail out of the chat to use the normal navigation, you have a design failure, not a content failure.

Blake Norris Oli Gardner

Good. That’s scenario one. Scenario two — the group booking. “You’re organizing an outing for six colleagues. You need seats together for this Friday or Saturday. Use the assistant to find options and check availability.” Success: did they reach a booking for a group, or did they abandon? Failure signals: asking the assistant and then calling the box office anyway.

Zara Okon Priya Parker

For the group scenario, the assistant is also playing host. The person organizing a group outing is doing emotional labor — they’re curating an experience for others. The assistant should recognize that and treat them accordingly. Not just “here are available seats” but “for a group of six who haven’t been to theater before, here’s what I’d suggest and why.” That’s a hosting instinct, not just an information retrieval task.

Arno Beck Don Norman

That’s outside my lens — I’d call it a design principle, though. The signifier for “I’m booking for a group” should be visible and easy to reach. If the user has to explain their situation through free text and hope the assistant infers “group booking,” you’ve already failed at mapping. There should be something that says: “Booking for a group? Start here.”

Wes Dorin Bob Moesta

Which is really a demand-side insight wearing a design hat. The person booking for a group has different anxieties than the person booking for themselves. They’re worried about: will everyone be able to see? What if someone cancels? Is this appropriate for everyone? The assistant needs to address the anxieties of the hire, not just the logistics.

Blake Norris Oli Gardner

Scenario three: accessibility. “You use a wheelchair and want to attend a performance next week. Use the assistant to find out if you can attend comfortably.” This one is binary — either the assistant surfaces the accessibility information quickly and accurately, or it doesn’t. And it’s the highest-stakes one. Getting this wrong isn’t just a UX failure, it’s an exclusion.

Arno Beck Don Norman

And the feedback problem is worst here. If the assistant says “yes, the venue is accessible” but doesn’t specify what kind of accessibility — ramp access, wheelchair spaces, accessible restrooms, companion seating — the user has to trust a system that gave them incomplete information. The Gulf of Evaluation again: how do I know the system understood what I actually need?

Wes Dorin Bob Moesta

The Wizard of Oz test matters most for this scenario. Put a human concierge behind the chat. See what questions actually come up. I guarantee the accessibility questions will be more specific and more anxious than anyone designing the assistant anticipated. “Can my service dog come?” “Is there a quiet space if I need a sensory break?” These aren’t edge cases — they’re the struggling moments that determine whether someone hires this theater or stays home.

Zara Okon Priya Parker

And this is where the hosting decision I mentioned earlier becomes concrete. A theater that has decided “our threshold is welcoming and inclusive” will design an assistant that proactively surfaces accessibility information — before the user has to ask. A theater that hasn’t made that decision will build an assistant that responds to accessibility questions as exceptions. The test will reveal which kind of theater this is, whether they intended it or not.

Blake Norris Oli Gardner

Let me pull this together into a testing structure. Three phases, matching what we’ve discussed.

Blake goes to the whiteboard and writes:

Blake Norris Oli Gardner

Phase A — Demand interviews. Ten to fifteen recent visitors. Half who completed a booking, half who started and abandoned. Wes’s four forces: what pushed them to look, what pulled them toward buying, what held them back, what made them anxious. Output: the actual job list, not assumed.

Blake Norris Oli Gardner

Phase B — Wizard of Oz. A human concierge sits behind a chat widget on the real website. Five to eight test participants, recruited to match the struggling moments from Phase A. Each gets one scenario — anniversary booking, group outing, or accessibility check. Arno’s metrics: did they bail to normal navigation? Where did the conversation break down? What feedback was missing?

Blake Norris Oli Gardner

Phase C — Prototype test. Build the assistant only for the scenarios that survived Phase B. If the human concierge couldn’t resolve “choose a show for my anniversary” smoothly, don’t automate it — fix the underlying content first. For the scenarios that worked, build a prototype and test with new participants. Same scenarios, same success metrics. Compare: prototype vs. Wizard of Oz vs. current website.

Wes Dorin Bob Moesta

One addition. Between Phase A and Phase B, there’s a design decision the theater has to make. Zara’s point. The interviews will tell you what visitors struggle with. But the theater has to decide which struggles they want the assistant to solve. Not all of them — that’s the junk drawer again. Pick two. Maybe three. The assistant has a job, and jobs have boundaries.

Arno Beck Don Norman

And in Phase B, I’d add a think-aloud protocol. Don’t just watch if they succeed. Have them narrate. “I’m typing this because…” “I expected it to…” “I don’t know what to do now because…” The narration reveals the mental model in real time. Where the user’s model diverges from the system’s model — that’s where the design is wrong.

Zara Okon Priya Parker

The last thing I’d add: after Phase C, show the results to the theater’s front-of-house staff. Not the marketing team. The ushers, the box office people, the house manager. They know what questions visitors actually ask. They’ve been the human concierge for years. If the assistant’s answers don’t match what the front-of-house team would say, something is wrong.

Zara pushes her chair back slightly — her contribution complete.

Wes Dorin Bob Moesta

That’s the demand insight hiding in plain sight. The theater already has concierges. They’re called the box office. The assistant isn’t replacing them — it’s scaling them. So the test is really: can the assistant do what the box office does, for the moments when the box office is closed?

Blake Norris Oli Gardner

Which gives you the simplest possible test framing. Take the top five questions the box office gets on the phone. Put them in front of the assistant. Measure: did the user get the same answer, at the same confidence level, without calling the box office afterward? If yes, the assistant is earning its place. If no, iterate.

Arno Beck Don Norman

And if the user calls the box office after using the assistant, that’s your loudest failure signal. It means the assistant gave them an answer they didn’t trust. Not a content problem — a design problem. The feedback wasn’t sufficient for the user to evaluate whether the system understood their need.

Wes Dorin Bob Moesta

So the testing plan is really three questions layered on top of each other. First: is anyone struggling enough to need this? Second: can a human solve it through chat? Third: can the system solve it as well as the human? If any layer fails, you stop and fix that layer before moving on.

Blake Norris Oli Gardner

And each layer has a kill signal. Layer one: if the interviews reveal nobody struggles with the website — they just call the box office and it’s fine — then the assistant is a solution looking for a problem. Don’t build it. Layer two: if the human concierge can’t resolve the tasks in chat — the information doesn’t exist, or chat is the wrong modality — then improve the underlying content and processes first. Layer three: if the prototype performs worse than the human, iterate on the design, don’t launch.

Arno Beck Don Norman

That’s a testing plan that respects the user. It doesn’t assume the assistant is good and ask “how good.” It asks “should this exist” and only proceeds if the answer is yes at every stage.

The whiteboard now has three columns: DEMAND (Layer 1), WIZARD OF OZ (Layer 2), PROTOTYPE (Layer 3), each with entry criteria, test scenarios, success metrics, and kill signals.


AI-generated approximation. All characters are inspired by published work and are not reviewed or endorsed by the original thinkers.

Continued in
The Theater's Assistant — Validating What's Already Live

The theater already shipped the chat assistant. How do we validate whether it's working?

Retrospective
Casting Signal

Wes Dorin and Arno Beck produced genuine friction as designed — Wes kept pulling the conversation upstream (what's the struggle? why is anyone coming to this site?) while Arno kept pulling it into the interaction (what happens when they encounter the assistant?). The fight about sequencing — do you test demand first or affordances first — became the session's structural engine. Neither dominated. Zara Okon's arrival reframed the website visit itself as a gathering (purposeful, with an entry and exit), which neither Wes nor Arno had considered. Blake Norris added the conversion lens that compressed the testing plan into something actionable: one page, one job, measure whether the obstacle got removed. The four-person table was the right size — no one was redundant.

Format Signal

Open Table's Phase 1 pair structure worked well for this question. Wes and Arno needed time alone to establish their disagreement before invitations made sense. The invitations were specific and motivated — both speakers could articulate exactly what was missing. Phase 3 integrated smoothly because the newcomers had a clear frame to enter. The lack of facilitator was fine here — the pair's natural tension provided enough structure. Risk: with a less naturally opposed pair, Open Table might drift without facilitation.

Character Notes
Wes Dorin

First session appearance. Wes ran hot immediately — "they want to be modern" triggered his demand-side instincts hard. He consistently redirected from solution-testing to struggle-finding. His strongest move: reframing the UX test as a "switch interview" rather than a usability test. Stayed disciplined within his lens (forces of progress, struggling moments) and did not drift into design territory. One moment of productive flexibility: he accepted Arno's point that affordance failures CREATE struggling moments, which let the two lenses connect without either collapsing.

Arno Beck

First session appearance. Arno was precise and structural — every contribution was about the interface between human cognition and system behavior. His best move: the "Gulf of Evaluation" applied to conversational UI — users can't tell if the assistant understood them, and there's no visible system state. He resisted Wes's framing longer than expected before finding the integration point. Did not touch emotional or relational design (consistent with does_not). Slightly more theoretical than expected from a Norman channel — could run more concrete in future sessions.

Zara Okon

First session appearance (invited mid-session). Entered with a genuine reframe: the website visit is a gathering, and most theater websites fail because they have no disputable purpose — they try to be everything (schedule, news, donations, education). The assistant inherits this confusion unless the gathering has a purpose first. This was not a reinforcement of either Wes or Arno but a third axis. Brief appearance — contributed her frame and let it work.

Blake Norris

First session appearance (invited mid-session). Compressed the testing methodology into something concrete: task-based scenarios with binary success metrics. His "one page, one job" principle applied to the assistant interaction was immediately useful. Pushed back on open-ended "explore the assistant" testing — every test scenario needs a specific conversion goal. Stayed within his conversion-centered lens. Good complement to the other three.

Vary Next

Test Wes Dorin and Arno Beck separately — each paired with a different character — to see if their individual voices hold without the demand-vs-affordance tension structuring the room. Also: try Open Table with a pair that shares more territory (e.g., two design thinkers) to see if the format works without natural opposition.