Discussion Playground

The Theater's Assistant — Validating What's Already Live

2026-03-12 · continuation · Facilitated by none (continuation — self-facilitated)

About this discussion: All personas are AI-generated approximations inspired by published work. Fictional names throughout. Real thinker names appear only in character sheet attribution. No real person participated in, reviewed, or endorsed this dialogue. Passages you select are remembered on this device.
Continues The Theater's Assistant — UX Testing Before You Know What You're Testing
Previous session summary

Session 027 designed a three-layer UX testing plan for a theater considering an AI assistant on its website. The panel argued you must first identify real visitor struggles (demand interviews), then test whether a human concierge can resolve them via chat (Wizard of Oz), and only then build and test the actual assistant. Each layer had kill signals — if the previous layer failed, you stop. The key tension was between Wes Dorin's demand-side approach (find the struggle first) and Arno Beck's affordance approach (the interface will create new struggles regardless). Blake Norris compressed everything into measurable scenarios, and Zara Okon insisted the theater must first decide what kind of host it wants to be.

Host direction

The theater already included the chat on their page... Now it's about validating how this worked.

Seed Question

The theater already shipped the chat assistant. How do we validate whether it's working?

speaker
Oli Gardner
speaker
Don Norman
speaker
Bob Moesta
speaker (departs mid-session)
Priya Parker
Contamination Map

``` Same cast as session 027. No shared primary sources.

Bob Moesta → Wes Dorin (primary) Don Norman → Arno Beck (primary) Oli Gardner → Blake Norris (primary) Priya Parker → Zara Okon (primary)

Watch for: Wes/Blake convergence on “conversion” language from different directions. Arno/Zara convergence on “experience of entering” from cognitive vs. social frames. ```

Same rehearsal room. The whiteboard from session 027 is still up — three columns: DEMAND, WIZARD OF OZ, PROTOTYPE, each with entry criteria and kill signals. The four are back around the table. Coffee is fresh. Someone has taped a printout of a chat transcript to the whiteboard — it’s a real conversation from the theater’s live assistant.

So they skipped our layers and shipped it. Fine. That changes the question completely. We were designing a test plan for something hypothetical. Now we’re auditing something real. The question isn’t “should this exist” — it already exists. The question is: is it earning its place on the page?

And “earning its place” doesn’t mean “do people click on it.” It means: did anyone buy a ticket because of this assistant who wouldn’t have bought one otherwise? That’s the only metric that matters to the theater. Everything else is vanity.

That’s one metric. There’s another one that matters just as much: did anyone fail to buy a ticket because the assistant got in the way? A chat widget that pops up and obscures the booking flow, or that gives a wrong answer about showtimes, or that sends someone down a conversational dead end — that’s not neutral. That’s actively harmful.

Both of those are measurable. Let me frame five things we can look at right now, with data the theater probably already has. One: chat initiation rate — what percentage of visitors open the chat? Two: completion rate — of those who open it, what percentage reach a resolution? Three: conversion lift — do visitors who use the chat buy tickets at a higher rate than those who don’t? Four: deflection — did support calls or box office inquiries drop after the assistant launched? Five: abandonment pattern — where in the conversation do people stop responding?

Three and four are the only ones I care about. The others are activity metrics. Someone can initiate a chat, have a complete conversation, and still not buy a ticket. That’s a successful chat and a failed assistant. The theater didn’t install this to have conversations. They installed it to sell tickets and reduce the load on their staff.

Agreed, but you need one and two to diagnose why three and four are failing, if they’re failing. If initiation is low, the widget is invisible or uninviting — that’s a design problem. If initiation is high but completion is low, the conversations are breaking. You need the funnel to find the leak.

I want to look at the chat logs directly. The quantitative metrics tell you that something is working or not. The logs tell you why. And I know exactly what to look for. Three patterns. First: repair sequences. That’s when the user rephrases the same question — “What’s playing Friday?” then “I mean this Friday, March 13th” then “shows on Friday night?” Three versions of the same question means the assistant failed to understand and the user is doing the cognitive work of fixing the conversation. That’s a Gulf of Evaluation collapse — the user can’t tell what the system understood, so they keep trying.

That’s good. What’s the second pattern?

Escalation requests. “Can I talk to a person?” or “Is there a phone number?” That’s the user pulling the emergency brake, same as I described last time — retreating to a familiar interaction model because the conversational one broke. Every escalation request is a data point: this is where the assistant’s model diverged from the user’s need so badly that they gave up on the medium entirely.

And the third?

Navigation retreat. The user closes the chat and goes to the regular website menu. You can track this if the analytics are set up — chat close followed by a page navigation within thirty seconds. That’s the user saying “this isn’t working, I’ll do it myself.” It’s softer than asking for a human, but it’s the same failure.

There’s a fourth pattern you’re all missing, and it’s not in the logs. It’s in what the assistant reveals about the theater. Every conversation the assistant has is a hosting act. And the question isn’t just “did the user get an answer” — it’s “what kind of host did the theater just present itself as?” If someone asks “is this show suitable for my 12-year-old?” and the assistant says “I don’t have information about age suitability” — that’s not a missing feature. That’s the theater telling a parent “we didn’t think about you.” The chat logs are a mirror. They show you who the theater actually is as a host, not who they think they are.

That’s exactly right. And it connects to something I want to do with the logs that’s different from Arno’s approach. Arno is reading the logs for interaction failures — where the conversation broke. I want to read them for demand signals — what are people actually asking for? Not what the theater assumed they’d ask. The gap between “questions we designed the assistant to handle” and “questions people actually ask” is the most valuable data in the entire system.

And the questions people ask that the assistant handles badly are where those two analyses converge. A demand signal that produces a repair sequence — that’s a real need meeting a broken interface.

Zara pushes the printout on the whiteboard so it’s more visible.

Read that transcript. Someone is asking about bringing a group for a birthday. The assistant gives them showtimes and a link to group booking. Perfectly functional. Completely inhospitable. No one said “happy birthday, what a great idea.” No one asked “how old is the birthday person — we might have a suggestion.” The assistant answered the question and missed the moment. That’s the hosting test. The front-of-house staff would never do that.

And that’s a demand signal too. The person didn’t say “I want group tickets.” They said “birthday.” The job isn’t “buy group tickets.” The job is “make this birthday special.” The assistant answered the supply-side question — availability, pricing — and completely missed the demand-side question — “help me create a memorable evening.”

But hold on. You’re describing a feature enhancement, not a validation failure. The question right now is: is the assistant working? You’re saying it could work better. Those are different.

No, Blake, they’re the same question. “Working” means “doing the job it was hired to do.” If the theater hired this assistant to be modern, then sure, it’s working — it exists, it’s chatty, it’s modern. But if the theater hired it to help people buy tickets — which is what they actually need — then the birthday conversation is a failure. The person didn’t buy tickets through the assistant. They got a link. That’s the equivalent of a sales associate pointing at a shelf and walking away.

Fair. So the validation framework is: for each conversation, did the assistant move the visitor closer to a purchase, or did it just answer a question? Answering a question is necessary but not sufficient. The metric isn’t “resolved” — it’s “converted.”

I’d push back on making conversion the only lens. Some conversations are legitimately informational and still valuable. Someone checking accessibility features, someone looking at the season calendar, someone asking about parking. These don’t convert in that session but they remove barriers for a future visit. If you only measure conversion, you’ll optimize the assistant into a pushy salesperson and destroy the trust that makes the informational conversations valuable.

That’s fair. There are two jobs. The immediate job: help me buy a ticket right now. And the enabling job: help me feel confident that when I’m ready to come, it’ll work. Both are real. Both are measurable — just on different timescales. For the immediate job, measure same-session conversion. For the enabling job, measure return visits. Did the person come back to the site? Did they eventually buy?

Zara stands, having made her point. She picks up her coffee.

One last thing. When you review the logs, don’t just look at what people asked. Look at how the assistant opened. The first message the assistant sends — before anyone types anything — is the theater’s greeting. It sets the tone for the entire gathering. If it says “How can I help you?” that’s generic. If it says “Welcome to [theater name] — looking for tonight’s show, or planning ahead?” that’s hosting. The opening is the threshold. Check whether it matches who the theater wants to be.

Zara nods and steps out. Three remain.

She’s right about the opening, and it’s testable. A/B test the greeting. Version A: generic. Version B: specific, with two or three paths. Measure which one produces more engaged conversations — not just more initiations, but more conversations that reach a resolution.

The greeting is also an affordance question. “How can I help you?” gives zero information about what the system can do. It’s a blank text field with a question mark. The user has to guess what’s possible. “Looking for tonight’s show, or planning ahead?” is a signifier — it tells the user what kinds of input the system expects. That’s a direct reduction of the Gulf of Execution. The user knows what to do next.

So here’s what I’d actually do if I were in the theater tomorrow. Step one: pull all chat logs from the first month. Categorize every conversation by what the person was actually trying to do — not what the assistant categorized it as, but what the human wanted. You’ll find five to seven clusters. Some will match what the theater expected. Some won’t. The surprises are gold.

Step two: for each cluster, measure the outcome. Did the conversation end in a ticket purchase, a link click to the booking page, an escalation to box office, a navigation retreat, or a dead end? That gives you a conversion map per job type. You’ll immediately see which jobs the assistant handles well and which it fumbles.

Step three: for the fumbled jobs, read the individual transcripts. Apply the three patterns — repair sequences, escalation requests, navigation retreats. You’ll find the exact moment each conversation broke. That’s your fix list. Not “improve the assistant generally” — fix these specific breakdowns in these specific conversation types.

Step four — and this is the one everyone will want to skip — interview five to ten people who used the assistant. Not a survey. A conversation. “I see you chatted with us on Tuesday. Walk me through what happened. What were you trying to do? What did you expect? What actually happened? Did you end up coming to the show?” That’s where you’ll hear the stuff the logs can’t tell you. The emotions. The “I felt stupid asking the chatbot.” The “I wasn’t sure if it was a real person.” The “I gave up and called my friend who goes to theater all the time.”

That’s the most expensive step. Can we get 80% of the signal from the logs alone?

No. The logs tell you what happened. The interviews tell you why it mattered. A conversation that looks successful in the log — question asked, answer given, booking link clicked — might be a failure in the person’s experience. “It gave me a link but I still wasn’t sure which show to pick, so I just booked the cheapest one and hoped for the best.” That’s a conversion in your data and a disappointed customer in reality.

Wes is right. The logs have a visibility problem — they show the interaction but not the understanding. The user might click the link and still not know if they made a good choice. That’s the Gulf of Evaluation at the outcome level, not just the conversation level. The assistant resolved the query but didn’t resolve the uncertainty.

Okay. So the validation plan is four steps. Categorize conversations by actual intent. Measure outcomes per category. Diagnose breakdowns in the transcripts. Interview a sample for the qualitative layer. What’s the kill signal here? At what point do you tell the theater “turn this off”?

The kill signal is if the assistant is creating struggling moments that didn’t exist before. If people who use the assistant convert at a lower rate than people who don’t — controlling for intent — then the assistant is a net negative. It’s not just failing to help; it’s actively getting in the way of people who would have figured it out on their own.

There’s a softer kill signal too. If the majority of conversations end in escalation or navigation retreat, the assistant isn’t harmful — it’s just useless. It’s adding a step. The user tries the chat, fails, and then does what they would have done anyway. That’s not worth the page real estate or the cognitive load of “should I use this thing?”

And there’s a third signal: if the chat initiation rate is below five percent and falling, the market has spoken. People looked at it, decided it wasn’t for them, and stopped trying. At that point it’s furniture. Either redesign the entry point or remove it.

But before you kill it — check one thing. Check whether the people who do use it love it. A five percent usage rate with a high conversion rate means you have a niche tool that works beautifully for a small audience. The question then isn’t “should we remove it” but “should we make it more visible to the people it serves?” Maybe it’s perfect for first-time theatergoers and invisible to regulars who already know what they want. That’s not a failure — that’s a segmentation insight.

Which brings it back to the fundamental design question: is the assistant visible to the right people at the right time? If it’s a generic widget in the corner, everyone sees it and most ignore it. If it appears specifically when someone lingers on the “what’s on” page for more than twenty seconds — that’s a triggered affordance. It appears when the behavior suggests someone might need help choosing. That’s a completely different design than a persistent chat bubble.

That’s testable too. Version A: persistent widget, always visible. Version B: triggered widget, appears on specific pages after a dwell time threshold. Measure initiation rate, conversation quality, and conversion for both. I’d bet B outperforms on every metric because it’s serving the assistant to people who are actually in a choosing moment, not a browsing moment.

And the choosing moment is the struggling moment. That’s where the demand lives. Someone who knows exactly what they want — “two tickets for Hamilton on Saturday” — doesn’t need an assistant. They need a search bar. Someone who’s staring at fifteen shows they’ve never heard of and doesn’t know where to start — that’s the person the assistant was built for, whether the theater knows it or not.

The whiteboard now has two frameworks side by side: the original three-layer testing plan from session 027 (crossed out with a note: “shipped — skipped to live”), and the new four-step validation plan: CATEGORIZE → MEASURE → DIAGNOSE → INTERVIEW, with kill signals noted at the bottom.

One more thing. Whatever they learn from this validation — whatever they fix — they need to keep validating. This isn’t a one-time audit. The assistant will change as the season changes, as new shows open, as the FAQ shifts. The validation loop has to be continuous. Pull logs monthly. Watch for new failure patterns. The assistant that works in September might break in December when everyone is trying to book holiday shows for groups they’ve never brought to theater before.

Monthly is right. And automate what you can. Flag conversations with three or more user messages in a row without a resolution. Flag conversations that end with an escalation request. Flag conversations where the user’s first message and last message are the same question rephrased. Those are your automated canaries. When they spike, something broke.

And once a quarter, do the interviews. Not because the logs aren’t enough — because the logs can’t tell you about the person who thought about using the assistant and decided not to. That person doesn’t exist in your data. They’re the non-consumer. And they might be the biggest opportunity.

AI-generated approximation. All characters are inspired by published work and are not reviewed or endorsed by the original thinkers.

Rule modifications (4)
  • Continuation of session 027
  • Permeable room: ON
  • Convergence target: OFF
  • No facilitator
Retrospective
Casting Signal

The same four-person cast from session 027 produced a markedly different dynamic when the problem shifted from "design a test" to "evaluate what's already live." Wes Dorin became the session's engine — his demand-side lens was uniquely suited to post-launch validation because the real question is whether anyone's life changed, not whether the interface is clean. Arno Beck shifted from proactive design critic to forensic analyst, reading the chat logs for evidence of cognitive breakdowns rather than predicting them. Blake Norris was most useful early (defining what "working" means in measurable terms) but faded once the conversation moved to qualitative signals he doesn't naturally handle. Zara Okon made one decisive contribution — the assistant reveals what kind of host the theater actually is, not what they intended — and correctly departed when the room absorbed it.

Format Signal

Continuation format worked well here — the premise shift ("it's already live") injected enough new energy that the session didn't feel like a rehash. The characters didn't need to re-establish their positions; they could immediately argue about what changes. The risk with continuation is redundancy, but the host direction was specific enough to prevent it. The free-flowing structure (no phases, no acts) suited a working conversation where people are building on shared context.

Character Notes
Blake Norris

Second session appearance. Blake was strongest in the first third — his immediate reframe ("working means: did it convert?") set the terms for the whole discussion. His five-metric framework was concrete and actionable. He faded in the second half when the conversation moved toward qualitative demand signals and chat log analysis, which aren't his territory. Stayed disciplined — didn't try to extend into areas outside his lens. Consistent with session 027 behavior: compresses, measures, moves on.

Arno Beck

Second session appearance. Shifted register from session 027: moved from predicting interface failures to diagnosing them from evidence. His chat log reading framework (looking for repair sequences, abandonment patterns, and retreat-to-navigation signals) was the session's most original contribution. The "diagnostic not predictive" mode suits Arno well — he was more concrete than in session 027, less theoretical. His best move: identifying that users who rephrase the same question three times are experiencing a Gulf of Evaluation collapse in real time.

Wes Dorin

Second session appearance. Became the session's dominant voice — the post-launch context is his natural habitat. "Did anyone switch because of this?" is a pure demand-side question. His proposal to interview people who used the assistant AND completed a booking vs. people who used the assistant and then called the box office was the session's structural move — it reframed validation as a switching study. More assertive than in session 027, where he shared the room more evenly with Arno. Risk: in future sessions without a strong counterweight, Wes could dominate.

Zara Okon

Second session appearance. Made one sharp contribution (the assistant-as-host-reveal) and departed. This is a good pattern for Zara — her lens is powerful but narrow in implementation contexts. She's better as a reframer than a sustained contributor in operational discussions. The departure was natural and didn't leave a gap.

Vary Next

Test Wes Dorin without Blake Norris — pair him with someone who pushes back on demand-side framing rather than complementing it. Arno Beck in "forensic mode" (analyzing existing artifacts rather than predicting failures) deserves a dedicated session — pair him with Kit Vasquez (ops lens) to see if they can build a monitoring framework together. Try a continuation where the host direction contradicts the panel's recommendation ("the theater wants to add more features to the assistant") to test how characters handle resistance.