```
Same cast as session 027. No shared primary sources.
Bob Moesta → Wes Dorin (primary)
Don Norman → Arno Beck (primary)
Oli Gardner → Blake Norris (primary)
Priya Parker → Zara Okon (primary)
Watch for: Wes/Blake convergence on “conversion” language from different
directions. Arno/Zara convergence on “experience of entering” from
cognitive vs. social frames.
```
Same rehearsal room. The whiteboard from session 027 is still up — three columns: DEMAND, WIZARD OF OZ, PROTOTYPE, each with entry criteria and kill signals. The four are back around the table. Coffee is fresh. Someone has taped a printout of a chat transcript to the whiteboard — it’s a real conversation from the theater’s live assistant.
So they skipped our layers and shipped it. Fine. That changes the question completely. We were designing a test plan for something hypothetical. Now we’re auditing something real. The question isn’t “should this exist” — it already exists. The question is: is it earning its place on the page?
And “earning its place” doesn’t mean “do people click on it.” It means: did anyone buy a ticket because of this assistant who wouldn’t have bought one otherwise? That’s the only metric that matters to the theater. Everything else is vanity.
That’s one metric. There’s another one that matters just as much: did anyone fail to buy a ticket because the assistant got in the way? A chat widget that pops up and obscures the booking flow, or that gives a wrong answer about showtimes, or that sends someone down a conversational dead end — that’s not neutral. That’s actively harmful.
Both of those are measurable. Let me frame five things we can look at right now, with data the theater probably already has. One: chat initiation rate — what percentage of visitors open the chat? Two: completion rate — of those who open it, what percentage reach a resolution? Three: conversion lift — do visitors who use the chat buy tickets at a higher rate than those who don’t? Four: deflection — did support calls or box office inquiries drop after the assistant launched? Five: abandonment pattern — where in the conversation do people stop responding?
Three and four are the only ones I care about. The others are activity metrics. Someone can initiate a chat, have a complete conversation, and still not buy a ticket. That’s a successful chat and a failed assistant. The theater didn’t install this to have conversations. They installed it to sell tickets and reduce the load on their staff.
Agreed, but you need one and two to diagnose why three and four are failing, if they’re failing. If initiation is low, the widget is invisible or uninviting — that’s a design problem. If initiation is high but completion is low, the conversations are breaking. You need the funnel to find the leak.
I want to look at the chat logs directly. The quantitative metrics tell you that something is working or not. The logs tell you why. And I know exactly what to look for. Three patterns. First: repair sequences. That’s when the user rephrases the same question — “What’s playing Friday?” then “I mean this Friday, March 13th” then “shows on Friday night?” Three versions of the same question means the assistant failed to understand and the user is doing the cognitive work of fixing the conversation. That’s a Gulf of Evaluation collapse — the user can’t tell what the system understood, so they keep trying.
That’s good. What’s the second pattern?
Escalation requests. “Can I talk to a person?” or “Is there a phone number?” That’s the user pulling the emergency brake, same as I described last time — retreating to a familiar interaction model because the conversational one broke. Every escalation request is a data point: this is where the assistant’s model diverged from the user’s need so badly that they gave up on the medium entirely.
And the third?
Navigation retreat. The user closes the chat and goes to the regular website menu. You can track this if the analytics are set up — chat close followed by a page navigation within thirty seconds. That’s the user saying “this isn’t working, I’ll do it myself.” It’s softer than asking for a human, but it’s the same failure.
There’s a fourth pattern you’re all missing, and it’s not in the logs. It’s in what the assistant reveals about the theater. Every conversation the assistant has is a hosting act. And the question isn’t just “did the user get an answer” — it’s “what kind of host did the theater just present itself as?” If someone asks “is this show suitable for my 12-year-old?” and the assistant says “I don’t have information about age suitability” — that’s not a missing feature. That’s the theater telling a parent “we didn’t think about you.” The chat logs are a mirror. They show you who the theater actually is as a host, not who they think they are.
That’s exactly right. And it connects to something I want to do with the logs that’s different from Arno’s approach. Arno is reading the logs for interaction failures — where the conversation broke. I want to read them for demand signals — what are people actually asking for? Not what the theater assumed they’d ask. The gap between “questions we designed the assistant to handle” and “questions people actually ask” is the most valuable data in the entire system.
And the questions people ask that the assistant handles badly are where those two analyses converge. A demand signal that produces a repair sequence — that’s a real need meeting a broken interface.
Zara pushes the printout on the whiteboard so it’s more visible.
Read that transcript. Someone is asking about bringing a group for a birthday. The assistant gives them showtimes and a link to group booking. Perfectly functional. Completely inhospitable. No one said “happy birthday, what a great idea.” No one asked “how old is the birthday person — we might have a suggestion.” The assistant answered the question and missed the moment. That’s the hosting test. The front-of-house staff would never do that.
And that’s a demand signal too. The person didn’t say “I want group tickets.” They said “birthday.” The job isn’t “buy group tickets.” The job is “make this birthday special.” The assistant answered the supply-side question — availability, pricing — and completely missed the demand-side question — “help me create a memorable evening.”
But hold on. You’re describing a feature enhancement, not a validation failure. The question right now is: is the assistant working? You’re saying it could work better. Those are different.
No, Blake, they’re the same question. “Working” means “doing the job it was hired to do.” If the theater hired this assistant to be modern, then sure, it’s working — it exists, it’s chatty, it’s modern. But if the theater hired it to help people buy tickets — which is what they actually need — then the birthday conversation is a failure. The person didn’t buy tickets through the assistant. They got a link. That’s the equivalent of a sales associate pointing at a shelf and walking away.
Fair. So the validation framework is: for each conversation, did the assistant move the visitor closer to a purchase, or did it just answer a question? Answering a question is necessary but not sufficient. The metric isn’t “resolved” — it’s “converted.”
I’d push back on making conversion the only lens. Some conversations are legitimately informational and still valuable. Someone checking accessibility features, someone looking at the season calendar, someone asking about parking. These don’t convert in that session but they remove barriers for a future visit. If you only measure conversion, you’ll optimize the assistant into a pushy salesperson and destroy the trust that makes the informational conversations valuable.
That’s fair. There are two jobs. The immediate job: help me buy a ticket right now. And the enabling job: help me feel confident that when I’m ready to come, it’ll work. Both are real. Both are measurable — just on different timescales. For the immediate job, measure same-session conversion. For the enabling job, measure return visits. Did the person come back to the site? Did they eventually buy?
Zara stands, having made her point. She picks up her coffee.
One last thing. When you review the logs, don’t just look at what people asked. Look at how the assistant opened. The first message the assistant sends — before anyone types anything — is the theater’s greeting. It sets the tone for the entire gathering. If it says “How can I help you?” that’s generic. If it says “Welcome to [theater name] — looking for tonight’s show, or planning ahead?” that’s hosting. The opening is the threshold. Check whether it matches who the theater wants to be.
Zara nods and steps out. Three remain.
She’s right about the opening, and it’s testable. A/B test the greeting. Version A: generic. Version B: specific, with two or three paths. Measure which one produces more engaged conversations — not just more initiations, but more conversations that reach a resolution.
The greeting is also an affordance question. “How can I help you?” gives zero information about what the system can do. It’s a blank text field with a question mark. The user has to guess what’s possible. “Looking for tonight’s show, or planning ahead?” is a signifier — it tells the user what kinds of input the system expects. That’s a direct reduction of the Gulf of Execution. The user knows what to do next.
So here’s what I’d actually do if I were in the theater tomorrow. Step one: pull all chat logs from the first month. Categorize every conversation by what the person was actually trying to do — not what the assistant categorized it as, but what the human wanted. You’ll find five to seven clusters. Some will match what the theater expected. Some won’t. The surprises are gold.
Step two: for each cluster, measure the outcome. Did the conversation end in a ticket purchase, a link click to the booking page, an escalation to box office, a navigation retreat, or a dead end? That gives you a conversion map per job type. You’ll immediately see which jobs the assistant handles well and which it fumbles.
Step three: for the fumbled jobs, read the individual transcripts. Apply the three patterns — repair sequences, escalation requests, navigation retreats. You’ll find the exact moment each conversation broke. That’s your fix list. Not “improve the assistant generally” — fix these specific breakdowns in these specific conversation types.
Step four — and this is the one everyone will want to skip — interview five to ten people who used the assistant. Not a survey. A conversation. “I see you chatted with us on Tuesday. Walk me through what happened. What were you trying to do? What did you expect? What actually happened? Did you end up coming to the show?” That’s where you’ll hear the stuff the logs can’t tell you. The emotions. The “I felt stupid asking the chatbot.” The “I wasn’t sure if it was a real person.” The “I gave up and called my friend who goes to theater all the time.”
That’s the most expensive step. Can we get 80% of the signal from the logs alone?
No. The logs tell you what happened. The interviews tell you why it mattered. A conversation that looks successful in the log — question asked, answer given, booking link clicked — might be a failure in the person’s experience. “It gave me a link but I still wasn’t sure which show to pick, so I just booked the cheapest one and hoped for the best.” That’s a conversion in your data and a disappointed customer in reality.
Wes is right. The logs have a visibility problem — they show the interaction but not the understanding. The user might click the link and still not know if they made a good choice. That’s the Gulf of Evaluation at the outcome level, not just the conversation level. The assistant resolved the query but didn’t resolve the uncertainty.
Okay. So the validation plan is four steps. Categorize conversations by actual intent. Measure outcomes per category. Diagnose breakdowns in the transcripts. Interview a sample for the qualitative layer. What’s the kill signal here? At what point do you tell the theater “turn this off”?
The kill signal is if the assistant is creating struggling moments that didn’t exist before. If people who use the assistant convert at a lower rate than people who don’t — controlling for intent — then the assistant is a net negative. It’s not just failing to help; it’s actively getting in the way of people who would have figured it out on their own.
There’s a softer kill signal too. If the majority of conversations end in escalation or navigation retreat, the assistant isn’t harmful — it’s just useless. It’s adding a step. The user tries the chat, fails, and then does what they would have done anyway. That’s not worth the page real estate or the cognitive load of “should I use this thing?”
And there’s a third signal: if the chat initiation rate is below five percent and falling, the market has spoken. People looked at it, decided it wasn’t for them, and stopped trying. At that point it’s furniture. Either redesign the entry point or remove it.
But before you kill it — check one thing. Check whether the people who do use it love it. A five percent usage rate with a high conversion rate means you have a niche tool that works beautifully for a small audience. The question then isn’t “should we remove it” but “should we make it more visible to the people it serves?” Maybe it’s perfect for first-time theatergoers and invisible to regulars who already know what they want. That’s not a failure — that’s a segmentation insight.
Which brings it back to the fundamental design question: is the assistant visible to the right people at the right time? If it’s a generic widget in the corner, everyone sees it and most ignore it. If it appears specifically when someone lingers on the “what’s on” page for more than twenty seconds — that’s a triggered affordance. It appears when the behavior suggests someone might need help choosing. That’s a completely different design than a persistent chat bubble.
That’s testable too. Version A: persistent widget, always visible. Version B: triggered widget, appears on specific pages after a dwell time threshold. Measure initiation rate, conversation quality, and conversion for both. I’d bet B outperforms on every metric because it’s serving the assistant to people who are actually in a choosing moment, not a browsing moment.
And the choosing moment is the struggling moment. That’s where the demand lives. Someone who knows exactly what they want — “two tickets for Hamilton on Saturday” — doesn’t need an assistant. They need a search bar. Someone who’s staring at fifteen shows they’ve never heard of and doesn’t know where to start — that’s the person the assistant was built for, whether the theater knows it or not.
The whiteboard now has two frameworks side by side: the original three-layer testing plan from session 027 (crossed out with a note: “shipped — skipped to live”), and the new four-step validation plan: CATEGORIZE → MEASURE → DIAGNOSE → INTERVIEW, with kill signals noted at the bottom.
One more thing. Whatever they learn from this validation — whatever they fix — they need to keep validating. This isn’t a one-time audit. The assistant will change as the season changes, as new shows open, as the FAQ shifts. The validation loop has to be continuous. Pull logs monthly. Watch for new failure patterns. The assistant that works in September might break in December when everyone is trying to book holiday shows for groups they’ve never brought to theater before.
Monthly is right. And automate what you can. Flag conversations with three or more user messages in a row without a resolution. Flag conversations that end with an escalation request. Flag conversations where the user’s first message and last message are the same question rephrased. Those are your automated canaries. When they spike, something broke.
And once a quarter, do the interviews. Not because the logs aren’t enough — because the logs can’t tell you about the person who thought about using the assistant and decided not to. That person doesn’t exist in your data. They’re the non-consumer. And they might be the biggest opportunity.
AI-generated approximation. All characters are inspired by published work and are not reviewed or endorsed by the original thinkers.