“Our chatbot will NOT stop talking about 🌭🌭🌭. Every response, a glizzy reference. Mustard this, Chicago dog that.”
“Why does our cookbot keep talking about hot dogs? What’s going on?”
“…oh. There’s a line in the instructions. I forgot that’s still in there.”
Okay, so maybe I added that hot dog line on purpose for this demo. But this sort of thing comes up all the time now.
In AI for Good, Josh Tyrangiel tells the story of Khan Academy building Khanmigo with OpenAI. The team expected the hard part to be software engineering. It wasn’t. The hard part was watching students talk to the tutor, catching where it failed, and rewriting the prompts, over and over. The instructions encoding how it should teach kept growing and changing. Prompt work felt less like coding and more like curriculum design. There was no permanent “done.”
As chat interfaces get more common, we see it ourselves: testing with our team, testing with clients, hitting an unexpected answer, then having to dig in and figure out why. In a lot of cases that’s much harder to track than a deterministic bug in code. But you can’t answer with a shrug.
We’ve been thinking of experiments that could help close that loop. So we built a thing.