I Wrote a Playwright Script to Test LLM Long-Term Memory — and Found 3 Critical Bugs
<p>It was 1 a.m. when the PM dropped a screenshot into the group chat: “Your AI assistant forgot the customer’s name again. Third time.” I zoomed in and dragged my finger across the words “Dear user, hello!”. My heart sank — the “long-term memory module” we had just shipped was completely unreliable in the real world. I had manually tested dozens of conversations. It always <em>felt</em> fine, but the moment it hit real users, everything fell apart. I closed Slack, opened my IDE, and decided to





