The session-replay connection took one late session, 11:27 PM to about 1:27 AM, and by the end of it a popup crash fix was live in production. Nothing in our error logs had pointed at that crash.
Every dashboard said healthy
The platform runs on a stack with error logging, alerting, and a work log that records every agent session. By any of those measures the two problems below did not exist. The logs were quiet, and we had been reading quiet as healthy.
Error logging only records what the code knows went wrong. A popup that fails to open for a user is not an exception if the failure happens in a way nothing catches. An onboarding flow that confuses people is not a failure at all. The request succeeds, the page renders, and the user leaves or flails.
What the session recordings showed
Once the platform’s session-replay data was connected, I used it to find real problems, not guesses.
Two things came back.
A popup crash. It was a live production bug, and the session data found it. I shipped a hotfix the same session.
Onboarding stalls on team changes. New users were getting stuck adding and removing teams during onboarding. The flow was doing exactly what it was built to do, and it was still failing the person using it.
One fix, one decision
The popup crash was a bug, so it got a fix. The onboarding stall was different. How far to fix the confusion is a product decision, and I left it open.
What I could do without that decision was measure. I started building funnel tracking for that step in the background.
What I would do differently
I would have connected this data source sooner. In about two hours it turned up a live crash and an onboarding step that was tripping people up.
Logs tell you what the system did. Behavior data tells you what happened to the person using it. I now treat a passing check as a statement about the code, not about the person using it.
The popup crash got its hotfix that night. The onboarding question is still open.
AI Skills
Use this lesson with the AI assistant you already use
A two-hour session-replay connection surfaced a live popup crash and a silent onboarding stall that had never triggered an error, an alert, or a log line despite the platform's full monitoring stack.
Paste the prompt, share only the context needed to answer it, and treat the result as a draft for your review. Do not include confidential information or let an AI assistant make changes without your approval.
Optional: for a visual report and saved memory, run /dxdev first.
Don’t have it? Get it at dxdev.com/skills/dxdev. The prompt works without it.
dxdev LESSON · paste into your AI coding agent
LESSON: Instrument Behavior Data Before You Trust Your Error Logs
SOURCE: dxdev.com/blog/2026-09-18_product-analytics-finds-silent-breaks
WHAT HAPPENED: The team connected session-replay data so an agent could query real user behavior instead of guessing. With the rule that every answer had to come from an actual recorded session, it found two problems invisible to error logging: a popup that crashed for real users without throwing a catchable exception, and new users stuck cycling through adding and removing teams during onboarding. The popup crash was a bug and got a same-session hotfix. The onboarding stall wasn't a bug since the flow worked as built, so it was left for the product owner with evidence attached, while funnel tracking for that step was started in the background.
THE RULE: Error logs and alerts only catch failures the code recognizes as failures; behavior data is required to find flows that succeed technically but fail the person using them, and bug fixes should be separated from product-scope decisions even when both surface from the same investigation.
CHECK MY CODE, then report PASS or FAIL with file:line for each:
1. Is there a way to inspect real user session behavior (replay, funnel, or event data), not just error logs and alerts, when investigating whether a flow works?
2. For flows with no caught exceptions, has anyone verified they actually complete for real users rather than assuming silence means success?
3. When an agent or investigation finds a UX/product issue versus a code bug, is the product-scope decision routed to the product owner instead of being resolved unilaterally?
THEN PRINT: a table (check, PASS/FAIL, evidence, fix) + a verdict (applies / partially / OUT_OF_SCOPE / no) + the single most important next action.