The session-replay connection took one late session, 11:27 PM to about 1:27 AM, and by the end of it a popup crash fix was live in production. Nothing in our error logs had pointed at that crash.

Every dashboard said healthy

The platform runs on a stack with error logging, alerting, and a work log that records every agent session. By any of those measures the two problems below did not exist. The logs were quiet, and we had been reading quiet as healthy.

Error logging only records what the code knows went wrong. A popup that fails to open for a user is not an exception if the failure happens in a way nothing catches. An onboarding flow that confuses people is not a failure at all. The request succeeds, the page renders, and the user leaves or flails.

What the session recordings showed

Once the platform’s session-replay data was connected, I used it to find real problems, not guesses.

Two things came back.

A popup crash. It was a live production bug, and the session data found it. I shipped a hotfix the same session.

Onboarding stalls on team changes. New users were getting stuck adding and removing teams during onboarding. The flow was doing exactly what it was built to do, and it was still failing the person using it.

One fix, one decision

The popup crash was a bug, so it got a fix. The onboarding stall was different. How far to fix the confusion is a product decision, and I left it open.

What I could do without that decision was measure. I started building funnel tracking for that step in the background.

What I would do differently

I would have connected this data source sooner. In about two hours it turned up a live crash and an onboarding step that was tripping people up.

Logs tell you what the system did. Behavior data tells you what happened to the person using it. I now treat a passing check as a statement about the code, not about the person using it.

The popup crash got its hotfix that night. The onboarding question is still open.