A customer’s session-timeout message named the wrong cause. It said their session had been set to expire in a year. The session that actually died belonged to a different account, one on a much shorter default. The dialog was reporting one account’s setting as the reason for another account’s logout.
That reads like a copy bug: swap the message, point it at the right value, close the ticket. I traced the whole path before touching anything, because a dialog confusing state between two accounts meant something upstream was doing the same thing.
What the trace found
The redirect that sends a person to the session-timeout page carries the real reason as a URL parameter. Nothing downstream ever reads it. The dialog just shows whatever the current session’s own setting is, and after a fresh login that is the new session, not the one that actually timed out. That explained the wrong message, but the same trace turned up something the ticket never mentioned.
The login page builds its form action from the entire incoming query string, unexamined, and hands the whole thing straight through to the authentication endpoint on submit. One of the values riding along in that query string is a redirect target, and after credentials are accepted, the endpoint sends the browser there with no check on host, scheme, or leading slash. The only thing standing between that value and the browser is a step that makes sure it is a string, which does not sanitize anything.
That is an open redirect. A link to the real login page, on the real domain, with the real certificate, can carry a hidden destination that fires the instant a real login succeeds. Every signal a person is taught to check, the domain, the padlock, the login form itself, is genuine right up to the moment they get sent somewhere else entirely.
Two smaller things came out of the same trace. Sessions that time out were never actually marked expired in storage, so if a timeout setting was raised later, an already-dead session could come back to life. And a settings page offering a fixed list of session lengths had no fallback for accounts holding a value outside that list, so opening the page and saving without changing anything silently dropped those accounts to the shortest setting on the list. 172 accounts held values outside the offered list.
Where I got it wrong, and how fast that showed up
I shipped a fix the same night: require a real path on the redirect target, write a true expiry flag instead of leaving the row valid, and put an allow-list on the session-length setting in both directions. I scoped the redirect fix to the login flow the original report pointed at, because that is where the bug lived from the customer’s side.
That scoping was wrong, and it did not take until the next day to find out. The identical redirect logic existed on a second login flow, for accounts with meaningfully more access than a customer account, completely unguarded, and the code that routes an expired session to the right page sends that account type through exactly that flow. Going back over the fix in the same sitting, minutes after shipping it, is what caught the gap. A second, narrower release closed it before the night was over.
The lesson wasn’t “check twice” in some vague sense. It was specific: when a fix addresses a pattern, not a single symptom, the fix has to be scoped to every place the pattern appears, found by searching the whole codebase for that shape of code, not to the one flow a bug report happened to name. That search would have taken a few minutes done first. Doing it after shipping instead of before is the only reason this closed in one night instead of staying open.