We’d already built our internal ops-tracking system’s Discord integration: daily, weekly, and monthly review cards, auto-posted on a schedule, pulled from the same session logs that drive the work log itself. First pass through, I backfilled the daily channel through Jul 22 and ran the weekly job across the last few weeks to seed history. That’s when it broke.

The alert I actually wanted to kill wasn’t the outage. It was the Discord bot posting a clean daily summary card into #reviews at 3:40pm on a Wednesday, for a week that wasn’t over yet.

What shipped first

The weekly job for Jul 5 to Jul 11 ran fine. Then I reran it, by accident, a minute later, same range, and it posted again. Then the Jul 12 to Jul 18 card posted twice too. Four weekly posts in ninety seconds, two of them exact duplicates, sitting in a channel that was supposed to be a clean historical record.

Duplicate posts are a nuisance. What made me stop and rethink the whole thing was the content of one of those cards: a “weekly review” for a week that included days with no session data yet, because the cron had fired mid week and the script didn’t know the difference between “no work happened” and “the data isn’t in yet.”

The wrong turn

My first fix was rate limiting. I figured the duplicate posts were a job-overlap problem, so I added a lock file, .state/weekly_post.lock, keyed on the date range, and made the poster bail if a post for that range had gone out in the last hour. I shipped it, reran the backfill, and it worked. No more duplicate weekly cards.

But the lock file didn’t touch the real problem. Two days later the daily card posted for “today” at 8am, before most of the day’s sessions had even started, showing a mostly-empty day that looked like a slow day when it was actually just an early snapshot. The lock file had solved duplication. It had done nothing about completeness. I’d fixed the symptom I noticed first instead of the mechanism underneath both symptoms, and I’d spent the better part of an hour building and testing a lock file that the eventual fix made irrelevant.

What actually needed to be true

The real invariant: a report card should never represent a day or week that isn’t done yet. Not “isn’t done” in the sense of business hours ending, but “isn’t done” in the sense that new session data could still land in that bucket. Two different things can make a bucket incomplete:

  • Not prepaid: the day/week is still in progress relative to wall clock. A Wednesday afternoon can’t summarize Wednesday.
  • Not complete-days-only: even a past day can still be mutated if a session logged against it hasn’t closed yet, since our work log entries get written after the session finishes, not when it starts, and long sessions (the 17h 29m platform-monitoring one, the 13h 8m CPU-leak one) routinely span midnight.

So the gate isn’t a lock file, it’s a precondition on the query itself. Before the daily job posts for date D, it checks that D < today in the reporting timezone, and that no open session has a start timestamp before D+1 00:00 and no end timestamp yet. Before the weekly job posts for a range, every day in that range has to independently pass the same daily gate. The monthly card requires every week in the month to have already posted its own weekly card, so completeness composes upward instead of getting re-derived at each tier.

Concretely, that’s why the backfill only went through Jul 22 (the day before the run, not the day of), and why the weekly channel only got the two ranges that were fully closed, Jul 5 to Jul 11 and Jul 12 to Jul 18, and not the partial week the backfill was actually run inside of. The gate rejected the current week outright, and it rejected it before any post attempt, not after a duplicate or a stale-looking card was already sitting in the channel for someone to notice.

The part that generalizes

The lock file is still in there. It’s cheap insurance against a cron double-fire. But it’s not load-bearing anymore. The load-bearing check is the completeness gate, and it runs on every scheduled post, not just backfills, because a cron that fires at a fixed wall-clock time has no idea whether the data behind that clock tick is actually settled. Silence is a fine failure mode for a report. A confidently-wrong summary sitting in a channel next to real ones is worse, because nothing about it signals that it’s wrong.