Between 10 and 20 percent of our signup codes and password resets to Yahoo, AOL and Gmail had been hard-rejected for about ten days, and nothing in our code had changed in that time.

We found out because a support person handed me a customer whose signup code hadn’t shown up. He ended up getting his code and creating his site. What matters is what I did first.

The wrong turn: the queue that drained itself

The staff email report joins our order rows to the mail server’s accounting records by an envelope id. An order row with no matching record renders as bounceCat inQueue/noData, shown as Unknown. The stuck send sat there as Unknown, so I went down that path.

I filtered the report to the day and switched to the Queue/Unknown view. The day totals looked alarming: newsletters at 4397 total with 1937 unknown, and an all-sources count of 1978 unknown. I re-polled the 2:50 PM send about two and a half minutes later, and it had flipped to Good on its own. The unknown count for our platform’s server went from 4 to 2 over the same window. The MTA is shared across job servers, so a big newsletter batch adds latency to everything. I wrote it up as batch lag and told the next session the customer’s mail path was healthy.

That was wrong, and it cost real time. Batch lag was true, but it explained the wrong messages. Three entries were still sitting hours old, including one from 12:09 AM, and they weren’t draining. I noted them as “small, none looked customer-facing” and moved on. The Unknown column can’t tell you whether the receiver refused a message. It only says the MTA record hasn’t been matched yet. I had picked the one view where a hard reject and a slow drain look alike, and then trusted the drain.

Reading the MTA log directly

What broke the loop was a separate observation. My partner’s own message to a Gmail address had shown up about seven minutes late. That made me stop reading our report and go read the mail server’s delivery log, which the MTA exposes over HTTP as a file fetch on its log path. The rejects were there in plain text, and one of them, from 00:26 that day, was a Gmail 550 5.7.25. That code means Gmail rejected the connecting IP for failing forward-confirmed reverse DNS (FCrDNS).

The check works like this. The receiver takes the connecting IP, looks up its PTR record to get a hostname, then looks that hostname up and confirms it resolves back to the same IP. If the pair doesn’t close, some receivers throttle you and some refuse you outright. Yahoo and AOL refused. Gmail did both, which is the likeliest reason my partner’s message arrived seven minutes late. I’m inferring that last part. The timing and the 5.7.25 fit, but I can’t prove causation on one message.

The PTR on our transactional sending IP pointed at a hostname in our sending domain, out.<our-mail-domain>. I ran the lookup against a public resolver and then against the authoritative Cloudflare nameserver. Both came back Non-existent domain. The forward half of the pair had no A record. Nothing existed at that name at all.

Why nothing would have caught it

The application has no dependency on that record. The code hands a message to the MTA and the MTA reports it accepted. The failure happens later, in a connection between our IP and someone else’s server, and the check runs on their side. A code audit, a dependency list, and a test send to a Gmail address (which mostly tolerates it) would all pass. This is also why the report said Unknown instead of Bad: when the rejection is at the SMTP conversation, the accounting record can be missing or late, so it reads as a queue.

We still don’t know when or why the record left the zone. I can’t tell you a cleanup removed it, because I couldn’t find the change. All I have is that it was gone and that the zone’s default TTL is 1800, so any receiver that cached the missing name could keep rejecting for up to about half an hour after the fix.

The fix, and what it did and didn’t prove

Our Cloudflare API tokens failed their auth check, so the record went in through the dashboard. Before touching anything I captured the full pre-change record set into a scratch folder alongside a rollback.py. Since nothing existed at that name, rollback is just deleting the new record. After the change I confirmed reverse DNS resolves correctly on all four of our sending IPs, not only the transactional one.

Then I checked the log for new rejects. There were none after 2:20 PM, which was before the change went in. That is an absence of new failures, and I was careful to say so. I could have forced a proof by triggering a real signup code to a Yahoo address, but that sends real mail through production, and the go I had covered a DNS record, not a test send. The actual confirmation is a real Yahoo delivery, and that is the check for the next day.

What the fix left open

Three follow-ups came out of the audit:

  • Whether to contact the customers who got no email for a week.
  • Messages that show as sent in our system but never reach the mail server at all.
  • The mail server’s clock, which runs 11 minutes behind. That one is a separate ticket and I have no path onto the box yet.

Ten days of silent loss on a path with no alarm is the cost here. The most useful thing we built afterward is a short note on how to read the MTA delivery log first, and to treat an Unknown row that is hours old as a failure until proven otherwise.