The alert landed at 12:53 on a Saturday, and it said ERROR six times.

5 Min Monitor: ERROR - THERE ARE ISSUES!
MTA Queue: ERROR
Item had High value: 1762
Item had High value: 1446
VMTA Servers: ERROR
(1762) newsletters
Domains: ERROR
(1446) freemail-a.com
(311) freemail-b.com - Error: "552 1 Requested mail action aborted, mailbox not found" ... [2026-08-15 12:52:38]
Relayed: 89% good (290 bounced out of last 2652 today)
Mailers: 100% good (0 had errors out of 66 in the past 60 minutes)
Emails: 100% good (0 were invalid out of 104 in the past 60 minutes)

I read it the next morning and spent 38 minutes on it. Nothing was wrong.

What the monitor actually measures

We run a commercial MTA, version 4.0r19, on an unmanaged Linux box, split into four virtual MTAs, one per sending IP. The newsletters vmta carries bulk runs. The 5-minute monitor polls each vmta’s queue and each destination domain’s queue, and it trips when a count crosses a fixed threshold.

A large newsletter send makes exactly that count. The 1,762 queued on newsletters was almost entirely the two domains named below it: 1,446 for the first freemail provider plus 311 for the second is 1,757. That is a batch sitting in the queue while the MTA works through it at the pace each receiving server allows. The queue is supposed to look like that mid-send.

The alert bundles two different kinds of check. “MTA Queue,” “VMTA Servers” and “Domains” are infrastructure state: how many messages are waiting. “Mailers” and “Emails” are closer to outcome: did requests fail, were addresses invalid. Both outcome lines said 100% good. Only the depth lines were red.

The wrong turn

I did not start with that reading. The second provider’s line has an SMTP error string in it, and we had a real delivery incident this summer. For roughly ten days, the DNS record behind one sending IP’s reverse-DNS pair was missing from our DNS host, and that provider rejected the stream with 550 5.7.25. That one was invisible until someone noticed mail not arriving.

So I went hunting for a repeat. I searched our notes for the monitor’s name and got seven stale reconciliation-queue history files, nothing about the monitor. I read the forward-confirmed reverse DNS section of our internal infrastructure inventory, which covers all four vmtas and which zone each A record lives in. I was checking whether the newsletters PTR still resolved back to its IP.

That was the wrong question, and it took a good part of the 38 minutes to see why. The error is a 552 ... mailbox not found. That is a recipient-level hard bounce: the receiving server was reached, accepted the connection, and said this one address does not exist. A reverse-DNS failure looks different. It is a 550 5.7.25 at connection time, and it hits every message on that IP rather than one address at a time. A 552 is what dead addresses on a newsletter list produce.

The inventory section was useful reading and it cleared the DNS question, but I only got there because I had matched an error string to the wrong incident. The alert’s own shape should have sent me to the bounce line first.

What settled it

Two facts closed it:

  1. The outcome lines were clean. The 290 bounced out of last 2652 line reads as a worrying 89% good, but the errors on it were recipient-level 552s, not the connection-level 5.7.x rejections that would point at the IP.
  2. The MTA drained the queue on its own. No restart, no intervention. The depth came back down after the batch finished.

Conclusion: queue-depth thresholds tripping during normal delivery. No action needed.

What I’d change in the monitor

The threshold is a static count on a quantity that legitimately swings depending on whether someone pressed Send. A count tells you how much work exists. It says nothing about whether the work is failing. The alerts I’d keep page on outcome:

  • Bounce rate over a trailing window per vmta, with the receiving domain’s rejection code attached. A jump in 5.7.x policy rejections pages someone. A scattering of 552 mailbox not found does not.
  • Age of the oldest queued message, not the number of messages. A queue of 1,762 that is all under a few minutes old is a healthy send. A queue of 40 that has been stuck for two hours is a real problem, and the current monitor would say OK to it.
  • Queue depth as context on the alert, never as the trigger. It is a useful number to read once you are already looking.

The reverse-DNS incident argues for this too. That failure would never have moved a depth counter much. It would have shown up as a rejection pattern across one vmta, which is the thing worth watching.

The cost of the current setup

The direct cost was 38 minutes on Sunday for a Saturday alert. The larger cost is that a monitor which goes red on every big send teaches us to read its ERRORs as noise. When the next real rDNS-style failure arrives, it will look like every other red alert we ignored on the way there. Alert fatigue is a delivery risk of its own.

The fix is small, and we haven’t made it yet: keep the depth numbers on the report page where they can be read, and move the paging to bounce rate and oldest-message age.