Six real bugs came out of the last day of the infrastructure-graph program, and one of them was an API returning credential data to a customer’s own AI assistant. Nobody had a ticket for it. It turned up while we were finishing the read side of a system whose write side we had already locked down.

What the write hardening covered

The program turns a business’s infrastructure register into a graph. Staff can log into any client account and see and edit exactly what the customer sees, with every action attributed to the staff member at the database level. Credential data sits behind a boundary customers cannot cross. Customers’ AI assistants can read and edit their own data.

Most of the effort went into writes, because writes are what scare you. Attribution at the database, a trigger on the flag that turns graph reads on, and a test that drives the real member path instead of calling the function directly.

The wrong turn: trusting a row policy to protect a column

The bizops_app role holds UPDATE on the businesses table. The policy, businesses_write, is keyed on id = app_current_business(). I read that as “a tenant can only touch their own business,” and I treated that as the end of the write question.

It protects the row. It says nothing about which columns on the row a member may change. Columns that were meant to move only through a platform command were assumed safe by comment, not by constraint. I found this during the phase 4 backfill by reading the live catalog, not by reasoning about the policy. Any member could update their own infra_graph_reads_enabled.

The fix for that flag was a BEFORE UPDATE trigger that refuses the member path, plus a test that goes through the same door a member would. The identical exposure sits on delivery_stage. We left it alone on purpose because it belongs to a different workstream, and filed it as its own task: “Any member can flip their own delivery_stage.”

The cost of the wrong turn was time and a false sense of closure. I had spent the effort on the write side and considered that side done, based on a policy I had read but not attacked.

Reading the read side with the same suspicion

Once the write side turned out to have a gap, the read side got the same treatment. The graph is now the register, backfilled in production with reads turned on. The question for reads is the one the write policy never answered: for each thing a caller can fetch, who is allowed to see it?

The credential leak turned up on that pass. An API endpoint returned credential data to a customer’s own assistant. The customer owned the assistant, and it was reading its own data, so the tenant boundary held. But the credential boundary is a second, separate line, and the endpoint crossed it. Row-level tenancy told us the caller was the right business. It told us nothing about whether that caller should see secrets.

That is the asymmetry. Write permissions get scrutiny because a bad write is loud. Read permissions get inherited from whatever the row policy allows, and a bad read is silent. Nothing errors. Nothing gets logged as wrong. The response is a valid 200 with too much in it.

What we do now

  • Every graph-facing endpoint gets a test from the outside, authenticated as a plain member or as a customer’s assistant, asserting on the response body and not just the status code.
  • Credential fields live behind their own boundary. A row policy on tenancy is never allowed to stand in for it.
  • Any column whose safety rests on “only the platform command touches this” gets a trigger or a constraint. A comment does not count.
  • Two tasks that had been marked done while still unresolved got reopened. Done status is a claim, and we checked it against the live system.

I filed four findings that day into the product’s own register, through its UI, and left all four as drafts. The first is the delivery_stage exposure, and it shows the whole problem in miniature: the fix for one column was clean, and the same hole was sitting next to it on another column, still open.