The channel already had the plan in it
My partner and I run a shared venture out of a Discord channel. Decisions get made there, ideas get dropped there, someone asks a question and three tangents follow before anyone answers it. None of that gets written down anywhere else. For months my answer to “where are we on X” was to scroll. The backlog was not missing. It was just sitting in a chat log, unshaped.
I set out to fix that with a task engine, and the first design decision mattered more than any of the code: the conversation itself had to become data, not just a place where data gets discussed.
Three layers, in that order
The system has three moving parts and they run in a strict sequence.
First, ingest. A channel gets pulled in as source records: every message, timestamped, attributed, stored once, and never edited after that. This layer does not interpret anything. It just makes the conversation quotable.
Second, decomposition. An agent reads the source and proposes atomic units: a decision, a request, an opinion, a question, an idea, a piece of work, a note. Each unit carries evidence, a literal pointer back to the message it came from, because a unit that cannot be traced to a real quote is not trustworthy enough to act on. On the first real decomposition run this caught its own agent: eleven units proposed, ten accepted, one rejected because the quote attached to it was not verbatim. That is the whole point of the evidence field. It is not decoration, it is a check the system runs on itself before a unit is allowed to exist.
Third, grouping. Units get clustered into tasks, and tasks into epics, and the result lands on a review surface. Nothing here writes itself into the record of truth. It proposes.
The tangent is not a bug
The thing that took me longest to accept, and that the design had to be built around, is that a real working conversation does not stay on topic and should not be forced to. My own working note during the spec process put it plainly: whatever we’re working on is going to reference other things, and I should be able to say whatever I want on any task without stopping to file it correctly in the moment. The friction people actually feel with task systems is being made to sort a thought at the instant they have it. So the system defers sorting. You talk. Later, something else reads the transcript and does the filing.
That single decision is why decomposition has to be a separate step from grouping, and why both have to be separate from the immutable source underneath them. Three layers, not one, because each one earns something different: the source layer earns trust, the unit layer earns traceability, the grouping layer earns organization. Collapsing them loses whichever property the collapsed layer was providing.
An early version of this tried to make task-close the forcing function: close a task, and whatever units under it never got resolved would get swept somewhere. A review at high effort killed that idea outright, because all it did was relocate the mess into a “parked” pile instead of actually resolving anything. The fix was to make parking require an explicit trigger and a reason, to track how often parked units actually get revived as a real kill metric, and to accept that if the revival rate stays near zero, this whole approach is an elaborate archive and should be shut down rather than expanded. I would rather build that kill criterion in on day one than discover two months in that I built a graveyard.
What the first real run turned up
The first time I ran this against a real batch of channel history, out of 270 decomposed units, 190 got filed into a working structure of 17 tasks and 6 epics. The rest stayed loose, which is correct. Not every message in a working channel is a task. Some units are genuinely just a note or an opinion that never needed a home, and forcing every fragment into a bucket would have meant grooming an unusable pile instead of a real backlog. Filing 190 out of 270 is not a bug, it is the decomposer doing its job: separate what wants structure from what does not.
Immutable source, reversible grouping
The property I care most about is the split between what never changes and what always can. The source records are permanent. Nobody edits history, not a person and not an agent, because the whole value of the layer is that a year from now a unit’s evidence still points at a real message that actually happened. But the grouping on top of that is meant to be pulled apart constantly. A unit can be pulled out of one task and dropped into another. A task can be re-scoped without touching the units under it. The design spec is explicit that promotion is a validated transition, not just flipping a type field, and that containment is an edge in a graph rather than a fixed parent column, specifically because a single unit can belong to a conversation and a task and a related task all at once. Grouping is a hypothesis about how the work is organized. The source is the fact of what was said. Only one of those two should ever be permanent.
Why a human still has to say yes
None of this writes to the record of truth on its own. Every promotion from unit to task and every unit placed under a task lands on a review surface first, and a person confirms it. An agent is never allowed to self-promote its own proposal, full stop. When I built the bulk-approve action for reviewing a full task’s worth of low-risk units at once, I put a hard rule in it too: units typed as a decision, a proposal, or a commitment never get swept up in a bulk action, no matter what the form claims. Those always get looked at one at a time, and the bulk action itself will not run until a reviewer has actually seen one real sampled quote from the batch, enforced with a required checkbox next to it so the button cannot be pressed unseen.
There is also a boundary the system has to respect on the way out, not just on the way in. My partner and I keep separate private trackers alongside the shared one, and early on a digest generator let my own private task IDs leak straight through into the shared channel because nothing was checking for that shape of string. The fix was a fail-closed rule: if a generated summary still contains one of those IDs after the model writes it, the card gets refused rather than posted. Same logic as the evidence check, just pointed outward instead of inward. Trust the process, but verify before anything crosses a boundary you cannot take back.
The channel was always the backlog. It just needed a layer underneath it that would not forget what was actually said, and a layer on top that a person still has to sign off on before it becomes real.