Stage 02 · User research
What they are trying to get done
One main job, four adjacent ones, and a matrix that answers two different questions: what to build first, and what not to build at all. Every job is solution-agnostic, so a completely different product could be hired for it.
- Main jobs in the product
- 1 a second is recorded, not merged
- Related jobs
- 4 a fifth was not written
- MVP core
- 3 two qualified, one is a bet
- Critique, two instruments
- 20 findings, one overlap
Wording here is canonical. personas.md carries the same formulations word for word, verified after the critique changed two of them. A job that exists in two editions is a job that gets built twice.
01 · Who owns the steps
A decision recorded before the jobs
Related jobs are separate adjacent tasks done in the same context. The steps of the journey toward the main job are normally CJM phases, and this project runs a shortened track where CJM is not included.
The steps of the journey are owned by the user flow at stage 03a. That flow already has to describe the path to the activation node, First Verdict, so it is the natural holder. The related jobs below are genuinely adjacent tasks, and stage 03a should not read them as a sequence.
02 · The main job
One, and it belongs to Rasha
When Clerk hands me a case it has already investigated, I want to decide whether its verdict holds, so that the decision is made and I can still defend it months later.
Solution-agnostic check. No function named. A peer review process, a checklist or a second analyst reading over a shoulder could all be hired for this. The job does not lock the answer.
P2 has a different main job, and it is recorded rather than merged
When I am accountable for what the agent did across all my clients, I want to know where its record has earned more latitude and where it has lost it, so that I can widen or narrow its scope without guessing.
Two main jobs is a signal of two products, so it is named here rather than glued into one formulation.
The reading taken. The product has one main job, Rasha’s, because she is primary and the activation node is First Verdict. Tomas’s job is real and is served by the same substrate read at a different scale: she reads one case, he reads the trail across many. It stays one product for as long as the record underneath them is one record.
The alternative, rejected with a reason. Treating them as co-equal makes the information architecture build two navigations, and the product splits into a case console and a fleet console that happen to share a database.
04 · The matrix
Jobs against personas, functions and the market
Importance is inferred, not measured. No importance number exists in the research, so each row names the signal it rests on. Where there is no signal the cell is [?] rather than an average.
| Job | P1 | P2 | Signal behind the importance | Function | Competitors |
|---|---|---|---|---|---|
| Main · rule on the case | 3 | 1 | Her whole shift, and the activation node. He reads outcomes rather than adjudicating each case | Case file, Verdict, Autonomy state | No. Nobody publishes the residual queue |
| R1 · the shift | 3 | 2 | An incoming team must read six reports to hold three days of context, and older ones stop being read. 79% of SOCs run 24/7 | empty in the canonical list | No. Verified across thirteen products |
| R2 · answer for a past decision | 2 | 3 | Rare per case, but it is why the record exists. The question arrives through the client relationship | Append-only decision log | Partly. An evidence trail is table stakes. A point-in-time snapshot is [?] |
| R3 · teach the agent | 2 | 3 | Under satisficing this happens only if it is cheap. Tuning quality is what lets him widen latitude | Verdict, reason routed to tuning | Partly. Dropzone shows a context-memory update; whether a human rejection routes to tuning is [?] |
| R4 · tell the client | 2 | 3 | She writes it, he owns it, and it is the revenue lever | Client summary | Partly. At Expel the client sees the same screen as the analyst |
| Main P2 · where latitude was earned | [?] | 3 | Whether an analyst needs the fleet at all is unverified. For him it is the job | Fleet view, Autonomy state, Review lane | Narrowed. See the correction below |
Emotional and social jobs are not scored here. They are satisfied by how the functional jobs are done rather than by functions of their own, and a row with no function cell would only be pretending to be a job.
05 · The MVP core
Two qualified. The third is a bet, and it is named as one
| In the core | The cell that proves it |
|---|---|
| Main | P1 = 3, competitors = No |
| R1, the shift | P1 = 3, competitors = No. Verified across all thirteen products, as is the residual-queue gap behind the main job |
The fleet view of trust enters the core as a bet, not on importance to the primary persona, which is [?]. A core of only Main and R1 is a good residual queue with a good handoff, and the market gap does not distinguish that product. The differentiation rests on the fleet, so it enters the core and is named as the bet it is.
This is also H1, the riskiest hypothesis in the project. If H1 falls, this is the job that leaves the core.
Narrowed a second time at Step 6. Simbian publishes “per-tenant autonomy configuration” with four named modes that progress per action and per alert type, and says latitude is earned in those words. So earned autonomy per tenant is not the gap. What nothing found publishes is the fleet as a readable operator surface: every tenant’s current latitude and accuracy trend legible at one glance where the analyst works, rather than on a configuration page. That is a design claim, not a capability claim, and it puts the whole weight on stages 04 and 07.
Why the rest are out, and why little is lost
- R3 does not need its own slot. Its function is the Verdict, already in the core through the main job, so design principle 3 does not end up without a feature behind it
- R2 likewise: the append-only log is a compliance requirement and ships regardless
- R4 is the only one genuinely deferred. The client summary is a revenue lever rather than something she does every shift
Orphan check, a different input
Not derived from the matrix. The eight planned features checked one by one against the jobs above. No orphans: all eight map to a formulated job.
What survives of the finding, smaller and true. The review lane maps to Tomas’s job on the audit side: sampling what the agent closed alone is how a record earns latitude. Its analyst side, an operator checking work she never saw, maps to no job anybody formulated. Either that job is written at stage 03a, or the analyst-side review lane is not built.
Idle control, reported honestly. This check returned zero. A check that returns zero proves nothing about itself. The only thing that tested it was a false positive it produced, and the critique caught that rather than the check.
06 · Critique and risks
Three instruments, twenty findings, one overlap
Claude ran a pass on broken chains from data to conclusion. Codex ran read-only against the research, looking for claims without support, drift from the source, and orphans in both directions. Neither saw the other’s table before both were complete. A third pass read the stage contract itself, because neither of the first two can see a step that never happened.
The divergences, which are what a second instrument is for
| Line | Claude when writing | Codex | Resolution |
|---|---|---|---|
| “It is the first and last screen of her shift” | A consequence of the 24/7 and remote data | No support. The research says nothing about first or last screens | Codex right |
| “69% report metrics manually. This is the person who does it” | Sourced from SANS | Figure absent from the research file; nothing says the lead does it | Both halves stand |
| “62% say their organisation is not doing enough” | Sourced from SANS | Figure absent from the research file | Real number, missing chain |
| R4 “without my rewriting it… does not cost more than doing it” | Grounded in the beneficiary | Neither motivation exists in the source | Codex right |
The heaviest finding, and it was ours
R1 cited the UCL handover study for the phrase “you spend your first hour rediscovering context someone else already had”. That phrase does not appear in the paper. It came from a search-result summary of a blog that was never opened, and it was signed with a source that had been. The job survives on three real quotes from the same study; the attribution did not. This is the rule of evidence failing in its purest form, and the failure is recorded above the job rather than quietly deleted.
What each instrument could not see
| Instrument | Found | Blind to |
|---|---|---|
| Claude, chains from data to conclusion | 10 | Line-by-line drift from the source. Needs no memory of what was meant |
| Codex, read-only against the source | 10 | Anything needing a reading of what the product is for. Also missed two cross-file contradictions inside its own radius |
| The stage contract as a checklist | 4 not done, 1 knowingly deferred | Defects in work that did happen |
One overlap out of twenty. All nineteen distinct findings were re-read in the file before being touched, all nineteen held, and all nineteen were fixed. None was dropped on verification, which is not a result to celebrate: it means the writing had that many real defects in it.
The three questions that would close the most dangerous gaps
| Question | What it decides | Status after Step 6 |
|---|---|---|
| Does an analyst carrying many tenants actually consult a cross-tenant view of trust? | Whether the third job in the MVP core exists at all. This is H1 | Open. Needs a working analyst; not closable from public sources |
| What does a Tier-2 analyst read in a queue row before opening it? | Behaviour 1, which decides what the first glance must deliver | Partly. Defender engineers the row name to carry scope, which is the industry’s belief, not a measurement of people |
| Is the person who moves autonomy a different behaviour, or a permission level? | Whether Tomas is a persona or a role | Open, and leaning toward permission. Authority sits with L3 analysts at Simbian |