Stage 02 · User research

What they are trying to get done

One main job, four adjacent ones, and a matrix that answers two different questions: what to build first, and what not to build at all. Every job is solution-agnostic, so a completely different product could be hired for it.

Main jobs in the product
1 a second is recorded, not merged
Related jobs
4 a fifth was not written
MVP core
3 two qualified, one is a bet
Critique, two instruments
20 findings, one overlap

Wording here is canonical. personas.md carries the same formulations word for word, verified after the critique changed two of them. A job that exists in two editions is a job that gets built twice.

01 · Who owns the steps

A decision recorded before the jobs

Related jobs are separate adjacent tasks done in the same context. The steps of the journey toward the main job are normally CJM phases, and this project runs a shortened track where CJM is not included.

Decision

The steps of the journey are owned by the user flow at stage 03a. That flow already has to describe the path to the activation node, First Verdict, so it is the natural holder. The related jobs below are genuinely adjacent tasks, and stage 03a should not read them as a sequence.

02 · The main job

One, and it belongs to Rasha

Main · P1
When Clerk hands me a case it has already investigated, I want to decide whether its verdict holds, so that the decision is made and I can still defend it months later.

Solution-agnostic check. No function named. A peer review process, a checklist or a second analyst reading over a shoulder could all be hired for this. The job does not lock the answer.

P2 has a different main job, and it is recorded rather than merged

Main · P2
When I am accountable for what the agent did across all my clients, I want to know where its record has earned more latitude and where it has lost it, so that I can widen or narrow its scope without guessing.

Two main jobs is a signal of two products, so it is named here rather than glued into one formulation.

The reading taken. The product has one main job, Rasha’s, because she is primary and the activation node is First Verdict. Tomas’s job is real and is served by the same substrate read at a different scale: she reads one case, he reads the trail across many. It stays one product for as long as the record underneath them is one record.

The alternative, rejected with a reason. Treating them as co-equal makes the information architecture build two navigations, and the product splits into a case console and a fleet console that happen to share a database.

04 · The matrix

Jobs against personas, functions and the market

Importance is inferred, not measured. No importance number exists in the research, so each row names the signal it rests on. Where there is no signal the cell is [?] rather than an average.

JobP1P2Signal behind the importanceFunctionCompetitors
Main · rule on the case31Her whole shift, and the activation node. He reads outcomes rather than adjudicating each caseCase file, Verdict, Autonomy stateNo. Nobody publishes the residual queue
R1 · the shift32An incoming team must read six reports to hold three days of context, and older ones stop being read. 79% of SOCs run 24/7empty in the canonical listNo. Verified across thirteen products
R2 · answer for a past decision23Rare per case, but it is why the record exists. The question arrives through the client relationshipAppend-only decision logPartly. An evidence trail is table stakes. A point-in-time snapshot is [?]
R3 · teach the agent23Under satisficing this happens only if it is cheap. Tuning quality is what lets him widen latitudeVerdict, reason routed to tuningPartly. Dropzone shows a context-memory update; whether a human rejection routes to tuning is [?]
R4 · tell the client23She writes it, he owns it, and it is the revenue leverClient summaryPartly. At Expel the client sees the same screen as the analyst
Main P2 · where latitude was earned[?]3Whether an analyst needs the fleet at all is unverified. For him it is the jobFleet view, Autonomy state, Review laneNarrowed. See the correction below

Emotional and social jobs are not scored here. They are satisfied by how the functional jobs are done rather than by functions of their own, and a row with no function cell would only be pretending to be a job.

05 · The MVP core

Two qualified. The third is a bet, and it is named as one

In the coreThe cell that proves it
MainP1 = 3, competitors = No
R1, the shiftP1 = 3, competitors = No. Verified across all thirteen products, as is the residual-queue gap behind the main job
The third slot, entered as an exception

The fleet view of trust enters the core as a bet, not on importance to the primary persona, which is [?]. A core of only Main and R1 is a good residual queue with a good handoff, and the market gap does not distinguish that product. The differentiation rests on the fleet, so it enters the core and is named as the bet it is.

This is also H1, the riskiest hypothesis in the project. If H1 falls, this is the job that leaves the core.

Narrowed a second time at Step 6. Simbian publishes “per-tenant autonomy configuration” with four named modes that progress per action and per alert type, and says latitude is earned in those words. So earned autonomy per tenant is not the gap. What nothing found publishes is the fleet as a readable operator surface: every tenant’s current latitude and accuracy trend legible at one glance where the analyst works, rather than on a configuration page. That is a design claim, not a capability claim, and it puts the whole weight on stages 04 and 07.

Why the rest are out, and why little is lost

  • R3 does not need its own slot. Its function is the Verdict, already in the core through the main job, so design principle 3 does not end up without a feature behind it
  • R2 likewise: the append-only log is a compliance requirement and ships regardless
  • R4 is the only one genuinely deferred. The client summary is a revenue lever rather than something she does every shift

Orphan check, a different input

Not derived from the matrix. The eight planned features checked one by one against the jobs above. No orphans: all eight map to a formulated job.

What survives of the finding, smaller and true. The review lane maps to Tomas’s job on the audit side: sampling what the agent closed alone is how a record earns latitude. Its analyst side, an operator checking work she never saw, maps to no job anybody formulated. Either that job is written at stage 03a, or the analyst-side review lane is not built.

Idle control, reported honestly. This check returned zero. A check that returns zero proves nothing about itself. The only thing that tested it was a false positive it produced, and the critique caught that rather than the check.

06 · Critique and risks

Three instruments, twenty findings, one overlap

Claude ran a pass on broken chains from data to conclusion. Codex ran read-only against the research, looking for claims without support, drift from the source, and orphans in both directions. Neither saw the other’s table before both were complete. A third pass read the stage contract itself, because neither of the first two can see a step that never happened.

The divergences, which are what a second instrument is for

LineClaude when writingCodexResolution
“It is the first and last screen of her shift”A consequence of the 24/7 and remote dataNo support. The research says nothing about first or last screensCodex right
“69% report metrics manually. This is the person who does it”Sourced from SANSFigure absent from the research file; nothing says the lead does itBoth halves stand
“62% say their organisation is not doing enough”Sourced from SANSFigure absent from the research fileReal number, missing chain
R4 “without my rewriting it… does not cost more than doing it”Grounded in the beneficiaryNeither motivation exists in the sourceCodex right

The heaviest finding, and it was ours

R1 cited the UCL handover study for the phrase “you spend your first hour rediscovering context someone else already had”. That phrase does not appear in the paper. It came from a search-result summary of a blog that was never opened, and it was signed with a source that had been. The job survives on three real quotes from the same study; the attribution did not. This is the rule of evidence failing in its purest form, and the failure is recorded above the job rather than quietly deleted.

What each instrument could not see

InstrumentFoundBlind to
Claude, chains from data to conclusion10Line-by-line drift from the source. Needs no memory of what was meant
Codex, read-only against the source10Anything needing a reading of what the product is for. Also missed two cross-file contradictions inside its own radius
The stage contract as a checklist4 not done, 1 knowingly deferredDefects in work that did happen

One overlap out of twenty. All nineteen distinct findings were re-read in the file before being touched, all nineteen held, and all nineteen were fixed. None was dropped on verification, which is not a result to celebrate: it means the writing had that many real defects in it.

The three questions that would close the most dangerous gaps

QuestionWhat it decidesStatus after Step 6
Does an analyst carrying many tenants actually consult a cross-tenant view of trust?Whether the third job in the MVP core exists at all. This is H1Open. Needs a working analyst; not closable from public sources
What does a Tier-2 analyst read in a queue row before opening it?Behaviour 1, which decides what the first glance must deliverPartly. Defender engineers the row name to carry scope, which is the industry’s belief, not a measurement of people
Is the person who moves autonomy a different behaviour, or a permission level?Whether Tomas is a persona or a roleOpen, and leaning toward permission. Authority sits with L3 analysts at Simbian