Stage 01 · Foundation research

Harrier

An adjudication console for analysts at managed detection and response providers. One analyst carries 40 or more client tenants. An AI agent named Clerk assembles the case and files a draft verdict. The human rules on it.

Competitors read live
13 across three groups
Screens captured
20 all consumed as evidence
Benchmark, outside security
4 products, 8 criteria
Critique findings
17 two instruments

Every competitor and benchmark page was opened live on 20 August 2026. Nothing here comes from model memory. Unverified claims carry [?]. Harrier is an invented product, so its operating premises are marked PREMISE: they are decisions, not measurements.

01 · Introduction

Five findings, and what they changed in the brief

FINDING 01

The category has one position and it is unanimous. Simbian: “your analysts review the 8% that need judgment, not the 92% that don’t.” Intezer: “human SOC teams review outcomes, not tickets.” Dropzone: “100% Software Execution. No vendor-provided human safety drivers.” Every one of them sells the number of cases the human no longer sees.

FINDING 02

Nobody shows the queue that remains. The residual work after auto-resolution is where every hard decision lives and where the shift is spent. No vendor selling auto-resolution publishes a view of it. The only working analyst queue published across the thirteen is Expel’s, and Expel is human-led.

FINDING 03

Multi-tenancy here means isolation or switching, never a fleet. Simbian isolates tenant data and markets one rate across all clients. Expel makes the tenant a dropdown in the top bar. Prophet is single-tenant by design and says so.

FINDING 04

Earned autonomy already exists, and that narrowed our claim. Prophet sells “Autonomy on your terms,” widened when the track record justifies it. So earned autonomy is not the differentiator. What nobody does is make earned trust per tenant and visible across a fleet.

FINDING 05

Confidence is a model property inside security and an operator property outside it. Alloy’s own customer quote: the AI “gives me the confidence to quickly review and close.” Sift calls its version “Clearbox control.” None of the five direct competitors said it that way.

What the research changed in the brief

In the original briefAfter research
Earned autonomy as the differentiatorCorrected. Prophet already sells it. The claim narrows to earned autonomy per tenant, visible across the fleet
Priced per analyst seat plus per monitored assetChallenged, not settled. Dropzone charges per investigation volume, up to 4,000 per year per AI analyst. Intezer charges per endpoint with no volume fees. The market prices machine work
“Clerk shows its work or it shows nothing”Sharpened. Every competitor publishes an evidence trail, so a trail is table stakes. The differentiator is that the number names its claim, scope and window
Confidence as a model propertyReversed. Confidence is what the operator gets. The interface exists for that one
Nothing about shift handoffAdded. No product across the thirteen addresses what an incoming analyst is told. That this is the riskiest hour is our hypothesis [?]; that it is unserved is verified

Lean UX Canvas

The whole strategy on one sheet

Format: Lean UX Canvas v2 by Jeff Gothelf. Detail in lean-ux-canvas.md.

1 · Business problem

An MDR provider cannot safely give an agent more work without a way to show, per client, how much latitude it has earned and what it did with it. Until that exists, growth is capped by how much risk a service delivery lead will carry blind.

2 · Business outcomes

+40% tenants per analyst at a flat reversal rate. 60% of tenants above entry autonomy after 90 days. 90% of client escalations answered by the original evidence. Median time to verdict under 4 minutes. 50% of rejections producing a tuning change within 14 days. All targets are hypotheses.

3 · Users

Tier-2 SOC analyst at an MDR provider, 40+ tenants, 6+ hours a day in the tool. SOC lead who owns SLA and autonomy. Detection engineer who consumes rejections. The client’s security contact, who never logs in.

4 · User outcomes

Know if this is normal at this client. See how hard Clerk looked before deciding how hard to look. Disagree in one action and have it matter. Answer in April about a decision made in February with the case file, not memory.

5 · Solutions

Fleet view. Cross-tenant case queue. Case file with narrative, evidence, per-tenant base rate and provenance. One-key verdict. Client summary drafted by Clerk. Per-tenant autonomy shown as armed and active. Review lane for autonomous closes. Append-only decision log.

6 · Hypotheses

We believe an analyst will carry more tenants without more reversals if they get context on an unfamiliar client without leaving the case, with a per-tenant base rate in the case header. Five more in the same form.

7 · Riskiest assumption

An analyst carrying 40 tenants makes faster and better-defended decisions when the agent’s latitude varies per client than when it is one flat policy. If varying trust is just 40 more things to hold in working memory, the differentiator becomes overhead and Harrier is a worse Simbian. A value risk, not a feasibility risk.

8 · First test

Prototype, no engineering. Eight to ten MDR analysts, two conditions, twelve cases from five fictional tenants. Hide the queue after 60 seconds and ask which clients Clerk may close on its own. Then measure time to verdict and how many verdicts they reverse once shown the full evidence.

02 · Strategy

One dimension the product has to be best at

Strategic dimension

Calibrated trust in an automated agent. The operator knows exactly how much to trust Clerk right now, on this tenant, and that trust is earned and visible rather than asserted. Stage 04 must carry it with a concrete element on the reference screen. Stage 07 must check it did not dissolve.

Defensibility of a decision and speed under volume both matter, but we treat both as consequences: an analyst is fast because they trust the case file in front of them, and a verdict is defensible because the evidence behind it was visible when it was made. That causal chain is reasoning, not a finding [?], and it is the argument for preferring this dimension over the other two.

Audience

Primary: Tier-2 SOC analyst at an MDR provider. Ten-hour shifts on two monitors, 40 or more tenants. Paid to be fast, judged on being right, slowed down by the fear of the one true positive closed as noise. PREMISE: no analyst has been interviewed. Stage 02 builds the real persona.

Secondary: SOC lead, who owns SLA and the decision to give Clerk more rope. Tertiary: detection engineer, who acts on rejection reasons. Non-user beneficiary: the client’s security contact, whose trust is built almost entirely out of summaries they receive, which is why the summary is a product surface and not an export.

03 · AARRR

One metric and one product decision per stage

Activation node · First Verdict

The analyst rules on a Clerk-assembled case, with the evidence in view, and files it. The atomic moment the product exists to produce, and the node the stage 03a user flow has to reach as directly as possible.

StageOne metricTargetOne product decision
AcquisitionBake-offs entered[?] baseline firstReplay. Rebuild the residual queue from 30 days of the prospect’s own alert history
ActivationFirst verdict within 30 min of first login. The threshold is arbitrary until a real distribution replaces it80%The first case is not live. A replayed case with a known outcome, so the first act of trust is checkable
RetentionFour-week analyst retention85%Shift handoff. What moved, what waits, which tenants changed autonomy, what Clerk closed unwatched
RevenueNet revenue retention on assets added125%Tenant trust report, white-labelled, showing both sides of the ledger including where latitude fell
ReferralEvaluations naming an existing customer30%Shareable case file. Redacted permalink, with the redaction visible in the artifact

Every stage of this funnel is answered by the same object, the case file. Replay shows a prospect a queue of them, activation is ruling on one, the trust report aggregates them, referral forwards one. That is where the centre of gravity sits, and it tells stage 03a that the detail screen is not beside the dashboard but underneath everything.

04 · Competitors

Thirteen products, and only two readable interfaces

Collection guardrail: public and pre-login pages only, no accounts created. All five direct competitors keep their working console behind login, so interface evidence comes from documentation and help centres. Two working interfaces were readable anywhere: Expel Workbench and the PagerDuty Operations Console.

Comparison matrix

SimbianProphetExpelCursorDatadog SIEM
Core objectAlert auto-resolvedAlert investigated end to endInvestigation → incidentTask with a diffSignal → case
Where the human sitsReviews the residual 8%Approves scope, not each caseWorks the case, client observesRules on every diffTriages every signal
How autonomy is setOne public rate. Per-tenant policy inside is [?]Per organisation, earnedAuto remediations per orgPer actionPer rule, with suppressions
How agent work is shown[? behind login]Every question and query documentedA log entry, with a checkbox to hide itVerb trail, diff, measured outcomeRule as readable clauses
What is charged for[?][?][?][?][?]

Evidence

Expel Workbench investigation screen
Expel WorkbenchTenant as a top-bar dropdown. Row columns describe a record, not a shape. And a checkbox: “Hide Ruxie actions.”
PagerDuty Operations Console
PagerDuty consoleLive as an explicit toggle. Named saved views. The active filter spelled out in editable chips. A “Last Note” column.
Simbian MSSP and MDR page
SimbianPer-tenant Context Lake, reasoning never crosses tenants, 92% auto-resolved, “review the 8% that need judgment.”
Prophet Security AI SOC analyst page
Prophet Security“Autonomy on your terms.” Scope widens when the track record justifies it. “Nothing is learned silently.”
Dropzone AI SOC analyst page
Dropzone AI“Glass Box, Not Black Box.” A worked case ending in a benign dismissal and a context-memory update.
Dropzone pricing page
Dropzone pricingPriced by investigation capacity, up to 4,000 per year per AI analyst. MSSP plan is a dedicated multi-tenant environment.
Intezer AI SOC page
Intezer“Human SOC teams review outcomes, not tickets.” Endpoint pricing, no volume fees. Trust framed as a threshold to cross.
Expel Workbench product page
Expel Workbench pageRuxie enriches every alert across 160+ tools before it reaches the queue. The client sees the same screen as the analyst.
Sift platform page
Sift“Clearbox control.” Evidence as named chips under “Why Sift decided.” One score routed to Allow, Step-up, Block.
Alloy Actionable AI page
AlloyConfidence framed as what the operator gets: “gives me the confidence to quickly review and close.”
Intercom Copilot page
Intercom CopilotEvery generated answer links its top sources so agents validate in the inbox, not in an audit tab.
Linear home page
Linear“Reviews” as a top-level destination. Agent actions in the shared feed with a named actor. “Worked for 7s.”
Cursor desktop interface
CursorA queue titled “READY FOR REVIEW 5.” A size measure per row. A compact verb trail. An outcome reported as a measurement.
Datadog Cloud SIEM signal panel
Datadog Cloud SIEMBase rate in the header. “What happened” first, in prose. A fixed “Next steps” rail splitting triage from action.
Superhuman home page
SuperhumanSplit Inbox, Daily Briefs, and inline suggestions that underline without taking the turn.
Prophet Security home page
Prophet, homeDeploys as a dedicated single-tenant environment, so multi-tenancy is explicitly not their shape.

Three shared patterns, three differences

PATTERN

The evidence trail is the product. “Glass Box, Not Black Box,” “Clearbox control,” every query documented, the rule rendered as readable clauses. Nobody in this set defends an unexplained verdict.

PATTERN

The list narrows rather than browses. PagerDuty spells the filter out in editable chips, Superhuman shapes the inbox before the operator arrives, Cursor names the list after the decision it wants.

PATTERN

Volume is the sales unit. 92% auto-resolved, 91% noise cut, 85% less manual investigation, 99.9% fewer leads, 70% fewer manual reviews.

DIFFERENCE

Multi-tenancy is isolation or switching, never a fleet. Nobody shows all clients at once as one working surface.

DIFFERENCE

Two opposite answers to agent output. Linear puts agent actions in the shared feed with a named actor. Expel gives the analyst a checkbox to hide them.

DIFFERENCE

Uncertainty is a number or a question. Security reports a confidence score. Cursor stops and asks a numbered question with options.

05 · Benchmark

Calibrated trust, judged outside security entirely

Four products from fields that have been calibrating human trust in automation on real people for decades. A fifth candidate, clinical triage AI, was opened and dropped: the public pages carry solution marketing and no examinable mechanism.

CriterionWaymoNWS forecastAviation FMAStockfish
Stated scope of competence5543
Mode legibility2353
Evidence on demand3425
Calibration published5513
Graceful handoff1252
Cost of override1355
Failure disclosure4442
Trust moves with evidence5412
SKYbrary article on the Flight Mode Annunciator
Flight Mode AnnunciatorArmed and active as two simultaneous colour-coded states in a fixed position. And OVRD: override as its own annunciated mode.
National Weather Service page on probability of precipitation
NWS probabilityA probability defined by three mandatory qualifiers: what event, at which point, over which window. Nothing less can be checked.
Waymo safety page
Waymo safetyEvery percentage paired with an absolute count: “94% fewer serious injury or worse crashes” beside “47 FEWER.”
Lichess analysis board with engine evaluation
Stockfish on LichessOne strip: estimate, engine identity, effort spent, provenance. +0.2 · SF 18 dev 85MB NNUE · Depth 75 · CLOUD.

Three mechanisms into the MVP

MECHANISM 01 · FROM AVIATION

Armed and active, in a fixed place, with override as a named state. Per-tenant autonomy renders two states, always in the same position: what Clerk is permitted to do here, and what it is doing now. Override becomes its own annunciated state, not the quiet absence of automation.

Why it works: mode confusion is a display failure, not a knowledge failure. Operators lose track of which mode is live because the state was inferable rather than readable. Aviation learned this by crashing aircraft.

MECHANISM 02 · FROM NWS AND WAYMO

A number that names its claim, scope and window, paired with an absolute count. Not “87% confident,” but “at Meridian Dental, 9 of the last 11 cases matching this pattern in 30 days were benign.”

Why it works: a bare probability cannot be wrong, so it cannot earn trust either. The count also defeats the small-sample illusion: 90% reads identically on 9 cases and on 900.

MECHANISM 03 · FROM LICHESS

A provenance strip: what produced this, how hard it looked, where it came from. Model, sources queried, time spent. Read once, ignored thereafter, present when it matters.

Why it works: high confidence from one source in two seconds and high confidence from six sources over four minutes deserve different attention, and no confidence score distinguishes them.

And one that will not work here. Waymo earns trust by removing the human, one bounded geography at a time. That works because the geography is mapped, one operator owns all liability, and the rider has no decision to make. Our operator answers to 40 clients with different risk appetites, the environments change weekly, and the premise is that a human still rules. Take Waymo’s honesty about a published track record. Do not take its endgame.

06 · Patterns

Five structurally different answers, one chosen

The key task: rule on cases the agent has already assembled, fast enough to keep up and well enough to defend, across forty clients whose normal is different.

PatternWhere it is usedWhen it breaksVerdict
Split-pane reviewMail clients, GitHub PR review, Datadog signal panel, Expel WorkbenchContent that genuinely needs full width. Small screens, where the split stops being itselfCHOSEN
Focused card stackRadiology worklists, moderation queues, spaced repetitionThe moment the operator needs to defer, skip or compare. It removes any sense of the wholeAlternative
Fleet map, drill-downDatadog host maps, NOC walls, air traffic controlSerial work. It answers where and then abandons the operator at the moment of decidingRejected
Conversational workspaceDropzone’s chatbot, general assistantsRepetitive adjudication. No scannable state, no keyboard rhythm, and it inverts who is workingRejected as spine
Command-driven consoleLinear’s command menu, Superhuman, developer toolsWhen the operator does not know what to ask for, which in triage is the whole problemKept as accelerator
The choice

Split-pane review, with the fleet as the resting state of the detail pane. A cross-tenant queue on the left that never leaves the screen. The right side holds the case when one is selected, and the fleet, meaning per-tenant autonomy state and accuracy trend, when nothing is. The empty state of the detail pane is the dashboard. One pattern, two screens.

  1. It matches the entry-point behaviour. The analyst pattern-matches first and reads second. A persistent list beside the detail lets them confirm or break a match against what is around it, then move on without rebuilding context.
  2. It makes cheap override structurally possible. One-key disagreement needs a keyboard rhythm: move, read, rule, move. Split-pane is the only one of the five where the list survives the decision, so the rhythm survives it too.
  3. It is the structural expression of the gap. Expel makes the tenant a dropdown. Simbian sells one policy. A cross-tenant list persisting beside the detail is what “one fleet, one queue” looks like in layout rather than in a sentence. The pattern is the argument.

What this rests on, honestly. Reason 1 rests on the pattern-matching behaviour and reason 2 on satisficing, and both behaviours are unverified inferences [?]. Only reason 3 stands on evidence collected this session. If stage 02 does not confirm those behaviours, the pattern has to be re-argued from reason 3 alone, which it can survive, but as a narrower argument. Stage 04 should not inherit this as settled.

07 · Conclusions

Gaps, six hypotheses, eight open questions

The openings we take

GapEvidence
Nobody sells a fleet view of trustSimbian isolates tenant data and markets one 92% rate across all clients; whether a per-tenant trust level exists inside the product is [?]. Prophet earns autonomy per organisation but deploys single-tenant. The gap is in what is sold and shown publicly, which is all a pre-login pass can establish
Everyone optimises for the human seeing less; nobody shows the moment they still have to decideFive vendors sell 92%, 91%, 85%, 99.9% and 70% reductions. None publishes a view of what remains after them
Base rate is per environment and should be per tenantDatadog puts “Past month signal count” in the signal header, answering “is this normal here” for one organisation. In a 40-tenant console that question has 40 answers
The accept, edit or reject atom has not moved from code review into securityCursor: a queue named for the decision, a size measure before opening, a compact work trail, a numbered fork when unsure. Security still ships confidence percentages
Agent output is treated as clutter rather than as a draftExpel’s log carries a “Hide Ruxie actions” checkbox. Linear puts agent actions in the shared feed with a named actor
Nobody addresses the shift handoffNo page across the thirteen describes what an incoming analyst is told. Superhuman’s Daily Briefs is the nearest pattern, and it is not a security product

Six hypotheses

H1 · RISKIESTIf per-tenant autonomy is shown as armed and active state in a fixed position, then an analyst carrying 40 tenants will decide faster and more defensibly than under one flat policy.

Because mode confusion is a display failure rather than a knowledge failure, and the Flight Mode Annunciator solved it by making state readable rather than inferable. If H1 is false the idea falls: per-tenant trust becomes 40 things to remember and the differentiator becomes overhead.

H2 · CONDITIONALIf the case header carries a per-tenant base rate, then an analyst will reach a verdict on an unfamiliar client without a research detour.

Because Datadog already proves the mechanism for one environment, and the cost here is tenant switching, which resets what counts as normal. Half-grounded: the mechanism is evidenced, the need for it is not [?].

H3If every Clerk output carries a provenance strip naming model, sources queried and time spent, then analysts will allocate attention on evidence rather than on order of arrival.

Because effort spent and confidence are different questions, and only the first predicts how much checking is warranted. Grounded in the Lichess strip: Depth 75, CLOUD.

H4 · CONDITIONALIf rejecting Clerk costs one keystroke and the reason routes to detection engineering, then rejections will keep being filed rather than quietly absorbed.

Because analysts satisfice under volume and take the cheapest sufficient action, so the cheapest action has to be the useful one. Rests entirely on an unverified behaviour [?].

H5 · CONDITIONALIf the fleet view shows autonomy state and accuracy trend per tenant, then latitude will grow on measured accuracy rather than on a forgotten setting.

Because trust in automation is set by the last failure rather than the average, and latitude never seen to fall stops being believed [?].

H6If a shift handoff is composed at the end of every shift, then less information will be lost across a 24/7 rotation.

Because no product across the thirteen addresses the handoff at all and the nearest pattern comes from an email client. That the handoff is the riskiest hour is itself unverified [?] and is the first thing stage 02 should ask.

Open questions

QuestionAddresseeWhat changes in the productStatus
One cross-tenant queue, or is switching client context one at a time a safety feature?MDR analyst; stands to product owner until stage 02If switching is safety, the dashboard stops being a merged queue and becomes a fleet view that hands off into one tenantOpen
Who moves a tenant’s autonomy level: analyst, SOC lead, or the client?SOC leadDecides whether the autonomy control lives in the operator console at allOpen
Can a product that routes work back to humans price the way this market prices?Product ownerIf not, the business model line in CLAUDE.md is wrong and the product must justify what a human decision is worthOpen
Named saved views, or one opinionated ordering the product owns and explains?Product owner, at stage 03aViews put the burden of a good queue on the operator. One ordering means the product must be right and show its reasoning. This decides what the dashboard isOpen
Will an analyst accept liability for a summary Clerk wrote and they approved?SOC leadIf not, the summary becomes a structured record the analyst assembles from partsOpen
Autonomy as one slider, or three named lanes the way Sift routes Allow, Step-up, Block?Product owner, at wireframesA slider asks how much you trust it. Lanes say where the work goes, which is easier to hold under pressure and easier to auditOpen
Do analysts own specific tenants, or the queue as a whole?SOC leadDecides whether the shift handoff is composed per tenant or per shiftOpen
Do MDR providers treat their tooling as a competitive secret?Product ownerIf yes, peer referral is suppressed and the channel is the only route, which rewrites the Referral stageOpen

08 · Before and after

Seventeen findings, two instruments, zero overlap

Codex ran read-only over research/docs and returned twelve findings. A separate pass ran on a class Codex cannot reach, conclusions whose chain back to a fact is broken, and returned five. The two sets did not overlap once. Full log in docs/decisions.md.

TypeFoundBecameWho found itStatus
Contradiction between files“Simbian sells one 92% rate across all of them,” stated as fact in two places, while the level file recorded [?] on whether any competitor has per-tenant trust inside the product“What Simbian sells publicly is one rate. Whether the product carries a per-tenant trust level internally is [?]: the console is behind login”CodexFixed
Contradiction between files“No competitor page examined publishes so much as a screenshot of it”“No vendor selling auto-resolution publishes a view of what is left after it. The only working analyst queue published across the thirteen is Expel’s, and Expel is human-led”CodexFixed
Fact without a sourceThe analyst persona given as data: “26 to 40, two to six years, 10-hour shifts on two monitors”Marked PREMISE: “an assumed profile we chose, not a researched one. No analyst has been interviewed”CodexFixed
Fact without a source“Finishing four cases per hour perfectly while the queue grows by forty”Numbers removed. They were illustration wearing the clothes of measurementCodexFixed
Conclusion without groundsThe chosen UX pattern presented as settled, with three reasons of equal weight“Only reason 3 stands on evidence collected this session… Stage 04 should not inherit this as settled”ClaudeFixed
Conclusion without groundsThe 30-minute activation threshold, presented as a target“The 30-minute threshold is arbitrary. It was chosen because it is roughly one coffee, not derived from anything”ClaudeFixed
Fact without a sourceExpel’s 15-month retention, flagged as having no citation at the lineLeft as written. The source is the docs page cited in the same table row and the text is visible in the cited screenshot. Codex read a text snapshot and could not see the screenshotCodexDropped on verification

What the self-audit systematically misses. Four of Codex’s findings were unsourced premises of our own product: tenants per analyst, how MDR margin works, who carries liability, the analyst profile. They are invisible to a self-audit because they were decisions when written and read as context afterwards. Codex does not know they are our decisions, so it asks for a source the same way it would for a competitor’s price. That is the whole argument for a second instrument.

Updated after publication, at stage 03a. Building the information architecture surfaced a gap in the data rather than in a screen: both dead ends in the product hang on how long the evidence under a case remains retrievable, and no retention window was ever chosen. It bounds the 90% target in Strategy before any design decision touches it, and it decides whether the evidence snapshot is a copy or a reference. Two new open questions are in research.md, section 10.

09 · Follow-up research

Added at stage 02, before personas were drafted

Stage 01 collected evidence about the market and a frame about people. Four public sources opened on 2026-08-20 replaced part of that frame with data. Three of the four are studies of practitioners, two of them with interviews, so several lines that stood as PREMISE now stand on measurement, and two stage 01 decisions turn out to be wrong.

Sources

SourceWhat it isWhy it counts
Kent, Tuptuk, Becker, Passing the Baton, UCL, Jan 2026Interviews with six incident responders, five of them at organisations providing cybersecurity as a serviceThe only study found that examines the shift handoff directly, in our exact setting
SANS SOC Survey 2025, Christopher Crowley, July 2025Ninth annual survey of how SOCs are built, staffed and runIndustry-scale numbers on coverage, tenure, outsourcing and satisfaction with tooling
Rastogi et al., Too Much to Trust?, RIT, ACM CCS 2025Survey of 248 analysts plus 24 in-depth interviewsThe largest body of analyst voice found anywhere, on exactly our question: what makes an AI explanation trustworthy under time pressure
Microsoft Defender multitenant management, updated Aug 2026Documentation of a shipping cross-tenant consoleThe platform most MDR providers actually operate on

Premise replaced by measurement

LineWasNow
A 24/7 rotationPREMISE79% of SOCs are operational 24/7
Two to six years in operationsPREMISEThree to five years is the most common tenure. 31% stay three to five years, 4% stay ten or more
MDR providers carry other companies’ triagePREMISE183 of 443 respondents outsource alerting, meaning triage and escalation, fully or partially
The analyst arrives trusting the agentnever statedFalse. AI/ML tools rank at the bottom of the SOC satisfaction list. Two of three AI/ML technologies measured ranked at the very bottom, generative language tools scored 2 out of 4, and 42% of SOCs use AI/ML out of the box with no customization. EDR/XDR is the only technology above 3 out of 4

The last row outranks the rest. Harrier’s operator does not arrive curious about Clerk. They arrive having already been sold an AI tool that underdelivered, and the tool they trust most is the boring one that works. Every trust mechanism in this product is arguing with that, not building on a blank slate.

What the analyst wants from an explanation

The headline of the RIT study: participants “were consistently willing to accept XAI outputs, even in cases of lower predictive accuracy, when explanations were perceived as relevant and evidence-backed”. Trust follows evidence quality, not the accuracy number.

FindingWhat it does to our design
A confidence score alone does not work. An explanation is more meaningful “if I have some context about how that percentage was generated”, such as which log data supports a 92% confidenceDesign principle 2 in CLAUDE.md, almost word for word, now stated by an analyst instead of inferred by us
More explanation is not more trust. No rationale gives low trust, a minimal clear one gives a large gain, and further detail has diminishing returns and “can even reduce trust if they introduce confusion or doubt”Sets a ceiling on Clerk’s evidence panel. Showing the work has an optimum, not a maximum
Organisational context is the named gap: a detection platform “doesn’t have the organizational perspective… if that is there then it is like wonders”This is the evidence H2 was missing. The mechanism was proven at Datadog; the need for it is now stated by analysts
Analysts think in incident narratives while tools explain single alerts: “we have to find a connection between the critical and the high alerts to determine if it’s an incident”That correlation is what Clerk is for. The closest thing to a validation of the product premise found in any source
Explanation depth varies by experience, not only by role. Access level changes what a person can see at allA behavioural split, not a demographic one. Personas have to decide whether it makes a second person

Two stage 01 decisions this breaks

DecisionWhat the evidence saysNow
Shift handoff composed at the end of every shiftThe end of the shift is exactly the failure mode. “Over-utilised analysts are just gonna be ready to just get out and head home. So they just wanna get it done fast, and rush.” Content goes stale inside the shift: “at 9am you’ve got something to put in the handover. By 9:30, that might have changed”Composed continuously through the shift and closed at the end. One participant had already solved it this way
The handoff is a documentFour of six participants described the same division: detail lives in the ticket and the handover points at it. “We put loads of ticket references in… Handovers are more for signposting”The handoff is signposting into case files, not prose

The five behaviours, restated

BehaviourStatus
1. Pattern-matching before readingStill [?]. The only one still holding a design decision on nothing
2. Satisficing under volumeSupported, by quotes in both studies
3. Trust set by the last failureStill [?] as stated, but the starting level is now known and it is low
4. Tenant switching resets contextSupported in structure, not in cost. Microsoft carries Tenant name as a row column, yet opens real work “in a new tab for that tenant” and refuses to assign across tenants
5. Writing for the future auditorCorrected. The reader is the next analyst, not an auditor

Still open. Who moves a tenant’s autonomy level, whether an analyst will accept liability for a summary Clerk wrote, and what tenant switching costs in seconds or errors. Forty tenants per analyst stays PREMISE: no public source found gives an analyst-to-client ratio at an MDR provider.

11 · The language of the category

Added at stage 05. Nine pages opened live on 2026-08-23

Updated after publication. Stage 01 asked what competitors do. Stage 05 re-read the same pages asking how they say it, which is a different reading and it produced a different answer. Full section in research/docs/research.md §11.

The source did not fall away, and it did not fully arrive either. Working consoles sit behind login across all five direct competitors, so the only interface language readable without an account is Expel’s, published in its own documentation. Everything else here is marketing language and is labelled as such.

Who the category addresses, and it is not the operator

VendorLine, verbatimAddressed to
Simbian“Scale Your MSSP Practice. Not Your Payroll.”The owner of the P&L
Simbian“The last analyst you’ll hire.”The owner of the P&L
Prophet“Free your SOC Analysts to focus on real threats”The manager, about the analyst
Dropzone“Your team can’t investigate every alert. Dropzone AI can.”The manager of the team
Dropzone“REINFORCEMENTS HAVE ARRIVED”Nobody in particular
Intezer“human SOC teams review outcomes, not tickets”The buyer, about the team

Not one of them speaks to the person who will sit in front of the screen. The category’s second person is the buyer. This is the sharpest finding of the pass and the cheapest to act on.

Evidence

Prophet Security AI SOC analyst page, read 2026-08-23
Prophet“Autonomy on your terms.” “Nothing is learned silently.” Their word for a verdict is determination.
Simbian MSSP and MDR page, read 2026-08-23
SimbianEvery line addressed to the P&L. And the one noun we share: the agent “applies the verdict.”
Dropzone AI SOC analyst page, read 2026-08-23
DropzoneA worked case with four numbered findings and one conclusion sentence: “Accepted behavior due to scheduled backup and requires no further action.”

What the incumbent’s own documentation gives

TAKE

Verb plus object, imperative. Expel’s nine auto remediations: Block Bad Hashes, Contain Hosts, Deactivate Access Keys, Delete Malicious Files, Disable Accounts, Kill Processes, Remove Malicious Email, Reset Credentials. Harrier’s action classes already read this way, so the rule is confirmed rather than invented.

TAKE

The record is answers to questions. Expel: “A Finding is where our SOC analysts document answers to questions like: What is it? Where is it? When did it get here? How did it get here?” Harrier’s block headings are the same move, and it is a pattern with a precedent, not an invention.

GAP

Disagreement is a free-text comment. Expel closes an investigation in five steps and the reason is step five: “You should be sure to leave a comment.” No reason taxonomy. Harrier’s 4.4 requires one of six and prints where each routes.

GAP

Two axes compressed into one label. Expel’s Verify Action: Not Authorized: Incident, Not Malicious: Close, Authorized: Close. The category has the problem; nobody has separated the axes.

GAP

Routing is stated once, off screen. “We may eventually create a Suppression for it” lives in a documentation aside. This product prints agent tuning, detection engineering, the tenant baseline, locked on the control itself.

POSITION

Ten invented nouns against one. Expel teaches its customer Assembler, DUET, Detection Strategy, Event, Expel Alert, Lead Alert, Finding, Remediation Action, Suppression, Auto Remediation. Harrier names exactly one thing: Clerk.

The vocabulary decision this settles

WordWho has settled itHarrier
verdictSimbian, verbatimUses it. Not a coinage
determinationProphet, verbatimDoes not use it. Same word, and the product needs one
findingExpel, verbatimDoes not use it. Harrier says signal and claim
contain, escalate, evidenceUniversalUses all three
latitudeNobody. Prophet says scope, Simbian says authority, Expel says accessHarrier’s own word, and the only invented term in the product. It carries the differentiator, so it earns its keep