Stage 01 · Foundation research
Harrier
An adjudication console for analysts at managed detection and response providers. One analyst carries 40 or more client tenants. An AI agent named Clerk assembles the case and files a draft verdict. The human rules on it.
- Competitors read live
- 13 across three groups
- Screens captured
- 20 all consumed as evidence
- Benchmark, outside security
- 4 products, 8 criteria
- Critique findings
- 17 two instruments
Every competitor and benchmark page was opened live on 20 August 2026. Nothing here comes from model memory. Unverified claims carry [?]. Harrier is an invented product, so its operating premises are marked PREMISE: they are decisions, not measurements.
01 · Introduction
Five findings, and what they changed in the brief
The category has one position and it is unanimous. Simbian: “your analysts review the 8% that need judgment, not the 92% that don’t.” Intezer: “human SOC teams review outcomes, not tickets.” Dropzone: “100% Software Execution. No vendor-provided human safety drivers.” Every one of them sells the number of cases the human no longer sees.
Nobody shows the queue that remains. The residual work after auto-resolution is where every hard decision lives and where the shift is spent. No vendor selling auto-resolution publishes a view of it. The only working analyst queue published across the thirteen is Expel’s, and Expel is human-led.
Multi-tenancy here means isolation or switching, never a fleet. Simbian isolates tenant data and markets one rate across all clients. Expel makes the tenant a dropdown in the top bar. Prophet is single-tenant by design and says so.
Earned autonomy already exists, and that narrowed our claim. Prophet sells “Autonomy on your terms,” widened when the track record justifies it. So earned autonomy is not the differentiator. What nobody does is make earned trust per tenant and visible across a fleet.
Confidence is a model property inside security and an operator property outside it. Alloy’s own customer quote: the AI “gives me the confidence to quickly review and close.” Sift calls its version “Clearbox control.” None of the five direct competitors said it that way.
What the research changed in the brief
| In the original brief | After research |
|---|---|
| Earned autonomy as the differentiator | Corrected. Prophet already sells it. The claim narrows to earned autonomy per tenant, visible across the fleet |
| Priced per analyst seat plus per monitored asset | Challenged, not settled. Dropzone charges per investigation volume, up to 4,000 per year per AI analyst. Intezer charges per endpoint with no volume fees. The market prices machine work |
| “Clerk shows its work or it shows nothing” | Sharpened. Every competitor publishes an evidence trail, so a trail is table stakes. The differentiator is that the number names its claim, scope and window |
| Confidence as a model property | Reversed. Confidence is what the operator gets. The interface exists for that one |
| Nothing about shift handoff | Added. No product across the thirteen addresses what an incoming analyst is told. That this is the riskiest hour is our hypothesis [?]; that it is unserved is verified |
Lean UX Canvas
The whole strategy on one sheet
Format: Lean UX Canvas v2 by Jeff Gothelf. Detail in lean-ux-canvas.md.
1 · Business problem
An MDR provider cannot safely give an agent more work without a way to show, per client, how much latitude it has earned and what it did with it. Until that exists, growth is capped by how much risk a service delivery lead will carry blind.
2 · Business outcomes
+40% tenants per analyst at a flat reversal rate. 60% of tenants above entry autonomy after 90 days. 90% of client escalations answered by the original evidence. Median time to verdict under 4 minutes. 50% of rejections producing a tuning change within 14 days. All targets are hypotheses.
3 · Users
Tier-2 SOC analyst at an MDR provider, 40+ tenants, 6+ hours a day in the tool. SOC lead who owns SLA and autonomy. Detection engineer who consumes rejections. The client’s security contact, who never logs in.
4 · User outcomes
Know if this is normal at this client. See how hard Clerk looked before deciding how hard to look. Disagree in one action and have it matter. Answer in April about a decision made in February with the case file, not memory.
5 · Solutions
Fleet view. Cross-tenant case queue. Case file with narrative, evidence, per-tenant base rate and provenance. One-key verdict. Client summary drafted by Clerk. Per-tenant autonomy shown as armed and active. Review lane for autonomous closes. Append-only decision log.
6 · Hypotheses
We believe an analyst will carry more tenants without more reversals if they get context on an unfamiliar client without leaving the case, with a per-tenant base rate in the case header. Five more in the same form.
7 · Riskiest assumption
An analyst carrying 40 tenants makes faster and better-defended decisions when the agent’s latitude varies per client than when it is one flat policy. If varying trust is just 40 more things to hold in working memory, the differentiator becomes overhead and Harrier is a worse Simbian. A value risk, not a feasibility risk.
8 · First test
Prototype, no engineering. Eight to ten MDR analysts, two conditions, twelve cases from five fictional tenants. Hide the queue after 60 seconds and ask which clients Clerk may close on its own. Then measure time to verdict and how many verdicts they reverse once shown the full evidence.
02 · Strategy
One dimension the product has to be best at
Calibrated trust in an automated agent. The operator knows exactly how much to trust Clerk right now, on this tenant, and that trust is earned and visible rather than asserted. Stage 04 must carry it with a concrete element on the reference screen. Stage 07 must check it did not dissolve.
Defensibility of a decision and speed under volume both matter, but we treat both as consequences: an analyst is fast because they trust the case file in front of them, and a verdict is defensible because the evidence behind it was visible when it was made. That causal chain is reasoning, not a finding [?], and it is the argument for preferring this dimension over the other two.
Audience
Primary: Tier-2 SOC analyst at an MDR provider. Ten-hour shifts on two monitors, 40 or more tenants. Paid to be fast, judged on being right, slowed down by the fear of the one true positive closed as noise. PREMISE: no analyst has been interviewed. Stage 02 builds the real persona.
Secondary: SOC lead, who owns SLA and the decision to give Clerk more rope. Tertiary: detection engineer, who acts on rejection reasons. Non-user beneficiary: the client’s security contact, whose trust is built almost entirely out of summaries they receive, which is why the summary is a product surface and not an export.
03 · AARRR
One metric and one product decision per stage
The analyst rules on a Clerk-assembled case, with the evidence in view, and files it. The atomic moment the product exists to produce, and the node the stage 03a user flow has to reach as directly as possible.
| Stage | One metric | Target | One product decision |
|---|---|---|---|
| Acquisition | Bake-offs entered | [?] baseline first | Replay. Rebuild the residual queue from 30 days of the prospect’s own alert history |
| Activation | First verdict within 30 min of first login. The threshold is arbitrary until a real distribution replaces it | 80% | The first case is not live. A replayed case with a known outcome, so the first act of trust is checkable |
| Retention | Four-week analyst retention | 85% | Shift handoff. What moved, what waits, which tenants changed autonomy, what Clerk closed unwatched |
| Revenue | Net revenue retention on assets added | 125% | Tenant trust report, white-labelled, showing both sides of the ledger including where latitude fell |
| Referral | Evaluations naming an existing customer | 30% | Shareable case file. Redacted permalink, with the redaction visible in the artifact |
Every stage of this funnel is answered by the same object, the case file. Replay shows a prospect a queue of them, activation is ruling on one, the trust report aggregates them, referral forwards one. That is where the centre of gravity sits, and it tells stage 03a that the detail screen is not beside the dashboard but underneath everything.
04 · Competitors
Thirteen products, and only two readable interfaces
Collection guardrail: public and pre-login pages only, no accounts created. All five direct competitors keep their working console behind login, so interface evidence comes from documentation and help centres. Two working interfaces were readable anywhere: Expel Workbench and the PagerDuty Operations Console.
Comparison matrix
| Simbian | Prophet | Expel | Cursor | Datadog SIEM | |
|---|---|---|---|---|---|
| Core object | Alert auto-resolved | Alert investigated end to end | Investigation → incident | Task with a diff | Signal → case |
| Where the human sits | Reviews the residual 8% | Approves scope, not each case | Works the case, client observes | Rules on every diff | Triages every signal |
| How autonomy is set | One public rate. Per-tenant policy inside is [?] | Per organisation, earned | Auto remediations per org | Per action | Per rule, with suppressions |
| How agent work is shown | [? behind login] | Every question and query documented | A log entry, with a checkbox to hide it | Verb trail, diff, measured outcome | Rule as readable clauses |
| What is charged for | [?] | [?] | [?] | [?] | [?] |
Evidence
















Three shared patterns, three differences
The evidence trail is the product. “Glass Box, Not Black Box,” “Clearbox control,” every query documented, the rule rendered as readable clauses. Nobody in this set defends an unexplained verdict.
The list narrows rather than browses. PagerDuty spells the filter out in editable chips, Superhuman shapes the inbox before the operator arrives, Cursor names the list after the decision it wants.
Volume is the sales unit. 92% auto-resolved, 91% noise cut, 85% less manual investigation, 99.9% fewer leads, 70% fewer manual reviews.
Multi-tenancy is isolation or switching, never a fleet. Nobody shows all clients at once as one working surface.
Two opposite answers to agent output. Linear puts agent actions in the shared feed with a named actor. Expel gives the analyst a checkbox to hide them.
Uncertainty is a number or a question. Security reports a confidence score. Cursor stops and asks a numbered question with options.
05 · Benchmark
Calibrated trust, judged outside security entirely
Four products from fields that have been calibrating human trust in automation on real people for decades. A fifth candidate, clinical triage AI, was opened and dropped: the public pages carry solution marketing and no examinable mechanism.
| Criterion | Waymo | NWS forecast | Aviation FMA | Stockfish |
|---|---|---|---|---|
| Stated scope of competence | 5 | 5 | 4 | 3 |
| Mode legibility | 2 | 3 | 5 | 3 |
| Evidence on demand | 3 | 4 | 2 | 5 |
| Calibration published | 5 | 5 | 1 | 3 |
| Graceful handoff | 1 | 2 | 5 | 2 |
| Cost of override | 1 | 3 | 5 | 5 |
| Failure disclosure | 4 | 4 | 4 | 2 |
| Trust moves with evidence | 5 | 4 | 1 | 2 |




+0.2 · SF 18 dev 85MB NNUE · Depth 75 · CLOUD.Three mechanisms into the MVP
Armed and active, in a fixed place, with override as a named state. Per-tenant autonomy renders two states, always in the same position: what Clerk is permitted to do here, and what it is doing now. Override becomes its own annunciated state, not the quiet absence of automation.
Why it works: mode confusion is a display failure, not a knowledge failure. Operators lose track of which mode is live because the state was inferable rather than readable. Aviation learned this by crashing aircraft.
A number that names its claim, scope and window, paired with an absolute count. Not “87% confident,” but “at Meridian Dental, 9 of the last 11 cases matching this pattern in 30 days were benign.”
Why it works: a bare probability cannot be wrong, so it cannot earn trust either. The count also defeats the small-sample illusion: 90% reads identically on 9 cases and on 900.
A provenance strip: what produced this, how hard it looked, where it came from. Model, sources queried, time spent. Read once, ignored thereafter, present when it matters.
Why it works: high confidence from one source in two seconds and high confidence from six sources over four minutes deserve different attention, and no confidence score distinguishes them.
And one that will not work here. Waymo earns trust by removing the human, one bounded geography at a time. That works because the geography is mapped, one operator owns all liability, and the rider has no decision to make. Our operator answers to 40 clients with different risk appetites, the environments change weekly, and the premise is that a human still rules. Take Waymo’s honesty about a published track record. Do not take its endgame.
06 · Patterns
Five structurally different answers, one chosen
The key task: rule on cases the agent has already assembled, fast enough to keep up and well enough to defend, across forty clients whose normal is different.
| Pattern | Where it is used | When it breaks | Verdict |
|---|---|---|---|
| Split-pane review | Mail clients, GitHub PR review, Datadog signal panel, Expel Workbench | Content that genuinely needs full width. Small screens, where the split stops being itself | CHOSEN |
| Focused card stack | Radiology worklists, moderation queues, spaced repetition | The moment the operator needs to defer, skip or compare. It removes any sense of the whole | Alternative |
| Fleet map, drill-down | Datadog host maps, NOC walls, air traffic control | Serial work. It answers where and then abandons the operator at the moment of deciding | Rejected |
| Conversational workspace | Dropzone’s chatbot, general assistants | Repetitive adjudication. No scannable state, no keyboard rhythm, and it inverts who is working | Rejected as spine |
| Command-driven console | Linear’s command menu, Superhuman, developer tools | When the operator does not know what to ask for, which in triage is the whole problem | Kept as accelerator |
Split-pane review, with the fleet as the resting state of the detail pane. A cross-tenant queue on the left that never leaves the screen. The right side holds the case when one is selected, and the fleet, meaning per-tenant autonomy state and accuracy trend, when nothing is. The empty state of the detail pane is the dashboard. One pattern, two screens.
- It matches the entry-point behaviour. The analyst pattern-matches first and reads second. A persistent list beside the detail lets them confirm or break a match against what is around it, then move on without rebuilding context.
- It makes cheap override structurally possible. One-key disagreement needs a keyboard rhythm: move, read, rule, move. Split-pane is the only one of the five where the list survives the decision, so the rhythm survives it too.
- It is the structural expression of the gap. Expel makes the tenant a dropdown. Simbian sells one policy. A cross-tenant list persisting beside the detail is what “one fleet, one queue” looks like in layout rather than in a sentence. The pattern is the argument.
What this rests on, honestly. Reason 1 rests on the pattern-matching behaviour and reason 2 on satisficing, and both behaviours are unverified inferences [?]. Only reason 3 stands on evidence collected this session. If stage 02 does not confirm those behaviours, the pattern has to be re-argued from reason 3 alone, which it can survive, but as a narrower argument. Stage 04 should not inherit this as settled.
07 · Conclusions
Gaps, six hypotheses, eight open questions
The openings we take
| Gap | Evidence |
|---|---|
| Nobody sells a fleet view of trust | Simbian isolates tenant data and markets one 92% rate across all clients; whether a per-tenant trust level exists inside the product is [?]. Prophet earns autonomy per organisation but deploys single-tenant. The gap is in what is sold and shown publicly, which is all a pre-login pass can establish |
| Everyone optimises for the human seeing less; nobody shows the moment they still have to decide | Five vendors sell 92%, 91%, 85%, 99.9% and 70% reductions. None publishes a view of what remains after them |
| Base rate is per environment and should be per tenant | Datadog puts “Past month signal count” in the signal header, answering “is this normal here” for one organisation. In a 40-tenant console that question has 40 answers |
| The accept, edit or reject atom has not moved from code review into security | Cursor: a queue named for the decision, a size measure before opening, a compact work trail, a numbered fork when unsure. Security still ships confidence percentages |
| Agent output is treated as clutter rather than as a draft | Expel’s log carries a “Hide Ruxie actions” checkbox. Linear puts agent actions in the shared feed with a named actor |
| Nobody addresses the shift handoff | No page across the thirteen describes what an incoming analyst is told. Superhuman’s Daily Briefs is the nearest pattern, and it is not a security product |
Six hypotheses
H1 · RISKIESTIf per-tenant autonomy is shown as armed and active state in a fixed position, then an analyst carrying 40 tenants will decide faster and more defensibly than under one flat policy.
Because mode confusion is a display failure rather than a knowledge failure, and the Flight Mode Annunciator solved it by making state readable rather than inferable. If H1 is false the idea falls: per-tenant trust becomes 40 things to remember and the differentiator becomes overhead.
H2 · CONDITIONALIf the case header carries a per-tenant base rate, then an analyst will reach a verdict on an unfamiliar client without a research detour.
Because Datadog already proves the mechanism for one environment, and the cost here is tenant switching, which resets what counts as normal. Half-grounded: the mechanism is evidenced, the need for it is not [?].
H3If every Clerk output carries a provenance strip naming model, sources queried and time spent, then analysts will allocate attention on evidence rather than on order of arrival.
Because effort spent and confidence are different questions, and only the first predicts how much checking is warranted. Grounded in the Lichess strip: Depth 75, CLOUD.
H4 · CONDITIONALIf rejecting Clerk costs one keystroke and the reason routes to detection engineering, then rejections will keep being filed rather than quietly absorbed.
Because analysts satisfice under volume and take the cheapest sufficient action, so the cheapest action has to be the useful one. Rests entirely on an unverified behaviour [?].
H5 · CONDITIONALIf the fleet view shows autonomy state and accuracy trend per tenant, then latitude will grow on measured accuracy rather than on a forgotten setting.
Because trust in automation is set by the last failure rather than the average, and latitude never seen to fall stops being believed [?].
H6If a shift handoff is composed at the end of every shift, then less information will be lost across a 24/7 rotation.
Because no product across the thirteen addresses the handoff at all and the nearest pattern comes from an email client. That the handoff is the riskiest hour is itself unverified [?] and is the first thing stage 02 should ask.
Open questions
| Question | Addressee | What changes in the product | Status |
|---|---|---|---|
| One cross-tenant queue, or is switching client context one at a time a safety feature? | MDR analyst; stands to product owner until stage 02 | If switching is safety, the dashboard stops being a merged queue and becomes a fleet view that hands off into one tenant | Open |
| Who moves a tenant’s autonomy level: analyst, SOC lead, or the client? | SOC lead | Decides whether the autonomy control lives in the operator console at all | Open |
| Can a product that routes work back to humans price the way this market prices? | Product owner | If not, the business model line in CLAUDE.md is wrong and the product must justify what a human decision is worth | Open |
| Named saved views, or one opinionated ordering the product owns and explains? | Product owner, at stage 03a | Views put the burden of a good queue on the operator. One ordering means the product must be right and show its reasoning. This decides what the dashboard is | Open |
| Will an analyst accept liability for a summary Clerk wrote and they approved? | SOC lead | If not, the summary becomes a structured record the analyst assembles from parts | Open |
| Autonomy as one slider, or three named lanes the way Sift routes Allow, Step-up, Block? | Product owner, at wireframes | A slider asks how much you trust it. Lanes say where the work goes, which is easier to hold under pressure and easier to audit | Open |
| Do analysts own specific tenants, or the queue as a whole? | SOC lead | Decides whether the shift handoff is composed per tenant or per shift | Open |
| Do MDR providers treat their tooling as a competitive secret? | Product owner | If yes, peer referral is suppressed and the channel is the only route, which rewrites the Referral stage | Open |
08 · Before and after
Seventeen findings, two instruments, zero overlap
Codex ran read-only over research/docs and returned twelve findings. A separate pass ran on a class Codex cannot reach, conclusions whose chain back to a fact is broken, and returned five. The two sets did not overlap once. Full log in docs/decisions.md.
| Type | Found | Became | Who found it | Status |
|---|---|---|---|---|
| Contradiction between files | “Simbian sells one 92% rate across all of them,” stated as fact in two places, while the level file recorded [?] on whether any competitor has per-tenant trust inside the product | “What Simbian sells publicly is one rate. Whether the product carries a per-tenant trust level internally is [?]: the console is behind login” | Codex | Fixed |
| Contradiction between files | “No competitor page examined publishes so much as a screenshot of it” | “No vendor selling auto-resolution publishes a view of what is left after it. The only working analyst queue published across the thirteen is Expel’s, and Expel is human-led” | Codex | Fixed |
| Fact without a source | The analyst persona given as data: “26 to 40, two to six years, 10-hour shifts on two monitors” | Marked PREMISE: “an assumed profile we chose, not a researched one. No analyst has been interviewed” | Codex | Fixed |
| Fact without a source | “Finishing four cases per hour perfectly while the queue grows by forty” | Numbers removed. They were illustration wearing the clothes of measurement | Codex | Fixed |
| Conclusion without grounds | The chosen UX pattern presented as settled, with three reasons of equal weight | “Only reason 3 stands on evidence collected this session… Stage 04 should not inherit this as settled” | Claude | Fixed |
| Conclusion without grounds | The 30-minute activation threshold, presented as a target | “The 30-minute threshold is arbitrary. It was chosen because it is roughly one coffee, not derived from anything” | Claude | Fixed |
| Fact without a source | Expel’s 15-month retention, flagged as having no citation at the line | Left as written. The source is the docs page cited in the same table row and the text is visible in the cited screenshot. Codex read a text snapshot and could not see the screenshot | Codex | Dropped on verification |
What the self-audit systematically misses. Four of Codex’s findings were unsourced premises of our own product: tenants per analyst, how MDR margin works, who carries liability, the analyst profile. They are invisible to a self-audit because they were decisions when written and read as context afterwards. Codex does not know they are our decisions, so it asks for a source the same way it would for a competitor’s price. That is the whole argument for a second instrument.
Updated after publication, at stage 03a. Building the information architecture surfaced a gap in the data rather than in a screen: both dead ends in the product hang on how long the evidence under a case remains retrievable, and no retention window was ever chosen. It bounds the 90% target in Strategy before any design decision touches it, and it decides whether the evidence snapshot is a copy or a reference. Two new open questions are in research.md, section 10.
09 · Follow-up research
Added at stage 02, before personas were drafted
Stage 01 collected evidence about the market and a frame about people. Four public sources opened on 2026-08-20 replaced part of that frame with data. Three of the four are studies of practitioners, two of them with interviews, so several lines that stood as PREMISE now stand on measurement, and two stage 01 decisions turn out to be wrong.
Sources
| Source | What it is | Why it counts |
|---|---|---|
| Kent, Tuptuk, Becker, Passing the Baton, UCL, Jan 2026 | Interviews with six incident responders, five of them at organisations providing cybersecurity as a service | The only study found that examines the shift handoff directly, in our exact setting |
| SANS SOC Survey 2025, Christopher Crowley, July 2025 | Ninth annual survey of how SOCs are built, staffed and run | Industry-scale numbers on coverage, tenure, outsourcing and satisfaction with tooling |
| Rastogi et al., Too Much to Trust?, RIT, ACM CCS 2025 | Survey of 248 analysts plus 24 in-depth interviews | The largest body of analyst voice found anywhere, on exactly our question: what makes an AI explanation trustworthy under time pressure |
| Microsoft Defender multitenant management, updated Aug 2026 | Documentation of a shipping cross-tenant console | The platform most MDR providers actually operate on |
Premise replaced by measurement
| Line | Was | Now |
|---|---|---|
| A 24/7 rotation | PREMISE | 79% of SOCs are operational 24/7 |
| Two to six years in operations | PREMISE | Three to five years is the most common tenure. 31% stay three to five years, 4% stay ten or more |
| MDR providers carry other companies’ triage | PREMISE | 183 of 443 respondents outsource alerting, meaning triage and escalation, fully or partially |
| The analyst arrives trusting the agent | never stated | False. AI/ML tools rank at the bottom of the SOC satisfaction list. Two of three AI/ML technologies measured ranked at the very bottom, generative language tools scored 2 out of 4, and 42% of SOCs use AI/ML out of the box with no customization. EDR/XDR is the only technology above 3 out of 4 |
The last row outranks the rest. Harrier’s operator does not arrive curious about Clerk. They arrive having already been sold an AI tool that underdelivered, and the tool they trust most is the boring one that works. Every trust mechanism in this product is arguing with that, not building on a blank slate.
What the analyst wants from an explanation
The headline of the RIT study: participants “were consistently willing to accept XAI outputs, even in cases of lower predictive accuracy, when explanations were perceived as relevant and evidence-backed”. Trust follows evidence quality, not the accuracy number.
| Finding | What it does to our design |
|---|---|
| A confidence score alone does not work. An explanation is more meaningful “if I have some context about how that percentage was generated”, such as which log data supports a 92% confidence | Design principle 2 in CLAUDE.md, almost word for word, now stated by an analyst instead of inferred by us |
| More explanation is not more trust. No rationale gives low trust, a minimal clear one gives a large gain, and further detail has diminishing returns and “can even reduce trust if they introduce confusion or doubt” | Sets a ceiling on Clerk’s evidence panel. Showing the work has an optimum, not a maximum |
| Organisational context is the named gap: a detection platform “doesn’t have the organizational perspective… if that is there then it is like wonders” | This is the evidence H2 was missing. The mechanism was proven at Datadog; the need for it is now stated by analysts |
| Analysts think in incident narratives while tools explain single alerts: “we have to find a connection between the critical and the high alerts to determine if it’s an incident” | That correlation is what Clerk is for. The closest thing to a validation of the product premise found in any source |
| Explanation depth varies by experience, not only by role. Access level changes what a person can see at all | A behavioural split, not a demographic one. Personas have to decide whether it makes a second person |
Two stage 01 decisions this breaks
| Decision | What the evidence says | Now |
|---|---|---|
| Shift handoff composed at the end of every shift | The end of the shift is exactly the failure mode. “Over-utilised analysts are just gonna be ready to just get out and head home. So they just wanna get it done fast, and rush.” Content goes stale inside the shift: “at 9am you’ve got something to put in the handover. By 9:30, that might have changed” | Composed continuously through the shift and closed at the end. One participant had already solved it this way |
| The handoff is a document | Four of six participants described the same division: detail lives in the ticket and the handover points at it. “We put loads of ticket references in… Handovers are more for signposting” | The handoff is signposting into case files, not prose |
The five behaviours, restated
| Behaviour | Status |
|---|---|
| 1. Pattern-matching before reading | Still [?]. The only one still holding a design decision on nothing |
| 2. Satisficing under volume | Supported, by quotes in both studies |
| 3. Trust set by the last failure | Still [?] as stated, but the starting level is now known and it is low |
| 4. Tenant switching resets context | Supported in structure, not in cost. Microsoft carries Tenant name as a row column, yet opens real work “in a new tab for that tenant” and refuses to assign across tenants |
| 5. Writing for the future auditor | Corrected. The reader is the next analyst, not an auditor |
Still open. Who moves a tenant’s autonomy level, whether an analyst will accept liability for a summary Clerk wrote, and what tenant switching costs in seconds or errors. Forty tenants per analyst stays PREMISE: no public source found gives an analyst-to-client ratio at an MDR provider.
11 · The language of the category
Added at stage 05. Nine pages opened live on 2026-08-23
Updated after publication. Stage 01 asked what competitors do. Stage 05 re-read the same pages asking how they say it, which is a different reading and it produced a different answer. Full section in research/docs/research.md §11.
The source did not fall away, and it did not fully arrive either. Working consoles sit behind login across all five direct competitors, so the only interface language readable without an account is Expel’s, published in its own documentation. Everything else here is marketing language and is labelled as such.
Who the category addresses, and it is not the operator
| Vendor | Line, verbatim | Addressed to |
|---|---|---|
| Simbian | “Scale Your MSSP Practice. Not Your Payroll.” | The owner of the P&L |
| Simbian | “The last analyst you’ll hire.” | The owner of the P&L |
| Prophet | “Free your SOC Analysts to focus on real threats” | The manager, about the analyst |
| Dropzone | “Your team can’t investigate every alert. Dropzone AI can.” | The manager of the team |
| Dropzone | “REINFORCEMENTS HAVE ARRIVED” | Nobody in particular |
| Intezer | “human SOC teams review outcomes, not tickets” | The buyer, about the team |
Not one of them speaks to the person who will sit in front of the screen. The category’s second person is the buyer. This is the sharpest finding of the pass and the cheapest to act on.
Evidence



What the incumbent’s own documentation gives
Verb plus object, imperative. Expel’s nine auto remediations: Block Bad Hashes, Contain Hosts, Deactivate Access Keys, Delete Malicious Files, Disable Accounts, Kill Processes, Remove Malicious Email, Reset Credentials. Harrier’s action classes already read this way, so the rule is confirmed rather than invented.
The record is answers to questions. Expel: “A Finding is where our SOC analysts document answers to questions like: What is it? Where is it? When did it get here? How did it get here?” Harrier’s block headings are the same move, and it is a pattern with a precedent, not an invention.
Disagreement is a free-text comment. Expel closes an investigation in five steps and the reason is step five: “You should be sure to leave a comment.” No reason taxonomy. Harrier’s 4.4 requires one of six and prints where each routes.
Two axes compressed into one label. Expel’s Verify Action: Not Authorized: Incident, Not Malicious: Close, Authorized: Close. The category has the problem; nobody has separated the axes.
Routing is stated once, off screen. “We may eventually create a Suppression for it” lives in a documentation aside. This product prints agent tuning, detection engineering, the tenant baseline, locked on the control itself.
Ten invented nouns against one. Expel teaches its customer Assembler, DUET, Detection Strategy, Event, Expel Alert, Lead Alert, Finding, Remediation Action, Suppression, Auto Remediation. Harrier names exactly one thing: Clerk.
The vocabulary decision this settles
| Word | Who has settled it | Harrier |
|---|---|---|
| verdict | Simbian, verbatim | Uses it. Not a coinage |
| determination | Prophet, verbatim | Does not use it. Same word, and the product needs one |
| finding | Expel, verbatim | Does not use it. Harrier says signal and claim |
| contain, escalate, evidence | Universal | Uses all three |
| latitude | Nobody. Prophet says scope, Simbian says authority, Expel says access | Harrier’s own word, and the only invented term in the product. It carries the differentiator, so it earns its keep |