Skip to content
CWS
CorovaAboutContact
Book a Call
All articles
Service Delivery

The Board Wants AI in the SOC. Start With the Queue Nobody Wants.

The instruction arrives with a date attached, and the demonstrations all point at the same work: the incident narrative, the executive summary, the parts of the job that read as judgment. Those are the hardest places to establish whether it worked. The first use case worth running is the one where somebody can check the answer before lunch.

CWSAugust 14, 20266 min read

The board is asking a fair question

A board reads what everybody reads, sits with people who sit on other boards, and carries responsibility for whether the organization is keeping pace. Asking what AI is doing inside the security operation is well within that remit. An answer starting with the word evaluating holds for about two meetings.

What comes back is a shortlist assembled from demonstrations, and a demonstration is built to show capability at its most impressive. Part of that shortlist will be autonomous alert triage, which several vendors now lead with, and that is exactly the kind of checkable work argued for below. The rest of it skews toward work that shows well: a model reading a messy investigation and producing a clean narrative, or summarizing an incident for an executive. Those are real capabilities too, and they sit in the part of the job where two experienced people disagree about what good looks like.

The shortlist then gets ordered by visibility, and the pilot starts on whichever item would impress most if it worked. Six weeks later somebody has to stand up and say whether it did. Nobody asked at the start how anyone would tell.

Part of that comes from the question being heard as a technical one. Taken plainly, a board is asking four things, and a particular capability answers none of them on its own.

  • Whether the organization is keeping pace with how this work gets done elsewhere.
  • Whether the money already committed to the security operation produces more than it did last year.
  • Whether the risk of adopting this is held by somebody with a name.
  • Whether the next twelve months produce a result or a further update.

The visible use case has no answer key

Take the incident narrative. A model reads the case and writes the summary, and somebody senior reads the summary. Was it right? One reviewer says the tone is wrong for the audience. Another says it left out the detail they would have opened with. A third says it beats what the team produces at five o'clock on a Friday. All three are reading honestly, because the work has no single correct output and never did when people did it by hand.

So a pilot on that work ends in a room, with senior people disagreeing about quality and no record of what the right answer would have been. It gets settled by whoever is most senior or most enthusiastic, and the organization comes out holding an opinion.

The cost of a wrong output runs the wrong way here too. A narrative that quietly omits something gets believed, because it is the artifact people read in place of the case. Four things about that pilot stay unanswerable for as long as it runs.

  • What the correct output would have been, written down before the model produced one.
  • How often the output was wrong, expressed as a rate somebody can quote.
  • Who checked it, against what, and how long one check took them.
  • What a wrong output cost, and how anybody came to notice.
A pilot on work with no right answer ends in a meeting where the most senior person present decides whether it worked.

Choose the first one for how cheaply you can check it

The property that makes a first use case useful has little to do with how impressive it is. It is whether somebody can establish, quickly and without an argument, that the output was right. That property belongs to the work, so you can assess it before anyone demonstrates anything.

The platform will run whatever use case you point it at. Choosing which one goes first is a judgment about your own queues, and it is made with what your analysts already know about them. Products built specifically for alert triage arrive claiming the checkable property directly, and the claim can be sound. It turns into evidence on the day it holds on your queue, against records your own people already keep, on a sample they drew.

Say the criteria out loud in front of a shortlist and the order changes on the spot. What scores highest looks like a queue arriving every day in the same shape, where an experienced person confirms or overturns an answer in two minutes using records that already exist.

This is a considered position and a recommendation, and it is the one we would defend in front of a skeptical operations lead. Score every candidate against the six things below.

  • Volume. A month of running should produce enough items that a sample describes the queue. Nine cases describe nine cases.
  • Repeatability. The same question, in the same shape, with the same kinds of evidence attached. Variation is what makes a wrong answer impossible to attribute afterward.
  • A checkable answer. Somebody establishes what the right answer was, in minutes, from records held today. Where that takes most of a working day, the checking quietly stops.
  • The cost of being wrong. A wrong answer here should produce rework or mild irritation for a colleague. Work where it can end in a missed intrusion belongs further down the list.
  • Speed of discovery. The error surfaces inside a period you can state in advance, and undoing it is an afternoon's task.
  • A record that already exists. Timestamps, dispositions, ticket history, who touched the item and when. Where the work leaves no trail, your first month produces impressions.
The first use case is a measuring instrument. Choose it for what it can tell you, and choose it before you choose a tool.

Read your own queues

The candidates are already in your own queues, and they are unglamorous by construction, because the glamorous work turns on judgment and judgment is what resists checking. Start with the queues nobody volunteers to own: user-reported suspicious messages, findings that have to reach whichever team owns the asset, cases describing one event three times over, evidence requests an auditor will accept or send back.

CWS has taken three Cortex XSIAM deployments through post-deployment work, and the case for a checkable first use case comes out of that. The queues that produced a verifiable answer were not the ones anybody demonstrated.

The governance frameworks converge on this without using these words. Stating the intended use of a system and the consequence of a wrong output, before it goes near production, is what they ask a deployer to do. Run that across four candidate queues honestly and the selection makes itself, because three of them produce a paragraph nobody is comfortable signing.

The exercise takes an afternoon and it belongs in a room with the people who work those queues. They know which one has a clean answer behind it and which one carries an argument running since 2023. Write five lines about each candidate before anybody opens a product page.

  • Items per week, read off the record itself.
  • Where the right answer comes from, and the name of the person or system holding it.
  • How long one check takes for somebody who knows the work well.
  • What a wrong answer costs, and the longest it could plausibly go unnoticed.
  • Whether the people working it would hand it over tomorrow, which tells you what the first month of adoption will feel like.

What goes back to the board

Pressure from above to show AI in the security operation is fair, and a twelve month evaluation answers it poorly. A result answers it, and results come from work that can be graded. Our view, offered as a view: look first at the queue that has sat on a list of things to fix for three years, worked by whoever drew the short straw. It has volume, it has a right answer behind it, and nobody minds much if it gets two of them wrong in the first month. Start there.

The visible use case survives. It moves to second and arrives in better shape for having waited, because by then the organization has a sampling routine and a route by which a disagreement becomes a change with an owner and a date. Who answers for a wrong automated answer, and how far its authority runs, are the questions that follow selection. Each is easier to settle for the first time on a queue where being wrong costs an hour.

The two choices also diverge at the meeting. A pilot on the visible work produces a demonstration and a recommendation, and the room is invited to take both on trust. A pilot on a checkable queue produces a sample anybody can inspect, and a sample can be questioned, which is what makes it evidence. Six lines carry the whole readout.

  • The queue, the volume over the period, and why this one went first.
  • The sample that was read, how it was drawn, and who read it.
  • The rate at which the automated answer and the human check agreed.
  • Every disagreement, and what happened to it.
  • What the organization now knows that it did not know three months ago, in one paragraph.
  • The second candidate, and the date it starts.
Bring the board a sample they can inspect and a rate they can question. A demonstration asks them to take your word for it.

Sources

  • CWS delivery corpus, three Cortex XSIAM deployments taken through post-deployment work
  • NIST AI Risk Management Framework and ISO/IEC 42001, on stating intended use and the consequence of a wrong output before deployment
  • Selecting a first AI use case by the cost of verifying its output: CWS opinion, stated here so it can be argued with.