The First Ninety Days of Alert Data Decide the Next Three Years
The alert stream after go-live gets handled one item at a time and then filed. Read once as a whole it describes the estate, the detections and the people working them, and it is CWS's position that the decisions taken from it in the first weeks are the ones still running three years later.
The alert stream is a record, and it gets filed unread
Cortex XSIAM stitches related issues, its term for detections that have fired, into cases by causality, and scores each case for urgency. What reaches an analyst is a case the platform has already assembled and scored. The case gets an owner and a state, and it gets closed. One at a time is how a case has to be handled. One at a time is also how the whole stream gets handled, with a tuning ticket raised off whichever case annoyed somebody that week.
Taken over a run of weeks those cases record two things. One half describes the estate: which systems produce the most, which sources are talkative, which detections have never fired. The other half describes the operation: which classes close in seconds, which wait for the second shift, what somebody suppressed in week three to survive a difficult fortnight. Both halves land in the same datasets the platform writes as it runs, and both wait for someone to ask.
A new deployment arrives with the vendor's half already built: Analytics BIOCs that Palo Alto Networks builds and maintains, content packs from the Marketplace carrying correlation rules, parsing rules, playbooks and dashboards, all written to hold true across every estate that runs them. That is what a platform of this class is for. The estate-specific half, the IOC, BIOC and correlation rules written against your naming, your crown jewels and your business hours, gets built after the data starts arriving, by the people who know which systems the business cannot afford to have quiet.
Treated as a workload the stream produces tuning tickets. Read once as evidence it produces an account of where the volume sits and which detections the team has already stopped believing.
Who has read the first ninety days as one document? The platform answers in XQL. Asking is the part that has to be scheduled.
The questions the distribution answers
The picture a security leader carries of the first months is assembled from anecdote, because anecdote is what reaches them. The bad Tuesday. The source somebody remembers as noisy. Every one of those is accurate, and every one is a single item out of thousands.
The stream answers better, and on this platform the read is a query. Issue and case data lands in datasets, XQL, the Cortex Query Language, reaches it, and the dashboards that ship with the product already carry part of what follows: volume by source, closure rates, coverage across the Analytics BIOCs. The questions worth an afternoon are the ones that cross those views and end up describing the organization.
A closure habit shows up in the data long before anyone will describe it in a meeting, and so does a detection the team has stopped acting on. The organization's own work has already answered these.
- Which sources account for most of what arrives, counted at the collector and the Broker VM as well as at the case, and whether the systems underneath them are the ones the business would name first.
- Which detections fire and close with no action taken, separated into the Analytics BIOCs that shipped with the platform and the rules your own team wrote. Trust gets withdrawn from those quietly, and the data shows it before anyone says it.
- Which classes of case close quickly with no note attached. A class being cleared looks different in the data from a class being worked.
- When the queue backs up, and whether those hours line up with a shift boundary, a batch window or one person's calendar.
- How long a case waits for its first human touch, sorted by class and set against the urgency score the platform gave it. Where the two disagree, the team's own ranking is showing.
- Every suppression and exclusion added since go-live: who added it, what it covers, and what reason was recorded at the time.
Why the early defaults outlive everyone who set them
Reading the post-deployment distribution at a fixed point, and the claim that early configuration decisions are the durable ones, are CWS positions. The argument rests on one mechanism: a running security operation leaves a settled decision settled, and the review that would reopen it has to be scheduled from outside.
A detection switched off in a noisy first month stays off. A detection that has gone quiet is silent in exactly the way a healthy one is, and the next engineer to open the configuration reads what they find as something somebody intended. A suppression written to make the first month survivable was the right call that month, and it holds until a person retires it. Nobody is scheduled to.
Habits are more durable still, because they pass by apprenticeship. An analyst joining in year two learns the queue from the analysts already on it, and that transfer happens in conversation, where it stays out of reach of any review against how the estate looks now. What is left in a configuration three years on is a short list.
- Suppressions and exclusions carrying no expiry date and no author.
- Detections disabled during the first weeks, with the reasoning in a chat thread that has since aged out.
- Thresholds set to make the volume survivable while the estate was still being fitted, never argued with once it settled.
- A triage order that formed in the first busy fortnight and became the way the queue is worked.
- Classes of case the team learned to skim, passed on since to every analyst hired.
A suppression written in week three runs until somebody is given the job of retiring it. The job needs a name on it and a date.
What a healthy ninety day picture looks like
Volume is the wrong test. A healthy ninety day stream and an unhealthy one can carry the same amount of alert. What separates them is whether somebody can account for the shape, and whether the changes made to it carry a name and a date.
In the healthy version the sources at the top of the distribution are there because somebody chose to prioritize those systems and can say why. The detections that have produced nothing are known, and are being worked through with a decision recorded against each. Suppressions have authors and expiry dates. Closure notes vary, because the cases varied. The unhealthy version shows up in the same places.
- One or two sources dominate what arrives, and the ranking got set by whichever collector was onboarded first.
- A large part of the detection set has produced no action since go-live, and naming those entries requires going to look.
- Closure notes identical across a whole class of case.
- The hours when alerts wait longest map onto one person's availability.
- Two people in the same team describe the stream differently, and both are describing memory.
One read, on a date somebody put in the calendar
Tuning is already continuous, and it should be. What this asks for is different in kind: one deliberate read of the whole stream, at a fixed point, with the people who can act on it in the room.
The read needs a stretch of settled data behind it, long enough to hold a quiet week and a bad one. It also needs a date. A standing intention slips a month at a time, and the event that would force it sits outside the queue entirely. Ninety days is roughly the earliest point where there is enough behind you and the defaults are still soft.
Who sits in the room decides what comes out: the analysts who work the queue, the engineer who made the early changes, and whoever can say which systems the business cannot afford to have quiet. Miss the last one and the read produces a tuning backlog the team already had.
- Every detection that has produced nothing, with either a reason to keep it or a date to retire it.
- Every suppression written since go-live, given an author, a reason and an expiry. The ones arriving without an author get one now.
- A ranking of the classes of case, agreed with the people who own the systems underneath them and set against the platform's own urgency scoring, written down so the next analyst inherits it in text.
- A named owner for the shape of the stream, separate from whoever owns the queue.
- The date of the next read, and the person who books it.
The record that composed itself
CWS's post-deployment work on three Cortex XSIAM deployments started, each time, from what the alert stream already said. The estates had little in common and the answers differed in each. The questions put to the stream were the same, and in every case it showed the team something they half knew and had never seen stated.
A status report is written by somebody with a view about how the deployment went. The alert data accumulated as a byproduct of the work, which is what makes it credible and what makes it uncomfortable.
The platform half of this arrived complete. A platform built to generalize, then fitted to one estate from the evidence that estate produces, is the division of labor working as designed. The fitting is yours, and the first ninety days are when it is cheapest.
Read once and deliberately, the ninety day picture turns into a set of decisions with names and dates on them. Left unread, it hardens into the configuration, and the configuration is what the next three years run on.
The stream is the only account of those first months that nobody composed.