case studies

Incident Déjà Vu Case Study: Every Returning Incident Named Within Ten Seconds

When a service starts failing at night, the person on call wants to know one thing first: has this happened before, and what fixed it? The answer usually exists somewhere. It sits in an old ticket or in a colleague’s memory, and finding it takes a search that only works if you already know what to search for. The Incident Déjà Vu demo asks the log file itself, every ten seconds.

The run continues the invented web shop from the Logs demo. Its file already holds the shop’s day before, 24 hours from noon on 29 September 2026, with every kind of log line counted a minute at a time and six incidents the on-call team wrote down. Then two more hours play out, from 12:00 to 14:00 UTC on 30 September, with five incidents on the schedule. One of them is a kind the file has never seen.

The two hours of the Incident Déjà Vu demo on one line. Above, five incidents: northpay at 12:18, the database pool at 12:52, an expired certificate at 13:14, a login attack at 13:34 and a second expired certificate at 13:48. Below, what the file said: incident #1 after 8 seconds, #4 after 10, new after 5, #2 after 10, and #7 after 10 for the second certificate, which the on-call team wrote down at 13:22.

The two hours: what happened to the shop above, and what the file said at each check below.

The Set-Up

Incident déjà vu
The day before 2,133,322 log lines of 25 kinds from five services, and six incidents written down, in a 3.7 MB SQLite file
Today 242,635 lines in two hours, about 34 a second
Policy One stream from logs: every kind of line counted by service and level, a minute at a time for a week and an hour at a time for 90 days, and each line kept whole for ten minutes
The comparison Plain SQL in the same file: views over the counts, and a trigger that saves each incident’s fingerprint when it is written down
Checks Every ten seconds, right after the Engine writes the file
Runtime The log reducer and the Engine, compiled to WebAssembly, writing into SQLite 3.53.4’s WebAssembly build

What Happened

  • 12:18:12. Card payments start failing at northpay. At the first check, 8 seconds in, the file says it looks like incident #1, yesterday’s northpay outage, and shows what fixed it: moving card payments to quickcard. The quickcard outage from that morning scores almost as high, so the page reads the raw lines, and provider=northpay settles it.
  • 12:52:40. The checkout database runs out of connections. The file names #4, the pool exhaustion from 02:30, 10 seconds in.
  • 13:14:05. An image server’s certificate expires. Two kinds of line the file has never seen appear, and 5 seconds in the page says nothing in the file looks like this. At 13:22 the on-call team writes it down: an expired certificate, fixed by renewing it and turning on expiry alerts.
  • 13:34:30. A login attack. The file names #2, yesterday’s botnet, 10 seconds in.
  • 13:48:20. A second image server’s certificate expires. Ten seconds in, the file names #7, the incident written down 26 minutes earlier.

How the File Tells

The log reducer turns each line into a template, such as charge failed provider=<*> result=<*> status=<*> amount=<*>, and the Engine counts templates by service and level, minute by minute. Everything else is SQL over those counts.

A template stands out when its rate over the last minute or two is at least three times its usual minute, higher or lower, and at least six lines a minute away from it. Usual means the median minute of the hour before. The median matters here. An incident earlier in the hour would drag an average up, and the quiet after it would then look like an incident of its own. A template going quiet also has to stay quiet over one more minute, because silence takes longer to be sure of than noise.

Each template that stands out gets a score, the logarithm of how far it moved, and together the scores form a fingerprint. Writing an incident down saves the fingerprint of its first two minutes beside it, with the newest raw line of each template, while the raw lines are still kept. A trigger does that on the insert. The view deja_vu then compares the fingerprint of now with every saved one by the cosine of the scores. At 0.7 or more, now looks like that incident; at 0.4 or more, it is partly like it.

Two of the incidents written down were outages at northpay and at quickcard. They leave the same templates, so their fingerprints match about equally well. Templates hide the words that differ, but the raw lines still have them. When two incidents score within 0.05 of each other, the page compares the hidden words in the newest line of each template with the example saved with each incident. At 12:18 the newest failed charge said provider=northpay, and so did the example saved with #1.

The Results

What Result
Kinds from the day before 3 of 3 named right, 8 to 10 seconds after they began
The new kind Flagged as new after 5 seconds. Once written down, named 10 seconds into its return
Named wrong Never
False alarms None in 721 checks
A check Two SQL views, 4 to 5 ms on average in Node and in Chromium, 35 ms at most
Counts All 1,661 counts by minute and hour, for each service and level, equal to a recount of the lines by separate code

Beyond the Script

A script can end up tuned to itself, so the same code ran on more than the scripted two hours. The day before was checked minute by minute, 1,379 minutes in all. Something stood out in 47 of them, every one during one of the six incidents. Twenty more days used other random seeds, with incidents started at random times: four of known kinds and one new kind each day. Known kinds were named right 80 times out of 80, 12 seconds after they began at the median and 50 seconds at most. The slowest were quickcard outages, since quickcard handles fewer of the payments. New kinds were flagged as new 20 times out of 20. None was named wrong, and nothing stood out outside an incident.

The file also stays the same whichever runtime writes it. The native Engine, precomputing put --lines, took the same 2,375,957 lines and wrote a file equal to the browser’s in all 328,955 values.

What It Shows, and What It Doesn’t

  • The shop and its incidents are invented, and the thresholds were chosen on this shop. Real systems are noisier, and a pilot on real logs is the test that counts.
  • The fingerprint only sees what the templates see. Two faults that break the same lines in the same proportions look alike, and the raw words decide only where they differ.
  • Incidents that overlap mix their fingerprints.
  • A new kind of incident is recognized only after someone writes it down.

Try It

Open the Incident Déjà Vu demo, press Play and watch 12:18: within seconds, Right now says it looks like the northpay outage. Then start the order queue incident, which the file has never seen, write it down and start it again. The comparison is on the page to read, and SELECT * FROM deja_vu asks the file directly.