We had something break this morning. From it being raised to working again took 20 minutes. The fix took four. Restart a service. That was it. The interesting bit is the other 16 minutes, because that's where the time normally goes. Not fixing the problem, but working out what to fix. And this one wasn't obvious. An infrastructure change in a different part of the stack exposed a defect nobody knew was there. The symptom was in one place, the cause somewhere else entirely. We’re currently trialling incident.io "Investigations", and this was a good example of where it could help. It kicked off automatically when the issue was raised and, by the time everyone was in the channel, there was already a hypothesis with the reasoning and evidence behind it. Instead of starting from scratch, the team had something to prove or disprove. The other important bit is that Investigations only had something useful to work with because of the groundwork underneath it. We've spent years of consolidating our monitoring, improving tracing and logging, and making infrastructure changes visible in a way it can understand. It's early in the trial, but that's probably the biggest takeaway for me so far. It's good at spotting patterns across far more information than any of us can digest at once. But it can only do that with what we've wired up for it. The agent is a genuine upgrade. The wiring is what makes it one. Big shout out to Tim Nicholls, Damon Marlow, Liam Fitzpatrick, Aparna Valsala, Patrick Bontoft and Michael B. for getting it resolved so quickly.
Thank you for the kind words Leigh! It's exactly what we built it for :)
Glad Investigations helped here!
Nice story Leigh! I mean why not take it further, and get AI fix it for you 👀