ORIGINAL REDDIT POST

how do you handle incidents when your team is scattered and panicking?

genuinely curious how other founders with teams to run with startups handle this now: something breaks at 2am. one person finds the logs, figures out what happened, fixes it. everyone else wakes up with no idea what occurred or why and next time the same…

Original postr/SaaS

genuinely curious how other founders with teams to run with startups handle this now: something breaks at 2am. one person finds the logs, figures out what happened, fixes it. everyone else wakes up with no idea what occurred or why and next time the same thing breaks you're starting from scratch again because nothing was documented in the chaos. i have unfortunately been sitting on this problem for a while. no elegant solution i've found yet, pagerduty is overkill, notion docs after the fact are useless, slack threads are a mess. how are you actually handling this?

Collected discussion

2 comments

u/hijinks

devops/sre eng here and been on call for 26 or so years now. You need to have an incident commander. This is the person that brings the people that can fix the problem in and gives them jobs. They also let non-engs know the status so they aren't bugging people. There are tons of solutions in the saas space for incident handling. look at rootly or incident.io for 2. There's like 10 others.

u/Sad-Spray3039

the reason notion-after-the-fact never works isn't discipline, it's timing. by the time someone sits down to write it up the adrenaline's gone and half the context left with it. willpower won't close that gap. what's worked for us is forcing capture at the moment of the fix instead. whoever's in the logs hits one thing, a slack shortcut or a webhook, and it dumps the raw stuff, last 50 log lines, the query they ran, whatever they were staring at, straight into a thread. no formatting, ugly is fine. tidy it at 9am or don't, still beats zero. and if it's the same thing breaking over and over, a dumb little bot that asks "what was the fix" with a dropdown of the usual suspects beats any real incident tool. match the size of the fix to how often it actually bites you.