Why is it burning?
Where incident management restores service, problem management looks for why it broke.
The distinction is not academic. An organization with only incident management solves the same fault over and over, and gets very good at solving that particular fault. The volume never falls.
Reactive and proactive
Reactive problem management starts from something that has already happened. After a major outage, or when the same ticket type has appeared for the fourth time in a month.
Proactive problem management looks for patterns in ticket data before they become disruptions. It depends on tickets being categorized consistently – one of the strongest arguments for keeping the category list short and usable.
Known errors are the process's most valuable output
When the cause is identified but the fix is delayed – because it needs a vendor update, a project or a budget – the problem is documented as a known error with a workaround.
The effect is immediate. Next time the ticket arrives, nobody needs to investigate anything. The first line resolves it directly, and resolution time drops from hours to minutes.
Known errors belong in the knowledge base and should be searchable where technicians already work.
How to start without a formal process
Pull the five most common ticket categories from last quarter. For each one, ask: why does this keep coming back?
That question alone is often enough to find something actionable – a faulty standard configuration, a training gap, a system that behaves differently from what users expect.
You don't need a formal method to begin. Asking why a few times in a row takes you further than most people think.
Someone has to own it
This is the ITIL practice most often skipped, and the reason is structural rather than anyone being lazy.
A problem has no impatient user waiting. There is never an urgent reason to work on it today, which means it gets postponed every day.
The answer is to set aside time, not to encourage it. A few hours a week for someone who isn't also in the ticket queue goes a long way.
The link to changes and dependencies
Many root causes lead back to something that was changed. A well-kept change log is therefore one of the most useful inputs in an investigation.
The same goes for dependencies. Without a CMDB it is hard to see that three seemingly unrelated incidents share a common component.
The metric
The number of recurring incidents per category over time. It is the only measure that shows whether the process actually removes work rather than adding it.
How the process fits with the rest of your delivery is covered under ITSM.


