Se alla lösningsområden

Problem Management

The difference between putting out fires and finding out why they start – and why the process is so often skipped.

Why is it burning?

Where incident management restores service, problem management looks for why it broke.

The distinction is not academic. An organization with only incident management solves the same fault over and over, and gets very good at solving that particular fault. The volume never falls.

Reactive and proactive

Reactive problem management starts from something that has already happened. After a major outage, or when the same ticket type has appeared for the fourth time in a month.

Proactive problem management looks for patterns in ticket data before they become disruptions. It depends on tickets being categorized consistently – one of the strongest arguments for keeping the category list short and usable.

Known errors are the process's most valuable output

When the cause is identified but the fix is delayed – because it needs a vendor update, a project or a budget – the problem is documented as a known error with a workaround.

The effect is immediate. Next time the ticket arrives, nobody needs to investigate anything. The first line resolves it directly, and resolution time drops from hours to minutes.

Known errors belong in the knowledge base and should be searchable where technicians already work.

How to start without a formal process

Pull the five most common ticket categories from last quarter. For each one, ask: why does this keep coming back?

That question alone is often enough to find something actionable – a faulty standard configuration, a training gap, a system that behaves differently from what users expect.

You don't need a formal method to begin. Asking why a few times in a row takes you further than most people think.

Someone has to own it

This is the ITIL practice most often skipped, and the reason is structural rather than anyone being lazy.

A problem has no impatient user waiting. There is never an urgent reason to work on it today, which means it gets postponed every day.

The answer is to set aside time, not to encourage it. A few hours a week for someone who isn't also in the ticket queue goes a long way.

The link to changes and dependencies

Many root causes lead back to something that was changed. A well-kept change log is therefore one of the most useful inputs in an investigation.

The same goes for dependencies. Without a CMDB it is hard to see that three seemingly unrelated incidents share a common component.

The metric

The number of recurring incidents per category over time. It is the only measure that shows whether the process actually removes work rather than adding it.

How the process fits with the rest of your delivery is covered under ITSM.

Problem management finds the root cause behind recurring incidents. It is the process that provides the most value and is the one most often skipped.

Frequently asked questions about problem management

What is the difference between an incident and a problem?
The incident is the individual disruption. The problem is the underlying cause. Ten incidents can belong to one problem.

What is a known error?
A problem where the cause has been identified but not yet fixed, documented together with a workaround so the service desk can act immediately next time.

How do we get started?
Look at the five most common ticket categories from last quarter and ask why they keep recurring. No formal process required.

Who should own the process?
Someone who isn't also sitting in the ticket queue. Whoever is putting out fires never has time to find out why things are burning.

What is the difference between reactive and proactive problem management?
Reactive starts from a disruption that has already happened. Proactive looks for patterns in ticket data to find what is about to go wrong.

Do we need a formal analysis method?
No. Asking why a few times in a row gets you a long way. The method matters less than someone actually doing the analysis.

Which metric shows whether it works?
The number of recurring incidents per category over time. If it falls, the process is working.

Why is the process skipped so often?
Because nobody calls to chase it. A problem has no impatient user waiting, so it gets deprioritized every single day.

Produkter inom området

Nordlo

Nordlo builds scalable IT delivery with MSP Nordics

37% more efficient service management through standardised processes and automation
Läs mer
Läs mer
With the right platform and clear processes, we can scale our delivery without increasing administration at the same rate.
Nordlo

Want to hear more?

We are happy to tell you more about how we have adapted and tailored long-term solutions for our customers.