SERVICE MANAGEMENT

Why incidents keep repeating: incident versus problem management

Many IT teams clear failures quickly and competently. The same incidents still return weeks later. The cause is rarely a lack of skill — it is that the organisation runs only one of the two processes it needs.

Abstract diagram: a repeating loop and the branch that breaks it

Two different goals

Incident management exists to restore service as fast as possible. Problem management exists to remove the cause so the incident stops recurring. Different work, different metrics, often different people.

With only incident management, every failure closes as resolved. The statistics look healthy — resolution time is short. But total incident volume does not fall, because nobody investigates why they arise.

How to recognise the gap

The signs are usually visible without any analysis:

  • The same incidents return every few weeks, each closed separately.
  • The team recognises a failure from its description and knows the fix by heart — meaning the cause is known but not removed.
  • There are “temporary” workarounds that have been running for a year.
  • Nobody can say which five failures generate the most tickets.

Where to start

No tool is required, and no formal ITIL programme. One habit is enough: review incidents once a week and identify the repeats.

Each repeat gets an owner and a deadline. It matters that this is a record separate from the incident — otherwise it closes along with the failure and the cause survives.

After a few months the first real metric appears: what share of incidents are repeats. If it falls, the process works. If not, it usually means problems are assigned but nobody has time for them — which is a management decision, not a technical one.

What the business gets

Fewer repeat incidents frees team capacity without additional headcount. It is usually the cheapest way to increase IT capacity — considerably cheaper than a new tool or another hire.

WORTH REMEMBERING

  • Incidents restore service; problems remove causes. You need both.
  • A short resolution time does not mean the process is working.
  • You can start without a tool — a weekly review of repeat incidents.
  • The metric that matters: what share of incidents are repeats.

RELATED

Where this leads next