Key takeaways

  • Monitoring should answer what is critical and who must respond.
  • Zabbix focuses on metrics, availability, performance and service status.
  • Wazuh analyses logs, system changes and security-related events.
  • The greatest value comes from linking an alert to an owner, priority and response instructions.
  • Hundreds of unfiltered notifications are not monitoring — they are a new source of problems.

Why does a business need monitoring?

Without monitoring, a business usually learns about a failure from a user. It is then difficult to determine when the problem began, what caused it and whether a similar event occurred before. Metrics, logs and alert history enable earlier responses and allow expansion to be planned on the basis of data.

Security monitoring has a different purpose from performance monitoring. Not every overload is an incident, and not every change in a log is an attack. Systems and rules must therefore suit the environment and the verification process.

Zabbix and Wazuh — two complementary perspectives

Zabbix

It measures CPU, memory, disk space, service availability, response time, network traffic and device parameters. It helps detect degradation before a failure occurs.

Wazuh

It collects and analyses logs, detecting selected file changes, unusual processes, security errors and events requiring verification.

Dashboards

They should show the status of services important to the business, not every possible metric. An administrator's view may differ from a report for a process owner.

Response

An alert must have a priority, owner, escalation channel and instructions. Without these, it remains merely information.

How should you assess the current situation?

  • Do we know which systems and services are critical?
  • Do we know the normal load and response-time values?
  • Do backup errors, disk space and temperature generate alerts?
  • Are logs from important servers and devices collected in one place?
  • Does every alert have an owner and an agreed response time?
  • Are alerts tested after configuration changes?
  • Is access to the monitoring system protected and audited?
  • Does the business know what to do after a suspicious event is detected?

Alerts that make sense

A threshold should follow from the impact on a service, not an arbitrary value. For example, a brief CPU spike may be normal, but a production system running out of disk space requires a specific response. Delays, dependencies and deduplication are worth using so that one problem does not create dozens of tickets.

Information

An observational event that does not require an immediate response.

Warning

A trend or condition requiring investigation within an agreed time.

Critical

A problem with a direct impact on a service or security.

Escalation

An alert that has not been acknowledged or resolved within the intended time.

Logs and security

Wazuh can support log analysis, file-integrity monitoring and the detection of events on hosts. Rules should be tailored to the business's systems, and the result of the analysis must be verified by a person. A log entry alone does not yet constitute a confirmed incident.

  • Restrict administrators' access to the monitoring system.
  • Protect logs against deletion and unauthorised changes.
  • Define the retention period and the scope of collected data.
  • Do not collect everything without a purpose — reduce noise and costs.
  • Prepare a procedure for isolating an account or device.
  • Document the verification and closure of events.

Common mistakes

Enabling every rule

An excessive number of alerts causes fatigue and leads to important events being ignored. Rules must be tuned to the actual environment.

No alert owner

A notification sent to a shared inbox often does not lead to a response. Every priority should have assigned responsibility and an escalation path.

Monitoring without testing

If you do not test an alert after a service failure, you may not know that the monitoring system itself has stopped working.

What determines the cost?

Cost is affected by the number of hosts and devices, log volume, required storage period, retention, high availability, integrations, notification channels and the time needed for analysis. A cheaper tool does not help if the business lacks a verification and response process.

Step-by-step implementation plan

  1. Define the objectives. Select the services and events that matter to the operation of the business.
  2. Inventory the sources. List the servers, devices, applications and logs available in the environment.
  3. Establish the normal state. Collect baseline values for load, availability and traffic.
  4. Design the alerts. Define thresholds, priorities, owners and escalation.
  5. Run a pilot. Test on a limited group and tune the rules before expanding.
  6. Document the response. Prepare short runbooks and regularly verify that monitoring is working.

Checklist

  • The list of critical systems is up to date.
  • Metrics and logs have defined purposes.
  • Alerts have priorities and owners.
  • An escalation channel exists.
  • Dashboards show important information.
  • Logs are protected against changes.
  • The monitoring platform has its own availability monitoring.
  • Runbooks are updated after changes.

What can the business do itself?

You can start by listing critical services, defining owners and identifying the events the business wants to know about. It is also worth checking whether current alerts are clear and whether someone acknowledges that they have been handled.

When is specialist support needed?

Support is useful when integrating multiple log sources, tuning security rules, implementing high availability, organising escalation and when an incident requires rapid isolation and analysis.

Frequently asked questions

Does monitoring place a load on servers?

It may generate a small load, but the scope and frequency of measurements are matched to the environment's capabilities. It is worth starting with the most important metrics.

Must Zabbix and Wazuh run separately?

Not always. In a small environment, they can initially share infrastructure, but as data volumes grow, separation improves performance and reliability.

Does an alert mean a security incident?

No. An alert is a signal requiring verification. Only an analysis of the context, source and impact allows the event to be classified.

Can monitoring create helpdesk tickets?

Yes. An integration can automatically create a ticket with a priority, description and link to a chart or log.

How often should alerts be tuned?

More frequently after implementation, and later following changes to the environment and reviews. Thresholds should reflect the current way of working.

Related service

Do you need to organise the monitoring of availability, performance and security events?

Explore IT systems monitoring

Read also

Backup and disaster recoveryHow to check whether backups can be restored.Essential IT securityThe most important protective measures.Business server roomsPower, environment and physical security.