My baseline starts with reachability, interface state, DNS response time, and a small number of service checks. Those four signals catch an unreasonable percentage of real problems.

I prefer tools I can understand from end to end. A lightweight metrics collector, a dashboard, and alerts that explain what changed are more valuable than hundreds of checks nobody trusts.

Collect signals that change decisions

The rule is simple: collect only what will change a decision. If a graph cannot tell me whether to wait, investigate, or roll back, it probably does not belong on the main dashboard.