My baseline starts with reachability, interface state, DNS response time, and a small number of service checks. Those four signals catch an unreasonable percentage of real problems.
I prefer tools I can understand from end to end. A lightweight metrics collector, a dashboard, and alerts that explain what changed are more valuable than hundreds of checks nobody trusts.
Collect signals that change decisions
The rule is simple: collect only what will change a decision. If a graph cannot tell me whether to wait, investigate, or roll back, it probably does not belong on the main dashboard.