Everything here exists because something broke.
TellHound is not a metrics explorer with an AI summary bolted on. Each rule encodes a specific failure — usually one that took a production system down — and each finding carries the fix.
It sits alongside whatever monitoring you already run rather than replacing it. Your dashboards keep telling you what happened; this tells you which of it needs doing something about.
Findings that end with an action
A finding states what is wrong, why it matters, exactly how to fix it, and what the fix costs. Recommendations that hide their price are recommendations you cannot act on, so every one carries a cost line — including when the answer is "this is free".
- Roughly 30 rules across reliability, security, cost, observability and maintenance
- Every finding: what / why / how / cost, plus the raw evidence
- Auto-resolve when the underlying problem is actually fixed, and reopen if it returns
- Acknowledge or mute individually — muted stays muted
A cost engine that refuses to guess
Ranked monthly savings with the pricing basis stated, and an explicit refusal to price anything it cannot justify. It will not tell you to shrink a cache tier because its CPU looks low.
- Unattached volumes, idle Elastic IPs, gp2 volumes, oversized instances
- Memory-optimised instances are reported but never priced — CPU is the wrong signal for them
- ECS container hosts are flagged as sized by task reservations, not CPU
- Anything unpriceable is shown as an unpriced remainder, never counted as $0
Commitment expiry, before the invoice
Reserved Instances and Savings Plans end on a fixed date and simply stop. AWS emails the account owner once and raises no alarm, so the first real signal is a bill that jumped.
- Warns at 60 days, escalates to P1 inside 14
- Reports commitments that ALREADY lapsed — a worse state than expiring
- Quantifies the exposure from the actual hourly commitment
- Flags a steady fleet running entirely on-demand
Alerting that admits uncertainty
A monitor that pages you for its own failure destroys trust faster than one that misses an alert. If our collection goes stale, alerting pauses and says so.
- Flap guard and cooldown, so nothing pages twice a minute
- Recovery is always sent — an alert with no "OK" trains people to ignore alerts
- Missing data can page (a stopped publisher is a real outage) but not when OUR pipeline is the thing that stopped
- Curated templates encoding real failure modes, with thresholds already set
Inventory and posture
A live inventory of what you actually run, with security posture checks that never present "we could not look" as "nothing wrong".
- EC2, RDS, ECS, Lambda, ELB, CloudFront, ElastiCache, ASGs, volumes, IPs, VPCs, security groups
- A failed collection sweep is reported as its own finding, not silence
- Per-resource pages with configuration, metrics and the findings for that resource
- Metric coverage advisor: what you can enable, how, what it costs, what it buys