TellHound
Read-only · No agents · No access keys

Your AWS dashboards show what happened. Not what to do about it.

Connect a read-only role and get findings a human can act on — the tier that cannot scale, the commitment about to lapse, the money going nowhere. What is wrong, why it matters, exactly how to fix it, and what the fix costs. Not a wall of graphs to interpret at 3am.

For the people who get paged — and the people who see the bill. No credit card, real findings in minutes.

Account health · production
72
Health score
31
Open findings
4
P1 — act now
$1.4k
Monthly waste
P1 Capacity provider reports scaling ENABLED but its ASG is pinned — scaling cannot act
media-transcode-cluster · ecs-capacity-provider-inert
P2 Reserved Instance expires in 11 days — renew, resize, or let it lapse
orders-db · commitment-expiring
P3 14 unattached volumes billing for nothing — ≈$41.20/month
account-wide · ebs-unattached
+ 28 more across reliability, security, cost and posture
EC2RDS & AuroraECS & FargateLambda CloudFrontApplication Load BalancersElastiCache Auto Scaling GroupsEBS & Elastic IPsVPC & Security Groups Savings Plans & RIs EC2RDS & AuroraECS & FargateLambda CloudFrontApplication Load BalancersElastiCache Auto Scaling GroupsEBS & Elastic IPsVPC & Security Groups Savings Plans & RIs

Not another wall of graphs

Three things every AWS estate is quietly losing: money it cannot see, capacity that cannot scale, and alerts nobody trusts. Here is what each looks like.

1 / 3
Cost opportunities · ranked by monthly saving
8 unassociated Elastic IPs billing hourly
Release the ones you no longer need.
$29.20
api-worker (t4g.large) peaked at 22% CPU
One size down saves ≈$24.53/month.
$24.53
14 unattached EBS volumes — 209 GB
Snapshot anything worth keeping, then delete.
$16.72
edge-cache-01 (t3a.small) peaked at 6% CPU
Rightsizing candidate — check memory first.
$6.86

It refuses to guess

Every figure states its pricing basis, and anything it cannot justify is reported as unpriced rather than counted as zero. A memory-optimised instance at 5% CPU looks like an easy win and is not one — shrinking a cache tier to save money is how a cost review becomes an outage. So we surface it and refuse to put a number on it.

Finding detail
P1 ecs-capacity-provider-inert
Capacity provider reports managed scaling ENABLED but its ASG is pinned at 2 — scaling cannot act
Why it matters

ECS asks the Auto Scaling Group for more instances, but an ASG with Min == Max has none to grant. The decision is made and silently discarded — more dangerous than no autoscaling, because every dashboard reports the tier as protected.

How to fix it

Raise Max above Min so scaling has headroom. If the ceiling is deliberate — a per-instance licence, a downstream connection limit — document it and alarm on saturation instead.

Cost impact: none. You are billed for instances running, not for the ceiling. Cost rises only when a real spike causes a scale-out that would otherwise have been dropped.

Written by someone who got paged

That finding is real. It described an image-processing tier where every dashboard showed autoscaling enabled and the tier still fell over, because the group beneath it had no room to grow. Neither resource looked wrong alone — only the join between them did.

Roughly 30 rules, each encoding a failure that actually happened. Every one ends with an action and its price.

Alert watches · data as of 40 seconds ago
Cache: hit ratio collapse
avg over 15m below 30
ok (3)
Cache: service down (or reporter stopped)
min over 5m below 1
firing (1)
Load balancer: 5xx surge
sum over 5m above 50
ok (24)
Database: CPU sustained high
avg over 15m above 80
pending (2)
Alerting paused. TellHound has not collected metrics for 24 minutes, so watches are held rather than firing "no data". This is our problem, not yours.

It will tell you when it can't see

Most monitoring treats missing data as a failure in your infrastructure. When the collector itself stops, that produces a page for every watch at once, at 1am, about services that are perfectly healthy.

TellHound checks its own pipeline first. If our data is stale, alerting pauses and says so in one message — because a monitor that cries wolf about its own outage is worse than one that stays quiet.

30
Rules across reliability,
security, cost & posture
9
AWS services
inventoried
0
Agents to install
or keys to hand over
5
Minutes from role
to first finding

Everything here exists because something broke

Not a checklist. A record of what actually goes wrong.

Findings, not metrics

What is wrong, why it matters, the exact remediation, and what acting costs — including when the answer is "nothing".

$

Cost that survives scrutiny

Ranked monthly savings with the pricing basis stated, and an explicit refusal to price what it cannot justify.

Commitment expiry

Reserved Instances and Savings Plans stop on a fixed date with no console banner. We warn at 60 days, escalate inside 14.

Honest alerting

Flap guards, cooldowns, guaranteed recovery notices — and a pause when our own collection goes stale.

Security posture

World-open groups, public databases, unencrypted storage, missing flow logs. A failed sweep is reported, never silently passed.

Explained health

One score per account with the arithmetic shown, so you can argue with it. Unassessed accounts are never graded.

Connected in three steps

1

Create a read-only role

One CloudFormation template, published in full so you can read every permission before you run it. Scoped to an ExternalId unique to you.

2

We read, we never write

TellHound holds no AWS keys and cannot accept them. It assumes your role for 15 minutes at a time and has no permission to change anything.

3

Get findings, not homework

Inventory, health score, ranked cost opportunities, and the findings that matter — each with a remediation you can hand to whoever owns it.

See what your account is hiding

14 days, one AWS account, no card. If it finds nothing worth fixing, that is worth knowing too.

Request early access