Series A–B teams burn $5K–50K/month on observability — up to a third of it on metrics nobody queries and logs nobody reads. We quantify the waste in dollars, in minutes.
Every audit cross-references your paid telemetry against what humans actually query — then ranks the waste.
Metrics billed every month that no dashboard, monitor or notebook has queried in 30+ days.
worker.job.retry_count_v2 → 900 seriesA single customer_id tag turns one metric into thousands of billable time series.
app.checkout.abandoned_cart_valueDebug sources indexed at 30-day retention with zero exclusion filters — read once a year, paid every month.
prod-debug · 168M events · ×2 rateHealthy traces retained at full volume. Errors and slow requests are what you actually debug.
180M spans/mo · tail-samplingNon-prod runs 40h/week but meters run 168h/week. Business-hours monitoring alone recovers ~73%.
stg-node-01..14 · infra + APM90-day retention applied org-wide when compliance only requires it on two audit indexes.
security-audit: keep 90d · rest: 15d*Ranges from the sample audit above. Your numbers come from your own usage data.
Paste API + App keys or connect via OAuth. We pull usage meters, metric inventory, dashboard definitions and log config — metadata only.
Billed metrics vs queried metrics. Indexes vs exclusion filters. Host counts vs actual hours anyone is watching.
Every finding ships with its monthly cost range and the exact fix. Apply it yourself or hand it to your platform team.
We request read-only scopes and hold keys in memory for the duration of one audit. We pull metadata — usage meters, metric names, dashboard definitions, index configs — never log contents. Revoke the app key right after if you like.
They exist and we recommend them — but your vendor will never tell you to leave or maximize your savings. We're vendor-neutral: our incentive is your bill going down, including by migrating signals off the meter.
The fixes target telemetry nobody queries. For everything else we favor sampling and tiered retention over deletion — you keep error and slow traces at 100%. Every recommendation links to Datadog docs so your team can verify.
Findings are labeled measured (from your Usage Attribution data) or estimated (heuristics), and always shown as ranges. Your contract discounts may shift absolute numbers; relative waste does not change.