Monitoring
The API, the metrics worth alerting on, and query logging.
The API listener
api = "127.0.0.1:8080"
bearertoken = "a-long-random-string"
api = "" disables it. When bearertoken is set, the blocklist, purge and
metrics routes require Authorization: Bearer <token>. The /debug/pprof
routes are the exception and are covered below.
| Endpoint | Method | Purpose |
|---|---|---|
/metrics |
GET | Prometheus exposition |
/api/v1/purge/:qname/:qtype |
GET | Drop one cache entry |
/api/v1/block/exists/:key |
GET | Is a name blocked |
/api/v1/block/get/:key |
GET | Read a blocklist entry |
/api/v1/block/set/:key |
GET | Add a name |
/api/v1/block/remove/:key |
GET | Remove a name |
/api/v1/block/set/batch |
POST | Add many |
/api/v1/block/remove/batch |
POST | Remove many |
The block endpoints are registered whether or not you configured a blocklist — the default chain always builds the handler. Note that the single-name mutations are GET requests, which means a browser or a link preview can trigger them.
That is one reason to keep this listener on loopback. The other is that it is
plain HTTP: bearertoken travels in the clear to a reachable address and can
be replayed, so it is a second layer rather than the protection. A reachable
deployment needs a TLS-terminating authenticating proxy, a VPN, or a firewall
that restricts the source.
/debug/pprof is served only when SDNS_PPROF=true is in the environment — and
those routes are the one exception to the token, since pprof tooling sends no
Authorization header. With pprof on, a token is not sufficient protection for
this listener; keep it on loopback or behind an authenticating proxy. See
diagnostics.
Metrics worth an alert
Is it answering?
dns_queries_total total queries, by type and rcode
dns_resolver_failures_total
dns_resolver_dnssec_failures_total
A rise in dns_resolver_dnssec_failures_total is either an upstream zone that
broke its signing or something interfering with your traffic. It is worth
alerting on because it is invisible to clients — they just see SERVFAIL.
Is the cache doing its job?
dns_cache_hit_rate
dns_cache_evictions_total
dns_cache_size
Evictions climbing while the hit rate falls means cachesize is below the
working set.
Is it being abused?
dns_accesslist_denied_total
dns_ratelimit_exceeded_total
dns_recursion_fanout_ratio
dns_recursion_firewall_exhaustions_total
dns_recursion_fanout_ratio — outbound queries per client query — is the single
most useful number for spotting a query pattern designed to cost you work. It
sits low and flat in normal operation.
Is anything being dropped at the door?
dns_udp_ingress_drops_total
dns_udp_ingress_overflow_total
dns_tcp_ingress_drops_total
dns_listener_errors_total
Overflow means queries arrived faster than the workers accepted them. That is a capacity signal, not a bug.
Everything else. All 64 metrics — with their types, labels, help strings, ready-made PromQL and the alerts worth having — are in the metrics reference.
Query logging
Two mechanisms, for two purposes.
accesslog = "/var/log/sdns/access.log"
Common Log Format, one line per query, human-readable. On a busy resolver this is the largest thing the process writes; it is off by default for that reason.
dnstapsocket = "/var/run/sdns/dnstap.sock"
dnstapidentity = "sdns"
dnstaplogqueries = true
dnstaplogresponses = true
dnstapflushinterval = 5
dnstap is a binary protocol over a Unix socket, meant for a collector rather than for reading. It is the right choice when you want to keep query data at volume.
Confirming which build is running
dig @resolver version.bind TXT CHAOS +short
Works while chaos = true. Useful across a fleet, where the answer to “did
that deploy land everywhere” is otherwise a guess.