Dear Knot Resolver users,
Knot Resolver 6.5.0 has been released!
Bugfixes:
- views: fix "unused tags" error when repeating a tag in views (!1883)
- watchdog: detect file-changes done via rename*() syscall (!1895)
This is relevant for RPZ files and TLS certs+keys.
- DoH fixes for sharing connections among different users (!1844)
- TCP-based connections used Nagle's algorithm since 6.2.0 (!1903)
This mistake significantly increased latencies of such transports.
- local-data: process upper-case names correctly (!1915)
Packaging:
- deb: migrate to sysusers and update packaging (!1857)
- rpm: fix source build against libknot >= 3.6 (!1907)
- python: 3.9+ (!1886), migrated to setuptools (!1889), removed
typing-extensions (!1910)
- Debian 11 is no longer supported
- Debian,Ubuntu: i386 and armhf architectures no longer supported
Improvements:
- decrease RAM usage by a lot (!1901)
- local-data: increase capacity to 63 GiB (logged MDB_MAP_FULL, !1885)
- systemd: use unit's {Runtime,State,Cache}Directory directives instead
of tmpfiles.d (!1853)
- controller: reimplement the notify socket in Python (!1875)
- python: improve Knot Resolver startup procedure (!1893)
Some startup actions previously handled by the manager have been
moved to a separate startup program.
- python: refactoring: utils (!1894), client (!1908)
- allow running under an unknown user, e.g. in containers (!1914)
Incompatible changes:
- remove legacy systemd unit files (kresd@.service,
kres-cache-gc.service) (!1853)
- remove $KRES_LOGGING_TARGET environment variable (!1893)
You can override the default by passing `--logtarget` to
`knot-resolver` command. (!1912)
- docker: remove cross-platform build for linux/arm/v7 (!1904)
Full changelog:
https://gitlab.nic.cz/knot/knot-resolver/raw/v6.5.0/NEWS
Sources:
https://knot-resolver.nic.cz/release/knot-resolver-6.5.0.tar.xz
GPG signature:
https://knot-resolver.nic.cz/release/knot-resolver-6.5.0.tar.xz.asc
Documentation:
https://www.knot-resolver.cz/documentation/v6.5.0/
--
Ales Mrazek
PGP: 3057 EE9A 448F 362D 7420 5A77 9AB1 20DA 0A76 F6DE
Hello,
I would like to report what appears to be a regression in Knot Resolver
6.4.2.
An RPZ rule loaded via local-data.rpz does not match queries if the rule's
owner name contains uppercase ASCII letters. Rules whose owner names are
entirely lowercase work correctly and match queries regardless of query
case.
Ordinary DNS lookups on the same resolver behave case-insensitively as
expected; for example, mixed-case queries such as GoGle.pl and WP.pl
resolve normally.
Affected version
Tested with:
-
knot-resolver6 6.4.2-cznic.1~bookworm
-
knot-resolver6-module-dnstap 6.4.2-cznic.1~bookworm
-
libknot16 3.5.8
-
libdnssec10 3.5.8
-
Debian 12 (bookworm)
Not affected
-
Knot Resolver 5.7.6 (cznic.1), with the same RPZ data loaded using Lua
policy.rpz()
Summary
DNS name matching is case-insensitive for ASCII letters. RFC 1034 §3.1 and
RFC 4343 specify that DNS name comparisons ignore ASCII case.
With local-data.rpz in Knot Resolver 6.4.2, however, a rule whose owner
name contains uppercase letters is never applied.
It does not match when the query:
-
is entirely lowercase,
-
uses exactly the same case as the RPZ owner name, or
-
is entirely uppercase.
Real-world example
Our RPZ blocklist contains eight entries with mixed-case owner names, for
example:
rabOna-7681.com IN A 10.15.20.254www.rabOna-7681.com IN A
10.15.20.254Italy9.cabaretclub.com IN A 10.15.20.254
totaIcasino.nl IN A 10.15.20.254
The same RPZ data is deployed on servers running Knot Resolver 5.7.6 and
6.4.2.
With 5.7.6 and policy.rpz():
$ dig www.rabona-7681.com @127.0.0.1
;; flags: qr aa rd ra;www.rabona-7681.com. 7200 IN A 10.15.20.254
With 6.4.2 and local-data.rpz:
$ dig www.rabona-7681.com @127.0.0.1
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN
$ dig rabOna-7681.com @127.0.0.1
(no RPZ answer)
None of the eight mixed-case entries receives the configured RPZ answer on
6.4.2. Lowercase entries from the same file work correctly.
Minimal reproducer
/etc/knot-resolver/case-test.rpz:
$TTL 300
@ IN SOA localhost. root.localhost. 1 3600 600 86400 300
@ IN NS localhost.
TeSt-case.example IN A 192.0.2.1
lower-case.example IN A 192.0.2.2
Relevant part of config.yaml:
local-data:
rpz:
- file: /etc/knot-resolver/case-test.rpz
kresctl validate --strict succeeds. I then run kresctl reload.
Results:
query status answer
lower-case.example NOERROR A 192.0.2.2
LOWER-CASE.example NOERROR A 192.0.2.2
test-case.example SERVFAIL none
TeSt-case.example SERVFAIL none
TEST-CASE.EXAMPLE SERVFAIL none
other.example SERVFAIL none
WP.pl NOERROR normal DNS answer
GoGle.pl NOERROR normal DNS answer
other.example, WP.pl, and GoGle.pl are control queries and are not present
in the RPZ file. WP.pl and GoGle.pl demonstrate that ordinary DNS
resolution works correctly and remains case-insensitive, while the
mixed-case RPZ owner name fails to match.
The lowercase rule matches queries regardless of case. The mixed-case rule
does not match any query, including a query using exactly the same case as
the owner name in the RPZ file.
Expected behavior
This rule:
TeSt-case.example IN A 192.0.2.1
should match:
test-case.example
TeSt-case.example
TEST-CASE.EXAMPLE
in the same way that the lowercase rule does, and in the same way that Knot
Resolver 5.7.6 with policy.rpz() behaves.
Actual behavior
An RPZ rule loaded through local-data.rpz is effectively ignored if its
owner name contains uppercase ASCII letters.
Workaround
Convert all owner names in the RPZ file to lowercase before loading the
file.
--
Best regards,
Mateusz Masłowski
Hello,I run a small public DNSSEC-validating recursive resolver on Knot Resolver
6.4.2 and hit an eleven-minute outage that I think is a bug in expiring
prefetch. I have aggregate metrics across the whole event but no
query-level data, for reasons I explain at the end. I would appreciate a
steer on whether this is known, and on the two questions at the bottom.
WHAT HAPPENED
-------------
On 2026-08-13 between 04:01 and 04:12 UTC:
- resolver_request_internal_total rose from a steady 1.34/s to a peak
of 10,901/s and stayed elevated for about eleven minutes
- the workers logged 285,480 "[system] error: stack overflow" messages,
peaking at 94,675 in a single minute
- the resolver stopped answering; my end-to-end probes failed on all
six transport/family combinations at 04:08, 04:10 and 04:14
- it ended on its own at 04:12 and has not recurred since
Client query rate never changed. Roughly 6.5 million internal requests
were generated in ten minutes against a real client demand of about 8.6
queries per second.
No worker crashed or restarted.
MEASUREMENTS
------------
At the peak, against the same counters ten hours later:
peak normal
resolver_request_internal_total 5,204/s 1.34/s
(max sample) 10,901/s
resolver_request_udp+tcp_total 8.6/s 8.9/s
dnsdist_queries (real client demand) 7.5/s 7.7/s
resolver_answer_total 4,800/s 10.3/s
cache hit ratio 0.06% ~30%
CPU busy 30% 3%
load1 2.06 0.25
p99 resolver_response_latency 1.5s 0.4s
Per worker, over the burst:
kresd0 0 overflow messages
kresd1 92,474
kresd2 90,143
kresd3 102,863
One worker in four was completely untouched, which I cannot explain.
RESOURCE EFFECT
---------------
Memory available fell from 2,862 MB to 622 MB and the kernel swapped
1,798 MB. The three affected workers were exactly the three that ended up
in swap (596, 537 and 524 MB) and each took roughly 240,000 major page
faults, against 30 for the entire process lifetime before the event. The
resolver was serving its LMDB cache from disk for the ten hours until I
restarted the container.
ENVIRONMENT
-----------
Version 6.4.2.dev1+f73d6f (kresd --version)
Image cznic/knot-resolver, tag v6.4.2
sha256:b589bfe67a61e2d3c1d6ea3836904e4ed2ec70c77c5db9127473
ed092d31d638
Deployed Docker, host networking, non-root, read-only root filesystem
Workers 4
Host 4 vCPU, 3852 MB RAM, Ubuntu 26.04 LTS, kernel 7.0.0-29
Cache LMDB, size-max 1536M, persistent
Role recursive resolver behind dnsdist 2.1.1 over loopback
Relevant configuration:
workers: 4
logging:
level: info
groups: [system, module, devel, io]
cache:
size-max: 1536M
prefetch:
expiring: true # prediction deliberately left disabled
options:
minimize: true
serve-stale is NOT enabled, so this is not issue #957.
WHAT I RULED OUT
----------------
- Not client-driven. dnsdist_queries is flat across the whole window.
No query spike, no new source, no rate-limit rule hit.
- Not a crash. No worker restarted; the only spawn lines in the
container log are from process start.
- Not serve_stale (#957). Not enabled.
- Not cache exhaustion. LMDB is 1536 MB and was not full.
- Not a one-off message. Isolated "stack overflow" lines occur at a
background rate of roughly 0.1/hour, on 6.4.1 and 6.4.2 alike,
without any internal-request spike. Only this event showed the
runaway. Possibly two related phenomena.
I also enabled debug logging for the system, module, devel and io groups
before this happened. They show no precursor at all: normal traffic, then
stack overflow messages at microsecond intervals.
HYPOTHESIS
----------
This is a guess about mechanism rather than a diagnosis.
Prefetch of expiring records is, as I understand it, answer-triggered: a
record is refreshed when the resolver answers with it at under 1% TTL or
under 5 seconds remaining. If a prefetch's own resolution itself answers
with a near-expired record, that would trigger a further prefetch, and
the loop could sustain itself with no client involvement.
What I can state from the data is only that internal request generation
became self-sustaining and decoupled from demand.
QUESTIONS
---------
1. Is "[system] error: stack overflow" a caught Lua or LuaJIT stack
limit? Is the affected request abandoned, or retried? A retry would
explain the self-sustaining behaviour.
2. Is there any rate limit or de-duplication on expiring prefetch, or a
guard preventing a prefetch from triggering further prefetches?
3. Does one worker of four being entirely unaffected suggest per-worker
state as the trigger?
WHAT I CAN PROVIDE
------------------
The service has a published no-query-logging policy, so query names and
client addresses were never captured and do not exist. I realise that is
the first thing you would normally ask for, and I am sorry not to have
it. I do have:
- full system/io/module/devel debug logs for the window, about 40 MB,
containing only upstream authoritative server addresses
- Prometheus series for any exported counter across the event
- the complete configuration
I am happy to run with cache.prefetch.expiring set to false to confirm
the association, or to carry a patch or an extra debug group if that
would help narrow it down. The resolver is low-traffic and I can
experiment on it freely.
Thanks for your time, and for the resolver.
Hi,
I just upgraded knot-resolver from 5.7.5 to 6.4.1 and now I'm seeing the
following error message regularly:
kresd[48937]: [cache ] [16128.01] stash failed, ret = 1
kresd[48937]: [cache ] [37316.01] stash failed, ret = 1
Can someone explain what that means?
--
Stefan Schweizer
Dear Knot Resolver users,
Knot Resolver 6.4.1 has been released!
Security:
- DNS-over-QUIC (DoQ) had severe issues, allowing even RCE
Many people reported (some of) these issues to us.
- DNSSEC correctness issues, acting mainly through the aggressive cache:
* dealing with Labels field in RRSIGs being smaller than the signer's
* dealing with NSEC's next-name pointing outside of the zone
Special thanks to Qifan Zhang from Palo Alto Networks.
Improvements:
- docker: upgrade to Debian 13 (!1856)
- update IANA's certificate for root trust anchor bootstrapping (!1845)
Bugfixes:
- /local-data/addresses*: make multiple addresses work (#808, #954)
- views: fix protocol-based matching for DoQ
Full changelog:
https://gitlab.nic.cz/knot/knot-resolver/raw/v6.4.1/NEWS
Sources:
https://knot-resolver.nic.cz/release/knot-resolver-6.4.1.tar.xz
GPG signature:
https://knot-resolver.nic.cz/release/knot-resolver-6.4.1.tar.xz.asc
Documentation:
https://www.knot-resolver.cz/documentation/v6.4.1/
--
Ales Mrazek
PGP: 3057 EE9A 448F 362D 7420 5A77 9AB1 20DA 0A76 F6DE
Dear Knot Resolver users,
Knot Resolver 5.7.7 has been released!
Security:
- DNSSEC correctness issues, acting mainly through the aggressive cache:
* dealing with Labels field in RRSIGs being smaller than the signer's
* dealing with NSEC's next-name pointing outside of the zone
Special thanks to Qifan Zhang from Palo Alto Networks.
Improvements:
- support cmocka 2.0.0
- avoid AD=1 in reply if ANSWER+AUTHORITY are empty (#914)
- packaging: rpm: require python3-setuptools (!1831)
- packaging: rpm: provide user/group (!1838)
This should also resolve the issue with user and group
configuration during installation (GH#130).
- make DoH cache-control header respect our cache's TTL limits (!1832)
- support libdnssec merged into libknot, as planned for knot >= 3.6 (!1833)
- update IANA's certificate for root trust anchor bootstrapping (!1862)
Bugfixes:
- respect disablement of QNAME case randomization even after TCP issues
- cache: fix wrong TTL in some cases, typically 32768
- reduce excessive caching of some uncommon failed answers (!1832)
- dns64: fix CNAME problems again (#797, !1862)
Full changelog:
https://gitlab.nic.cz/knot/knot-resolver/raw/v5.7.7/NEWS
Sources:
https://knot-resolver.nic.cz/release/knot-resolver-5.7.7.tar.xz
GPG signature:
https://knot-resolver.nic.cz/release/knot-resolver-5.7.7.tar.xz.asc
Documentation:
https://www.knot-resolver.cz/documentation/v5.7.7/
--
Ales Mrazek
PGP: 3057 EE9A 448F 362D 7420 5A77 9AB1 20DA 0A76 F6DE
Hi,
is it somehow possible to disable RFC 8198 - aggressive caching in knot-resolver 6?
I cannot find any information about that in docs. Unnamed chatbot is sure that it can be done with "cache.aggressive = false" in LUA script. But since it is undocumented, I would like to confirm that.
Regards
Jiri Masek