All posts

DNS from my containers timed out, and only from my containers

·5 min read

AdGuard Home runs on a Raspberry Pi on my home network and does DNS for the whole house. The router hands out 192.168.1.43 as the resolver, every phone and laptop uses it, and the query log shows ad domains getting blocked. As far as anyone in the house could tell, it worked.

Then I added Uptime Kuma, another container on the same Pi, and gave it a DNS monitor pointed at 192.168.1.43. It timed out every time. dig @192.168.1.43 from my Mac answered instantly. The same query from inside the container got nothing back.

The setup

AdGuard was published the way most compose files publish it:

services:
  adguard:
    image: adguard/adguardhome:v0.107.79
    ports:
      - "53:53/udp"
      - "53:53/tcp"

No host IP on the port, so Docker binds 0.0.0.0:53.

Where the reply comes from

LAN traffic and container traffic take different paths to that port.

A query from my laptop comes in on eth0, and Docker's iptables rules DNAT it straight to the AdGuard container. Conntrack remembers the translation and rewrites the reply on the way out, so the laptop gets an answer from 192.168.1.43:53, the address it asked.

A query from another container goes to the Pi's own IP. That traffic gets handed to docker-proxy, the small userland process Docker starts for each published port. docker-proxy takes the packet, opens its own connection to AdGuard, and sends the answer back to the asking container itself.

With the port bound to 0.0.0.0, docker-proxy's UDP socket isn't tied to any particular address. When it replies, the kernel picks a source address from the route back to the container, and that route goes through the Docker bridge. So the answer leaves from the bridge gateway, something like 172.18.0.1, not from 192.168.1.43.

The container asked 192.168.1.43 and got a reply from 172.18.0.1. Plenty of resolvers check that the reply comes from the server they queried, and Node's does. Uptime Kuma is a Node app, so it threw the reply away and waited until it timed out. dig doesn't check, which is why every manual test I ran inside a container looked fine.

Fix one: bind to the real address

Publish the port on the Pi's LAN IP instead of the wildcard:

    ports:
      # On 0.0.0.0, docker-proxy answers containers' UDP queries from their
      # bridge gateway (172.x.0.1), and resolvers such as Uptime Kuma's drop the reply.
      - "192.168.1.43:53:53/udp"
      - "192.168.1.43:53:53/tcp"

Now docker-proxy's socket is bound to 192.168.1.43, so its replies leave from that address. The Uptime Kuma monitor went green within a minute. LAN clients didn't notice anything, because their traffic never went through the proxy.

Fix one causes problem two

A few days later the Pi rebooted after a power problem, and DNS was down for the whole house. Every device got timeouts until I ran docker start on the AdGuard container by hand.

The container log showed:

cannot assign requested address

At boot, Docker started before DHCP had given eth0 its address. Binding a socket to 192.168.1.43 fails when the machine doesn't own that address yet. With 0.0.0.0 this never mattered, because the wildcard always binds. Docker treats a failed port bind as a failed start and doesn't keep retrying it, so restart: unless-stopped didn't save me. The container just stayed stopped.

I'd swapped a bug that only one monitor could see for one that took down DNS for the whole house after a reboot.

Fix two: let the kernel bind an address it doesn't have yet

Linux has a sysctl for this:

echo "net.ipv4.ip_nonlocal_bind = 1" >/etc/sysctl.d/60-rpi-dock.conf
sysctl -q -p /etc/sysctl.d/60-rpi-dock.conf

With ip_nonlocal_bind=1, a process can bind to an IPv4 address the host hasn't been assigned. The bind succeeds right away, and the socket starts getting traffic once DHCP hands over the address. HAProxy and keepalived setups use the same setting for floating IPs.

It does loosen one safety check: a typo in a bind address now fails silently instead of loudly. On a single-purpose box where the address comes from a DHCP reservation, that's a trade I'm happy with.

After a reboot, AdGuard came up on its own, before the Pi even had its IP.

Keeping both fixes in place

Each of these is one line that's easy to "clean up" later without realising what it does, so both have a check in my make verify target. The container check runs Node's resolver rather than dig, because dig passed even while the bug was there:

ssh pi "docker exec uptime-kuma-uptime-kuma-1 node -e \"
  const r = new (require('dns').promises.Resolver)({timeout: 3000, tries: 1});
  r.setServers(['192.168.1.43']);
  r.resolve4('example.com').then(() => process.exit(0), () => process.exit(1))
\"" \
  && echo "ok: containers resolve via 192.168.1.43" \
  || { echo "FAIL: DNS from a container times out (UDP reply from the wrong source?)"; exit 1; }

ssh pi 'cat /proc/sys/net/ipv4/ip_nonlocal_bind' | grep -qx 1 \
  && echo "ok: ip_nonlocal_bind=1" \
  || { echo "FAIL: ip_nonlocal_bind is not 1"; exit 1; }

Why only UDP

The TCP side of port 53 never had this problem. A TCP reply always comes back on the connection the client opened, so its source address is fixed when the connection is set up. With UDP, every reply is a separate packet, and the kernel chooses its source address again each time. Anything you publish over UDP with Docker can hit this, not just DNS: syslog, WireGuard, game servers. It only shows up when the client checks where the reply came from, and that's why it took a Node app to notice.

Share this post

Post on X