Bridge closes long-lived connections without logging the cause #11
Labels
No milestone
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
carvers/silta#11
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Long-lived postgres connections through a bridge get reset after a few minutes.
The client sees "SSL error: unexpected eof while reading"; silta logs only a
plain
INFO closedline, so the log does not say which side closed or why.Setup
127.0.0.1:15432->svc/timescale5432 in contextnetzlive-test (bridged to pod timescale1-1).
Observed
Three connection losses within about 30 minutes on 2026-08-14, all
SSL error: unexpected eof while readingon the client:no bytes on the wire).
VIEW.
Counter-examples on the same forward, same period: short queries always work,
and a deliberate
select pg_sleep(270)survived 4.5 minutes of complete wiresilence. So it is not an idle timeout; it looks like sporadic resets of the
upstream stream (kubelet port-forward leg or API-server path), passed through
as a normal close.
The eof arriving without a TLS close-notify says the remote leg died; silta
itself may only be the messenger. But from silta's log alone this is
indistinguishable from the client hanging up.
Ask
stream error/eof), and the error if there was one.
INFO closedfor both agraceful client disconnect and an upstream failure hides exactly this class
of problem.
rather than folding it into the same close path.
Re-establishing the stream cannot rescue a mid-flight database session, so
honest close reporting is the valuable part; reconnect behavior is secondary.
Root cause found, and it is not the bridge: the postgres container behind the forward was OOMKilled (exit 137, 14:00:13Z) by the heavy statement each session was running. The close at 13:54:41 and the others line up with container kills, and silta correctly logged 'WARN resolve timescale/svc/timescale: no Ready pod behind service' during the restart window.
So the reset behavior in this report is explained and no bridge bug is implied there. What stands is the logging ask: the client-visible failure ('SSL error: unexpected eof') was indistinguishable in silta's log from a graceful client hang-up. An 'upstream closed (eof/reset)' vs 'client closed' distinction on the close line would have cut this investigation from an hour to a minute.