Bridge closes long-lived connections without logging the cause #11

Closed
opened 2026-08-14 13:58:52 +00:00 by aav · 1 comment
Owner

Long-lived postgres connections through a bridge get reset after a few minutes.
The client sees "SSL error: unexpected eof while reading"; silta logs only a
plain INFO closed line, so the log does not say which side closed or why.

Setup

  • silta 0.1.0, forward 127.0.0.1:15432 -> svc/timescale 5432 in context
    netzlive-test (bridged to pod timescale1-1).
  • Client: psql 17 over TLS (sslmode require), running a long SQL script.

Observed

Three connection losses within about 30 minutes on 2026-08-14, all
SSL error: unexpected eof while reading on the client:

  1. ~13:2x UTC, during a multi-minute CREATE MATERIALIZED VIEW (server busy,
    no bytes on the wire).
  2. ~13:4x UTC, seconds into a fresh session, on a trivial CREATE OR REPLACE
    VIEW.
  3. ~13:54:41 UTC, mid-script; matches this bridge log line:
2026-08-14T13:54:41.162015Z  INFO silta::bridge: [127.0.0.1:15432/62060] closed

Counter-examples on the same forward, same period: short queries always work,
and a deliberate select pg_sleep(270) survived 4.5 minutes of complete wire
silence. So it is not an idle timeout; it looks like sporadic resets of the
upstream stream (kubelet port-forward leg or API-server path), passed through
as a normal close.

The eof arriving without a TLS close-notify says the remote leg died; silta
itself may only be the messenger. But from silta's log alone this is
indistinguishable from the client hanging up.

Ask

  1. Log the close cause: which side initiated (local client eof, upstream
    stream error/eof), and the error if there was one. INFO closed for both a
    graceful client disconnect and an upstream failure hides exactly this class
    of problem.
  2. If the upstream stream reports an error, consider logging it at WARN/ERROR
    rather than folding it into the same close path.

Re-establishing the stream cannot rescue a mid-flight database session, so
honest close reporting is the valuable part; reconnect behavior is secondary.

Long-lived postgres connections through a bridge get reset after a few minutes. The client sees "SSL error: unexpected eof while reading"; silta logs only a plain `INFO closed` line, so the log does not say which side closed or why. ## Setup - silta 0.1.0, forward `127.0.0.1:15432` -> `svc/timescale` 5432 in context netzlive-test (bridged to pod timescale1-1). - Client: psql 17 over TLS (sslmode require), running a long SQL script. ## Observed Three connection losses within about 30 minutes on 2026-08-14, all `SSL error: unexpected eof while reading` on the client: 1. ~13:2x UTC, during a multi-minute CREATE MATERIALIZED VIEW (server busy, no bytes on the wire). 2. ~13:4x UTC, seconds into a fresh session, on a trivial CREATE OR REPLACE VIEW. 3. ~13:54:41 UTC, mid-script; matches this bridge log line: ``` 2026-08-14T13:54:41.162015Z INFO silta::bridge: [127.0.0.1:15432/62060] closed ``` Counter-examples on the same forward, same period: short queries always work, and a deliberate `select pg_sleep(270)` survived 4.5 minutes of complete wire silence. So it is not an idle timeout; it looks like sporadic resets of the upstream stream (kubelet port-forward leg or API-server path), passed through as a normal close. The eof arriving without a TLS close-notify says the remote leg died; silta itself may only be the messenger. But from silta's log alone this is indistinguishable from the client hanging up. ## Ask 1. Log the close cause: which side initiated (local client eof, upstream stream error/eof), and the error if there was one. `INFO closed` for both a graceful client disconnect and an upstream failure hides exactly this class of problem. 2. If the upstream stream reports an error, consider logging it at WARN/ERROR rather than folding it into the same close path. Re-establishing the stream cannot rescue a mid-flight database session, so honest close reporting is the valuable part; reconnect behavior is secondary.
Author
Owner

Root cause found, and it is not the bridge: the postgres container behind the forward was OOMKilled (exit 137, 14:00:13Z) by the heavy statement each session was running. The close at 13:54:41 and the others line up with container kills, and silta correctly logged 'WARN resolve timescale/svc/timescale: no Ready pod behind service' during the restart window.

So the reset behavior in this report is explained and no bridge bug is implied there. What stands is the logging ask: the client-visible failure ('SSL error: unexpected eof') was indistinguishable in silta's log from a graceful client hang-up. An 'upstream closed (eof/reset)' vs 'client closed' distinction on the close line would have cut this investigation from an hour to a minute.

Root cause found, and it is not the bridge: the postgres container behind the forward was OOMKilled (exit 137, 14:00:13Z) by the heavy statement each session was running. The close at 13:54:41 and the others line up with container kills, and silta correctly logged 'WARN resolve timescale/svc/timescale: no Ready pod behind service' during the restart window. So the reset behavior in this report is explained and no bridge bug is implied there. What stands is the logging ask: the client-visible failure ('SSL error: unexpected eof') was indistinguishable in silta's log from a graceful client hang-up. An 'upstream closed (eof/reset)' vs 'client closed' distinction on the close line would have cut this investigation from an hour to a minute.
aav closed this issue 2026-08-14 14:09:48 +00:00
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
carvers/silta#11
No description provided.