From: Aris Konstantoulas <arist.kon@gmail.com>
To: passt-dev@passt.top
Cc: Stefano Brivio <sbrivio@redhat.com>,
David Gibson <david@gibson.dropbear.id.au>,
Aris Konstantoulas <aris@ariscodes.com>
Subject: [PATCH] tcp: Don't fast re-transmit if only our FIN is outstanding
Date: Mon, 28 Sep 2026 13:31:20 +0300 [thread overview]
Message-ID: <20260928103120.233586-1-arist.kon@gmail.com> (raw)
From: Aris Konstantoulas <aris@ariscodes.com>
In the TAP_FIN_RCVD path of tcp_tap_handler(), a bare segment from the
guest acknowledging exactly seq_ack_from_tap, with an unchanged window,
is taken as a duplicate ACK and triggers a fast re-transmit.
If the only unacknowledged sequence number is our own FIN, that's
harmful: tcp_rewind_seq() rewinds seq_to_tap and clears TAP_FIN_SENT,
so tcp_data_from_sock() immediately sends the FIN again. If the guest
answers that FIN with the same bare ACK, as a socket in TIME-WAIT will,
we loop at packet rate:
- conn->retries is never incremented on this path, so we never reach
TCP_MAX_RETRIES and tcp_rst()
- ACK_FROM_TAP_DUE is re-armed on every iteration, so the backed-off
re-transmission in tcp_timer_handler() never fires
- TAP_FIN_ACKED can't be set, as it requires TAP_FIN_SENT, which the
rewind just cleared
On an idle Podman host (rootless, pasta), this showed up as a single
flow exchanging ~45,000 54-byte segments per second between pasta and
a container whose socket was in TIME-WAIT, with pasta using ~75% of one
core, until the socket was killed by hand. It recurred on the idle
teardown of an HTTP/2 connection to an ACME server.
With a raw-socket peer driving the same sequence against pasta at
f8df3f1, pasta re-sent the FIN 727,509 times in 10 seconds. With this
change it's re-transmitted by the timer at 1, 3, 7, 15, 31, 63 and 127
seconds, and the connection is reset once TCP_MAX_RETRIES is reached.
Don't consider a duplicate ACK as a fast re-transmit trigger if the
only outstanding sequence number is the FIN, and leave it to the timer.
Fast re-transmit of data is unaffected, with or without a FIN queued
after it.
Fixes: bde1847960cf ("tcp: Fast re-transmit if half-closed, make TAP_FIN_RCVD path consistent")
Link: https://bugs.passt.top/show_bug.cgi?id=125
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Aris Konstantoulas <aris@ariscodes.com>
---
Notes:
Notes for reviewers, not for the commit message:
- Reproducer: a raw AF_PACKET TCP peer in pasta's namespace completes a
handshake, closes first, then answers each FIN from pasta with a bare
ACK for seq_ack_from_tap, with an unchanged window. I can post it
(Python, ~200 lines) if useful.
- Open question: I haven't established how a real flow first reaches
the state where the peer ACKs exactly up to, but not past, our FIN.
The live capture showed the TIME-WAIT socket ACKing the sequence
number of pasta's FIN itself. That would fit a FIN re-sent after the
first one was already acknowledged, one sequence number further on,
that is, tcp_rewind_seq() clearing TAP_FIN_SENT with nothing
outstanding. Candidates are the zero-window rewind and the
fast-retransmit check in tcp_data_from_tap(), neither of which checks
that anything is outstanding. This patch breaks the loop in either
case, but not that possible entry path. I'm happy to look into it
further, possibly as part of bug 125.
tcp.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
diff --git a/tcp.c b/tcp.c
index 3b78d2e..e279f8a 100644
--- a/tcp.c
+++ b/tcp.c
@@ -2417,15 +2417,22 @@ int tcp_tap_handler(const struct ctx *c, uint8_t pif, sa_family_t af,
/* Established connections not accepting data from tap */
if (conn->events & TAP_FIN_RCVD) {
+ bool fin_only, retr;
size_t dlen;
- bool retr;
if ((dlen = tcp_packet_data_len(th, l4len))) {
flow_dbg(conn, "data segment in CLOSE-WAIT (%zu B)",
dlen);
}
- retr = th->ack && !th->fin &&
+ /* If only our FIN is outstanding, rewinding on a duplicate ACK
+ * would re-send it on every ACK, bypassing the backoff and
+ * retry limit of tcp_timer_handler(): let the timer do it.
+ */
+ fin_only = (conn->events & TAP_FIN_SENT) &&
+ conn->seq_to_tap == conn->seq_ack_from_tap + 1;
+
+ retr = th->ack && !th->fin && !fin_only &&
ntohl(th->ack_seq) == conn->seq_ack_from_tap &&
ntohs(th->window) == conn->wnd_from_tap;
--
2.43.0
reply other threads:[~2026-09-28 10:31 UTC|newest]
Thread overview: [no followups] expand[flat|nested] mbox.gz Atom feed
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928103120.233586-1-arist.kon@gmail.com \
--to=arist.kon@gmail.com \
--cc=aris@ariscodes.com \
--cc=david@gibson.dropbear.id.au \
--cc=passt-dev@passt.top \
--cc=sbrivio@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
Code repositories for project(s) associated with this public inbox
https://passt.top/passt
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for IMAP folder(s).