From mboxrd@z Thu Jan 1 00:00:00 1970 Authentication-Results: passt.top; dmarc=fail (p=none dis=none) header.from=suse.de Received: from mail.bmwiedemann.de (mail.bmwiedemann.de [IPv6:2a01:4f8:221:b52:fcfd:ff:fe00:ec04]) by passt.top (Postfix) with ESMTPS id EDE125A061A for ; Fri, 02 Oct 2026 18:40:20 +0200 (CEST) Received: from mail.bmwiedemann.de (localhost [127.0.0.1]) by mail.bmwiedemann.de (Postfix) with ESMTP id 31ED28C4; Fri, 02 Oct 2026 16:40:20 +0000 (UTC) X-Spam-Checker-Version: SpamAssassin 4.0.1 (2024-03-26) on vm4c.zq1.de X-Spam-Level: X-Spam-Status: No, score=-1.0 required=5.0 tests=ALL_TRUSTED shortcircuit=no autolearn=unavailable autolearn_force=no version=4.0.1 Received: from zq1.de (unknown [10.8.5.117]) by mail.bmwiedemann.de (Postfix) with ESMTP; Fri, 02 Oct 2026 16:40:20 +0000 (UTC) Received: by zq1.de (Postfix, from userid 1000) id EC6D435AC90; Fri, 02 Oct 2026 18:40:19 +0200 (CEST) From: "Bernhard M. Wiedemann" To: passt-dev@passt.top Subject: [PATCH] tcp_splice: avoid delay on certain transfers Date: Fri, 2 Oct 2026 18:39:25 +0200 Message-ID: <20261002163925.1338988-1-bwiedemann@suse.de> X-Mailer: git-send-email 2.55.0 MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-MailFrom: bernhard@zq1.de X-Mailman-Rule-Hits: nonmember-moderation X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation Message-ID-Hash: DXKPXB4HFKPBF6P2KB4KQJZJJQAK53OY X-Message-ID-Hash: DXKPXB4HFKPBF6P2KB4KQJZJJQAK53OY X-Mailman-Approved-At: Fri, 02 Oct 2026 19:17:11 +0200 CC: "Bernhard M. Wiedemann" X-Mailman-Version: 3.3.8 Precedence: list List-Id: Development discussion and patches for passt Archived-At: Archived-At: List-Archive: List-Archive: List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: without this patch, sending buffers of certain sizes caused a delay of 200ms because tcp_splice_forward() passes SPLICE_F_MORE to the writer when a read filled at least 90% of the pipe. This is easy to hit when pipes are small: once a user exceeds fs.pipe-user-pages-soft (64 MiB by default, which a few long-running pasta instances with large pipes reach on their own), tcp_set_pipe_size() settles on the 8 KiB minimum, and every message whose length is 7372 to 8192 bytes (90% of the pipe and more) past a multiple of 8 KiB stalls. This change leaves bulk throughput unchanged. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Bernhard M. Wiedemann --- Notes: The change was slightly tested and benchmarked. Results look decent. Not sure if we actually need the flow_trace for error handling there. This issue was accidentally found by running PostgreSQL in a rootless podman container for benchmarking. tcp_splice.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-) diff --git a/tcp_splice.c b/tcp_splice.c index 4b01f1a..f7a3913 100644 --- a/tcp_splice.c +++ b/tcp_splice.c @@ -488,6 +488,7 @@ static int tcp_splice_forward(struct ctx *c, { uint8_t lowat_set_flag = RCVLOWAT_SET(fromsidei); uint8_t lowat_act_flag = RCVLOWAT_ACT(fromsidei); + bool corked = false; while (1) { ssize_t readlen, written; @@ -517,8 +518,17 @@ static int tcp_splice_forward(struct ctx *c, * there's nothing in the pipe so there's nothing to do * write side either. */ - if (!conn->pending[fromsidei]) + if (!conn->pending[fromsidei]) { + /* Setting TCP_NODELAY again flushes data held + * back by SPLICE_F_MORE + */ + if (corked && + setsockopt(conn->s[!fromsidei], SOL_TCP, + TCP_NODELAY, &((int){ 1 }), + sizeof(int))) + flow_trace(conn, "failed to push data"); break; + } } else { conn->pending[fromsidei] += readlen; @@ -549,6 +559,9 @@ static int tcp_splice_forward(struct ctx *c, if (written < 0) break; + if (written > 0) + corked = more; + conn->pending[fromsidei] -= written; if (!conn->pending[fromsidei] && readlen <= 0) { -- 2.55.0