public inbox for passt-dev@passt.top
 help / color / mirror / code / Atom feed
From: "Bernhard M. Wiedemann" <bwiedemann@suse.de>
To: passt-dev@passt.top
Cc: "Bernhard M. Wiedemann" <bwiedemann@suse.de>
Subject: [PATCH] tcp_splice: avoid delay on certain transfers
Date: Fri,  2 Oct 2026 18:39:25 +0200	[thread overview]
Message-ID: <20261002163925.1338988-1-bwiedemann@suse.de> (raw)

without this patch, sending buffers of certain sizes
caused a delay of 200ms because tcp_splice_forward()
passes SPLICE_F_MORE to the writer
when a read filled at least 90% of the pipe.

This is easy to hit when pipes are small: once a user exceeds
fs.pipe-user-pages-soft (64 MiB by default, which a few long-running
pasta instances with large pipes reach on their own),
tcp_set_pipe_size() settles on the 8 KiB minimum, and every message
whose length is 7372 to 8192 bytes (90% of the pipe and more) past a
multiple of 8 KiB stalls.

This change leaves bulk throughput unchanged.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Bernhard M. Wiedemann <bwiedemann@suse.de>
---
Notes:
  The change was slightly tested and benchmarked. Results look decent.
  Not sure if we actually need the flow_trace for error handling there.
  This issue was accidentally found by running PostgreSQL in a rootless
    podman container for benchmarking.

 tcp_splice.c | 15 ++++++++++++++-
 1 file changed, 14 insertions(+), 1 deletion(-)

diff --git a/tcp_splice.c b/tcp_splice.c
index 4b01f1a..f7a3913 100644
--- a/tcp_splice.c
+++ b/tcp_splice.c
@@ -488,6 +488,7 @@ static int tcp_splice_forward(struct ctx *c,
 {
 	uint8_t lowat_set_flag = RCVLOWAT_SET(fromsidei);
 	uint8_t lowat_act_flag = RCVLOWAT_ACT(fromsidei);
+	bool corked = false;
 
 	while (1) {
 		ssize_t readlen, written;
@@ -517,8 +518,17 @@ static int tcp_splice_forward(struct ctx *c,
 			 * there's nothing in the pipe so there's nothing to do
 			 * write side either.
 			 */
-			if (!conn->pending[fromsidei])
+			if (!conn->pending[fromsidei]) {
+				/* Setting TCP_NODELAY again flushes data held
+				 * back by SPLICE_F_MORE
+				 */
+				if (corked &&
+				    setsockopt(conn->s[!fromsidei], SOL_TCP,
+					       TCP_NODELAY, &((int){ 1 }),
+					       sizeof(int)))
+					flow_trace(conn, "failed to push data");
 				break;
+			}
 		} else {
 			conn->pending[fromsidei] += readlen;
 
@@ -549,6 +559,9 @@ static int tcp_splice_forward(struct ctx *c,
 		if (written < 0)
 			break;
 
+		if (written > 0)
+			corked = more;
+
 		conn->pending[fromsidei] -= written;
 
 		if (!conn->pending[fromsidei] && readlen <= 0) {
-- 
2.55.0


                 reply	other threads:[~2026-10-02 16:40 UTC|newest]

Thread overview: [no followups] expand[flat|nested]  mbox.gz  Atom feed

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261002163925.1338988-1-bwiedemann@suse.de \
    --to=bwiedemann@suse.de \
    --cc=passt-dev@passt.top \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
Code repositories for project(s) associated with this public inbox

	https://passt.top/passt

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for IMAP folder(s).