public inbox for passt-dev@passt.top
 help / color / mirror / code / Atom feed
* [PATCH] tcp_splice: avoid delay on certain transfers
@ 2026-10-02 16:39 Bernhard M. Wiedemann
  0 siblings, 0 replies; only message in thread
From: Bernhard M. Wiedemann @ 2026-10-02 16:39 UTC (permalink / raw)
  To: passt-dev; +Cc: Bernhard M. Wiedemann

without this patch, sending buffers of certain sizes
caused a delay of 200ms because tcp_splice_forward()
passes SPLICE_F_MORE to the writer
when a read filled at least 90% of the pipe.

This is easy to hit when pipes are small: once a user exceeds
fs.pipe-user-pages-soft (64 MiB by default, which a few long-running
pasta instances with large pipes reach on their own),
tcp_set_pipe_size() settles on the 8 KiB minimum, and every message
whose length is 7372 to 8192 bytes (90% of the pipe and more) past a
multiple of 8 KiB stalls.

This change leaves bulk throughput unchanged.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Bernhard M. Wiedemann <bwiedemann@suse.de>
---
Notes:
  The change was slightly tested and benchmarked. Results look decent.
  Not sure if we actually need the flow_trace for error handling there.
  This issue was accidentally found by running PostgreSQL in a rootless
    podman container for benchmarking.

 tcp_splice.c | 15 ++++++++++++++-
 1 file changed, 14 insertions(+), 1 deletion(-)

diff --git a/tcp_splice.c b/tcp_splice.c
index 4b01f1a..f7a3913 100644
--- a/tcp_splice.c
+++ b/tcp_splice.c
@@ -488,6 +488,7 @@ static int tcp_splice_forward(struct ctx *c,
 {
 	uint8_t lowat_set_flag = RCVLOWAT_SET(fromsidei);
 	uint8_t lowat_act_flag = RCVLOWAT_ACT(fromsidei);
+	bool corked = false;
 
 	while (1) {
 		ssize_t readlen, written;
@@ -517,8 +518,17 @@ static int tcp_splice_forward(struct ctx *c,
 			 * there's nothing in the pipe so there's nothing to do
 			 * write side either.
 			 */
-			if (!conn->pending[fromsidei])
+			if (!conn->pending[fromsidei]) {
+				/* Setting TCP_NODELAY again flushes data held
+				 * back by SPLICE_F_MORE
+				 */
+				if (corked &&
+				    setsockopt(conn->s[!fromsidei], SOL_TCP,
+					       TCP_NODELAY, &((int){ 1 }),
+					       sizeof(int)))
+					flow_trace(conn, "failed to push data");
 				break;
+			}
 		} else {
 			conn->pending[fromsidei] += readlen;
 
@@ -549,6 +559,9 @@ static int tcp_splice_forward(struct ctx *c,
 		if (written < 0)
 			break;
 
+		if (written > 0)
+			corked = more;
+
 		conn->pending[fromsidei] -= written;
 
 		if (!conn->pending[fromsidei] && readlen <= 0) {
-- 
2.55.0


^ permalink raw reply	[flat|nested] only message in thread

only message in thread, other threads:[~2026-10-02 16:40 UTC | newest]

Thread overview: (only message) (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 16:39 [PATCH] tcp_splice: avoid delay on certain transfers Bernhard M. Wiedemann

Code repositories for project(s) associated with this public inbox

	https://passt.top/passt

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for IMAP folder(s).