From mboxrd@z Thu Jan 1 00:00:00 1970 Authentication-Results: passt.top; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: passt.top; dkim=pass (1024-bit key; unprotected) header.d=redhat.com header.i=@redhat.com header.a=rsa-sha256 header.s=mimecast20190719 header.b=K2HXxANU; dkim-atps=neutral Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) by passt.top (Postfix) with ESMTPS id 75E9A5A0271 for ; Mon, 05 Oct 2026 21:37:17 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1791229036; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=4uRzLy5n11S7XQnMKNF8wxcduaRl0dV6lNYr3xHb6mE=; b=K2HXxANUIFCVfdRlilLGZJ+xEKv+wWE7hlj4b4JYmsnRuS5Qmvu15EWkwRozBD2v4L0qG7 GAV066e8C2/9YGyzHOstc9DYqykYnyUV9iic2YHYvaaLP7RaA6v7/gyl+xYRhEmLZL1OnA V5fHPilisCh4ahcTBu1Y/JdYrGYt9d8= Received: from mail-wr1-f72.google.com (mail-wr1-f72.google.com [209.85.221.72]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-615-ewYnhH56NECRdIAA2v2IDg-1; Mon, 05 Oct 2026 15:37:15 -0400 X-MC-Unique: ewYnhH56NECRdIAA2v2IDg-1 X-Mimecast-MFC-AGG-ID: ewYnhH56NECRdIAA2v2IDg_1791229034 Received: by mail-wr1-f72.google.com with SMTP id ffacd0b85a97d-48c511d9e97so1047793f8f.0 for ; Mon, 05 Oct 2026 12:37:14 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791229034; x=1791833834; h=date:content-transfer-encoding:content-type:mime-version :organization:references:in-reply-to:message-id:subject:cc:to:from :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=4uRzLy5n11S7XQnMKNF8wxcduaRl0dV6lNYr3xHb6mE=; b=tqR3YAf6EcwnkuaaMXakssSqCkTKZsIoWsIpLzQlypUl38ALjX0ku7oZu4GWLsc34x QrYe60QdiPHxkcRtbAGXF7l2J8YCG0ElH9Z8ibo7+S/NRfYN13j+OVwtAS5fS0E7aY4y 4zh6Yoa1T7Va7+OzE1gwcwx00TYFp901h5VIAUqsonpyC5HNZA6R+bx03HI3bVplLrKQ PvcCSFRp/Af6Vcu6s/WFZSHkbIgD/68fxYM0paADkIYlwbphHX/fQo/o0p4VvvMyQ6OF WGuLcS5dGI1aBs81Y10M2vw//Mj8hJ12DtRE2CWEk7KpjoeiYhHZvJtilHtsQ2rPzuyh mBzA== X-Gm-Message-State: AFq9FYL/Vc4+CPNLgwNVrQvtFA7roISufAY+70e0t6zCiYwGbU+BXkb6 vTDUFv2o7I7obTVfrofj7GaQA3/rX1LXpSoP/kEAbwLURI19kVX9OzrmCaqkMKFPr6wOA4LXaR3 L6fnWIUC3nnuzqLUnWEA9Ub2pJMjFX9PoJkCclw550avpbVex2pi7drXQBbja1hNs3odbSBJ1nt jh+iL5CCMi0dNhf+icxh/VhWaTu+QY+1fjThog X-Gm-Gg: AYBFou2eJsZRgU53yyv+fjftctRRs5r1c9GAsLOuh/Ni6IOZ06eudGYI4J8s/oF1tdu twz5YhdzYcPCQ4snW7JjnfqKKamuEv9/NMxSt06eYyxwYCzqNEhdqEolLdhjmxzJnLDm6IxW7rS 5W8cVVvxGaTFQpcdV+MvjyBt5Hj6W8N405xXvs7QaPUB8tKf56vXcReghkI+fzBi6o79VGjjbsq MYwvlMeUdI2Tcb6aS4XqAWQKri4DEgnDqHOlP/OxgpZT+bG292UDhRmwlHd3UNrCLAb3TKLsdoj KEoj3kLSTt1QNe6PT8EMu4+MkETypXjEBQtXclDUgbZpJStTxIW+X+7wXCPA3DoU0L/w9wruBEq 2tCouLFZI5w== X-Received: by 2002:a05:6000:2086:b0:48b:174:eab1 with SMTP id ffacd0b85a97d-48b1271f6f5mr22072133f8f.36.1791229033727; Mon, 05 Oct 2026 12:37:13 -0700 (PDT) X-Received: by 2002:a05:6000:2086:b0:48b:174:eab1 with SMTP id ffacd0b85a97d-48b1271f6f5mr22072069f8f.36.1791229033030; Mon, 05 Oct 2026 12:37:13 -0700 (PDT) Received: from maya.myfinge.rs (ifcgrfdd.trafficplex.cloud. [2a10:fc81:a806:d6a9::1]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-48c63bccd36sm5359366f8f.5.2026.10.05.12.37.12 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 05 Oct 2026 12:37:12 -0700 (PDT) From: Stefano Brivio To: "Bernhard M. Wiedemann" Subject: Re: [PATCH] tcp_splice: avoid delay on certain transfers Message-ID: <20261005213711.23b92b3c@elisabeth> In-Reply-To: <20261002163925.1338988-1-bwiedemann@suse.de> References: <20261002163925.1338988-1-bwiedemann@suse.de> Organization: Red Hat X-Mailer: Claws Mail 4.2.0 (GTK 3.24.49; x86_64-pc-linux-gnu) MIME-Version: 1.0 Date: Mon, 05 Oct 2026 21:37:12 +0200 (CEST) X-Mimecast-Spam-Score: 0 X-Mimecast-MFC-PROC-ID: hPKPGXeeFwq0ZhbT1D4HmZc0In2j9rfOCjdJqJTIvSY_1791229034 X-Mimecast-Originator: redhat.com Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit Message-ID-Hash: OR75P3NSLTT2TUNZLSCX2EH7WC34I3LA X-Message-ID-Hash: OR75P3NSLTT2TUNZLSCX2EH7WC34I3LA X-MailFrom: sbrivio@redhat.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header CC: passt-dev@passt.top X-Mailman-Version: 3.3.8 Precedence: list List-Id: Development discussion and patches for passt Archived-At: Archived-At: List-Archive: List-Archive: List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Bernard, thanks for the investigation and for the patch. Just one doubt: On Fri, 2 Oct 2026 18:39:25 +0200 "Bernhard M. Wiedemann" wrote: > without this patch, sending buffers of certain sizes > caused a delay of 200ms because tcp_splice_forward() > passes SPLICE_F_MORE to the writer > when a read filled at least 90% of the pipe. > > This is easy to hit when pipes are small: once a user exceeds > fs.pipe-user-pages-soft (64 MiB by default, which a few long-running > pasta instances with large pipes reach on their own), > tcp_set_pipe_size() settles on the 8 KiB minimum, and every message > whose length is 7372 to 8192 bytes (90% of the pipe and more) past a > multiple of 8 KiB stalls. > > This change leaves bulk throughput unchanged. > > Assisted-by: Claude:claude-opus-5-5 > Signed-off-by: Bernhard M. Wiedemann > --- > Notes: > The change was slightly tested and benchmarked. Results look decent. > Not sure if we actually need the flow_trace for error handling there. > This issue was accidentally found by running PostgreSQL in a rootless > podman container for benchmarking. > > tcp_splice.c | 15 ++++++++++++++- > 1 file changed, 14 insertions(+), 1 deletion(-) > > diff --git a/tcp_splice.c b/tcp_splice.c > index 4b01f1a..f7a3913 100644 > --- a/tcp_splice.c > +++ b/tcp_splice.c > @@ -488,6 +488,7 @@ static int tcp_splice_forward(struct ctx *c, > { > uint8_t lowat_set_flag = RCVLOWAT_SET(fromsidei); > uint8_t lowat_act_flag = RCVLOWAT_ACT(fromsidei); > + bool corked = false; > > while (1) { > ssize_t readlen, written; > @@ -517,8 +518,17 @@ static int tcp_splice_forward(struct ctx *c, > * there's nothing in the pipe so there's nothing to do > * write side either. > */ > - if (!conn->pending[fromsidei]) > + if (!conn->pending[fromsidei]) { > + /* Setting TCP_NODELAY again flushes data held > + * back by SPLICE_F_MORE ...nice, I didn't know about that trick. But wouldn't it be more natural to not use SPLICE_F_MORE if the pipe is small enough? Do you have a stand-alone reproducer that could help figuring this out? I'm a bit worried we might cause unnecessary setsockopt() calls in some corner cases if we go this way, even though it's a rather minor concern (we just called splice(), and that setsockopt() is not _that_ expensive in comparison). > + */ > + if (corked && > + setsockopt(conn->s[!fromsidei], SOL_TCP, > + TCP_NODELAY, &((int){ 1 }), > + sizeof(int))) > + flow_trace(conn, "failed to push data"); > break; > + } > } else { > conn->pending[fromsidei] += readlen; > > @@ -549,6 +559,9 @@ static int tcp_splice_forward(struct ctx *c, > if (written < 0) > break; > > + if (written > 0) > + corked = more; > + > conn->pending[fromsidei] -= written; > > if (!conn->pending[fromsidei] && readlen <= 0) { -- Stefano