From mboxrd@z Thu Jan 1 00:00:00 1970 Authentication-Results: passt.top; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: passt.top; dkim=pass (2048-bit key; unprotected) header.d=gmail.com header.i=@gmail.com header.a=rsa-sha256 header.s=20251104 header.b=Wn4fv805; dkim-atps=neutral Received: from mail-wm1-x331.google.com (mail-wm1-x331.google.com [IPv6:2a00:1450:4864:20::331]) by passt.top (Postfix) with ESMTPS id 164D15A0274 for ; Sun, 02 Aug 2026 15:23:51 +0200 (CEST) Received: by mail-wm1-x331.google.com with SMTP id 5b1f17b1804b1-49545ba3d4eso4973155e9.3 for ; Sun, 02 Aug 2026 06:23:51 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785677030; x=1786281830; darn=passt.top; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=C0BtjQsSktFMa0oahNejWy9tQuHZb+aDiOJUvKkscUE=; b=Wn4fv805TwIjBYh1l3+i4S+3LMfGwtvlgcPtaHSscJcLLaGu5waNd0gSc7Sc7q+gys rE4QxaDlqRfIUAXe+OYCBfudm4KN93uPTVlOuxW5rmIz1eGAf/PvDFwN6DBcyeAaqXB3 dE/rj9R1Ply13OEhMf3m+rf78v8qmYoFkCEoYP6A5VgVVEaSR4BOZNL/GQJDHTRfyH6R cXFjBBhnNnKnwIFX9r0Y6od2weItTEYw6b4y6UcuuAE0BLOn9TlojmyfmmueEuJnmez1 uN3Oi8GGwnmWmbwKVxEJy+MojXFt7xm776dRIY2/jiZzHeANqN2oh+B4ABBHD9HZD33m i+yQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785677030; x=1786281830; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=C0BtjQsSktFMa0oahNejWy9tQuHZb+aDiOJUvKkscUE=; b=D+BsAU5Z14g59TiTXsyVleXzefGVzymhfvU7JCtuXJq4iMxDwDAL3m/jY+YTVjL+vX 6HchWmT9ifOBQG+LWF6ZmryYZag2K7SmFEZsSQtyGxmV9APFcPE7/RYbC8BRrv2bN05H tJRYKyuh1G8h9qsM/uCuwxdlk7azYm+UVvrquhLMtxWEHl3XUyyzCj0dT/ncLicWYIqF xFJgoYXngfx8MZRCA4okNCzVd1k+G7n83n1wIkk/yUmml5Csgvnd2OHxFGFsogCqrcGV hqghQJJHYra/sgtMf2RP19X/XbCegHgOtRszXNTmfz1gbRnzpE0v3GxogvgFWIjdoRww OevQ== X-Gm-Message-State: AOJu0YxKa2D8z4zAlGTG6BHG6b1DNDH79nMGszcRXHOa+PcLWg2GS9MX k2fLjCX+OLAqQbPffRHbV5lgP4xBcJAYjBWr6PlZVw5U0rdDSFZIRIjpFJfdc52e X-Gm-Gg: AR+sD13/+QD1owYuIBx6bW1WYklWfXybGF7BTbI7SkJApUnnx/VPDoQ+7wsh+ydWCsA Rknvuv+HPCTNA7ogZOD7Qq5dtwC++wS6u98aglMkOqXTxTLEv9k1TRplmbRCNoBrIPzQ/S0ZOrz GlqkTDhCn/Hd6Kw34j4ki52JeOFR9+dloc2D6BSo4y5KT2RIjZMKnvUeOIuF441bvT8leEI+BDg CwIO+Oqy/F9vQQJSWO0fznICkeGZymxahBfn7kyVyT3b3oVGssAEzOrwohFyVKLhC+UedV+rubI lcyojPitDwf/26Yntv8DcdWsIjtu4e/0NINPRgo9D3L4i5AJwEBzdNS3LtH8DUjm4b26+kxOXPq yY6X6BjJaq8Vv3j0n55nsI0c+0tooGL3m2QJlqh3SPY8d6qtHf0jERqc8NpxIwW7Warn+IxaaTw 6Z/iEIcXOKqm+OmALbUEl/2wev/qGmx9p8g8AVkNxHjVUuDDlXt6xRX/F89b3swi1k+sJIKY2XU M/US83K8+NnkvJaXxTWnVc1hRJ7mGNQJEwx X-Received: by 2002:a05:600c:348a:b0:495:6b55:f938 with SMTP id 5b1f17b1804b1-4980c674e72mr128303975e9.10.1785677030422; Sun, 02 Aug 2026 06:23:50 -0700 (PDT) Received: from localhost.localdomain ([196.137.15.128]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4980869106csm180745155e9.9.2026.08.02.06.23.49 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 02 Aug 2026 06:23:50 -0700 (PDT) From: Ammar Yasser To: passt-dev@passt.top Subject: [RFC v3 6/8] virtio: Implement pasta vhost acceleration guest->pasta path Date: Sun, 2 Aug 2026 13:21:53 +0000 Message-Id: <20260802132155.870796-7-aerosound161@gmail.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260802132155.870796-1-aerosound161@gmail.com> References: <20260802132155.870796-1-aerosound161@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Message-ID-Hash: H2THX4ANSO4I6PVEJXKOCMFKEPSMTDGC X-Message-ID-Hash: H2THX4ANSO4I6PVEJXKOCMFKEPSMTDGC X-MailFrom: aerosound161@gmail.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header CC: eperezma@redhat.com, Ammar Yasser X-Mailman-Version: 3.3.8 Precedence: list List-Id: Development discussion and patches for passt Archived-At: Archived-At: List-Archive: List-Archive: List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Add a new function called tap_vhost_input that will be used to indicate that the guest is trying to send data to us through the rx queue and the kernel is informing us to handle this data. We will however still recieve epoll events on the fd normally. We should explicitly use either the vhost path of the tap_handler_pasta path and so if vhost was requested the tap_handler_pasta path gets skipped. Add vhost bootstrapping to tap_sock_tun_init in case vhost was requested. consume_one_rx_descriptor: fetch one descriptor from what the kernel has advertised to us was written to, advance the tracking data structure to reflect the last index we used, and the count of descriptors we are ready to hand back to the kernel. return a void pointer to the start of this descriptor's data tap_vhost_input: fetch descriptors using consume_one_rx_descriptor, and skip over the vnet header, then build an iov tail from data and queue it for processing, process the frames, and lastly, advertise to the kernel that the descriptors are free to reuse Signed-off-by: Ammar Yasser --- passt.c | 6 ++++ tap.c | 109 +++++++++++++++++++++++++++++++++++++++++++++++++++++--- tap.h | 1 + 3 files changed, 112 insertions(+), 4 deletions(-) diff --git a/passt.c b/passt.c index 865b331..fdf9070 100644 --- a/passt.c +++ b/passt.c @@ -305,6 +305,12 @@ static void passt_worker(void *opaque, int nfds, struct epoll_event *events) case EPOLL_TYPE_CONF: conf_handler(c, eventmask); break; + case EPOLL_TYPE_VHOST_CALL: + tap_vhost_input(c, ref, &now); + break; + case EPOLL_TYPE_VHOST_ERROR: + die("Error on vhost-kernel socket"); + break; default: /* Can't happen */ assert(0); diff --git a/tap.c b/tap.c index dfa66c7..9d601a2 100644 --- a/tap.c +++ b/tap.c @@ -13,6 +13,7 @@ * */ +#include "common.h" #include #include #include @@ -38,6 +39,7 @@ #include #include #include +#include #include #include @@ -61,6 +63,7 @@ #include "vhost_user.h" #include "vu_common.h" #include "epoll_ctl.h" +#include "virtio.h" /* Maximum allowed frame lengths (including L2 header) */ @@ -1359,7 +1362,8 @@ void tap_handler_pasta(struct ctx *c, uint32_t events, if (events & (EPOLLRDHUP | EPOLLHUP | EPOLLERR)) die("Disconnect event on /dev/net/tun device, exiting"); - if (events & EPOLLIN) + /* don't proceed with the normal tap processing in case vhost acceleration was required */ + if (events & EPOLLIN && !c->vhost) tap_pasta_input(c, now); } @@ -1514,6 +1518,90 @@ void tap_listen_handler(struct ctx *c, uint32_t events) tap_start_connection(c); } +/** + * consume_one_rx_descriptor() - Consume one used RX descriptor from the kernel + * @len: Set to the length of data written by the kernel + * + * Pops a single entry from the used ring. Advances vqs[0].last_used_idx + * (the number of entries we have consumed) and vqs[0].num_free (the count + * of descriptors awaiting refill announcement). + * + * NOTE: This function assumes the kernel is going to post single descriptors + * always, No chains. If that changes, we would need to increment num_free + * as we advertise back to the kernel the free descriptors by the length of the chain. + * + * Return: pointer to the packet buffer, or NULL if no data is available + */ +static void *consume_one_rx_descriptor(unsigned *len) +{ + struct vring_used *used = &vring_used_all[0].used; + uint32_t i; + uint16_t used_idx, last_used; + + used_idx = le16toh(used->idx); + + smp_rmb(); + + /* if the kernel's last_used index matches our last_used_idx, then + * we have finished consuming data. + */ + if (used_idx == vqs[0].last_used_idx) { + *len = 0; + return NULL; + } + + /* read the last index we consumed */ + last_used = vqs[0].last_used_idx % VHOST_NDESCS; + /* read the index of what the */ + i = le32toh(used->ring[last_used].id); + *len = le32toh(used->ring[last_used].len); + + if (i != last_used) { + die("vhost: id %u at used position %u != %u", i, last_used, i); + } + + /* the kernel has queued for us something we cannot receive */ + if (*len > PKT_BUF_BYTES/VHOST_NDESCS) { + die("vhost: id %d len %u > %zu", i, *len, PKT_BUF_BYTES/VHOST_NDESCS); + } + + vqs[0].last_used_idx++; + vqs[0].num_free++; + return pkt_buf + i * (PKT_BUF_BYTES/VHOST_NDESCS); +} + + +/** + * tap_vhost_input() - Handler for new data on the tun socket to hypervisor vq + * @c: Execution context + * @ref: epoll reference + * @now: Current timestamp + */ +void tap_vhost_input(struct ctx *c, union epoll_ref ref, const struct timespec *now) +{ + eventfd_read(ref.fd, (eventfd_t[]){ 0 }); + + tap_flush_pools(); + + struct virtio_net_hdr_mrg_rxbuf *hdr; + struct iov_tail data; + unsigned len; + + while ((hdr = consume_one_rx_descriptor(&len))) { + if (len < sizeof(*hdr)) { + warn("vhost: invalid len %u", len); + continue; + } + + /* skip over the vnet header, we wanna add the packet without it*/ + data = IOV_TAIL_FROM_BUF((void *)(hdr+1), len - sizeof(*hdr), 0); + tap_add_packet(c, &data, now); + } + + tap_handler(c, now); + rx_descriptor_handoff(c); +} + /** * tap_ns_tun() - Get tuntap fd in namespace * @c: Execution context @@ -1524,16 +1612,15 @@ void tap_listen_handler(struct ctx *c, uint32_t events) */ static int tap_ns_tun(void *arg) { - struct ifreq ifr = { .ifr_flags = IFF_TAP | IFF_NO_PI }; - int flags = O_RDWR | O_NONBLOCK | O_CLOEXEC; struct ctx *c = (struct ctx *)arg; + struct ifreq ifr = { .ifr_flags = IFF_TAP | IFF_NO_PI }; int fd, rc; c->fd_tap = -1; memcpy(ifr.ifr_name, c->pasta_ifn, IFNAMSIZ); ns_enter(c); - fd = open("/dev/net/tun", flags); + fd = open("/dev/net/tun", O_RDWR | O_NONBLOCK | O_CLOEXEC); if (fd < 0) die_perror("Failed to open() /dev/net/tun"); @@ -1561,6 +1648,20 @@ static void tap_sock_tun_init(struct ctx *c) die("Failed to set up tap device in namespace"); } + /* initialize the vhost-net dev file descriptor */ + if (c->vhost) { + setup_vhost_net(c); + + for (int i = 0; i < ARRAY_SIZE(c->vq); i++) + setup_eventfds(c, i); + + if (setup_memory_table(c) < 0) + die_perror("VHOST_SET_MEM_TABLE ioctl on /dev/vhost-net failed"); + + for (int i = 0; i < ARRAY_SIZE(c->vq); i++) + set_vring_for_queue(c, i, c->fd_tap); + } + pasta_ns_conf(c); if (!c->splice_only) diff --git a/tap.h b/tap.h index 1625975..eb02da8 100644 --- a/tap.h +++ b/tap.h @@ -66,6 +66,7 @@ static inline void tap_hdr_update(struct tap_hdr *thdr, size_t l2len) thdr->vnet_len = htonl(l2len); } +void tap_vhost_input(struct ctx *c, union epoll_ref ref, const struct timespec *now); unsigned long tap_l2_max_len(const struct ctx *c); void *tap_push_l2h(const struct ctx *c, void *buf, const void *src_mac, uint16_t proto); -- 2.34.1