From mboxrd@z Thu Jan 1 00:00:00 1970 Authentication-Results: passt.top; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: passt.top; dkim=pass (1024-bit key; unprotected) header.d=redhat.com header.i=@redhat.com header.a=rsa-sha256 header.s=mimecast20190719 header.b=fXk1RBGB; dkim-atps=neutral Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) by passt.top (Postfix) with ESMTPS id 8E47A5A0265 for ; Fri, 14 Aug 2026 13:55:30 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1786708529; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=9kdweg/q7bOrZ1C6SKz21movWoV2dezojyY5DyeSGhI=; b=fXk1RBGBtAwkMjXijVlkwyXwqMr2t8JX7KW5dvZlFyXRAGg+P8D6Fjq4XXNsN2Vuyzguz6 /F2JcuAEg4N5GTH0PjVQLUhzJvFEjQ5G8ZwliDqhcdlnYothVpCzRCiUXTpZrdSsEhtCVy 2KoVE+oNZ9EY/gssrnJbVpon1NRs63k= Received: from mail-wm1-f70.google.com (mail-wm1-f70.google.com [209.85.128.70]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-130-tpjISZGvNmi7VQaehvB7Ug-1; Fri, 14 Aug 2026 07:55:28 -0400 X-MC-Unique: tpjISZGvNmi7VQaehvB7Ug-1 X-Mimecast-MFC-AGG-ID: tpjISZGvNmi7VQaehvB7Ug_1786708527 Received: by mail-wm1-f70.google.com with SMTP id 5b1f17b1804b1-4954b1c6310so8997875e9.0 for ; Fri, 14 Aug 2026 04:55:27 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786708527; x=1787313327; h=date:content-transfer-encoding:content-type:mime-version :organization:references:in-reply-to:message-id:subject:cc:to:from :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=9kdweg/q7bOrZ1C6SKz21movWoV2dezojyY5DyeSGhI=; b=lRYm8CRBAplxhIKATIw2Zrt9vJdnMHoYnaMDjf2AE+RFIU559os71s8GP/swbNn1TD 9da4wJCMyFFTLhIvaJBv88NZhKFeVncyk7ChIAF7kg1eJV1ut+xJ/bEk6aqeqeuJpX2b 9uBMM+tu9B7boxknlhMa/yRy8gjbQmBpfDTGBR7V3DmT2Anrja2wHPbhAUZ0+doTPJcJ w3szlQ9NT6RzFGd+w4S8na5Il66S4nYLYLA+TbEieNSZjg/2t0n5J+b9dK9qTARCS0IQ XtTYmdGReZzfvBYyknPOKP086BLgoblm7p8KM6U9NIfnJ7nJ+fGumZ2MraK5W9/ozMF/ xiwQ== X-Forwarded-Encrypted: i=1; AHgh+RqmijK9iD4WGEeylN4QL+rGukz3Y72/HMZxj26PsHvhwuNTm98r4zdW53AteRIsJ0+Vvd9Ea1lceoQ=@passt.top X-Gm-Message-State: AOJu0Ywly0ALHXEKmD2AZFvSEYkpF7RADaKIpohpwtkjj9LCzrZpcV/X AN6U0zK5F87mUgQSxrB7bYNXd5zOVhKMn/c7+hWtPDwa7FLMLER5FPoOkiQEMAObpxCobFKSZdv z0WoJoTFZZkT7vIwZlbbGXUF6o3ALrooYY0XaUN8OYq5l0bcqeAZX/w== X-Gm-Gg: AR+sD12lPJ7VA+mVeyuHUqtOCgzYNJsM22AJ12ThuBeQ8Tfd5YqsLxsHXyw6nr1VfGp 2yTn3q+p113dxA97y6KTFDTmTGcSkKYuJ2KBDRhRSECxKT9XOo14GQJFRnAOO7R3CCpbU8CKCX0 WXru2nxkHPYMBcvkZ9CFbXxd+FtlhwLmRc1Q01Qr1GkISY0tJdOibpRghKYUzaVqn2CKWBz4iKR 4Lys9WjSgI7kRcVIjHRIl6u/8lqbUN/1PI4IWZ3mVQBaW1FlIItYi09e6EQ2G9W+NCUnay/HSFJ qeJxN6DexdxBf27icMqqNqEuwDbb+J6toCZwrNiXmx7ActxiJtzDjQPpASYHAF06Ad7zIaPsgKy EHFYQNd4raszgCHKD63UeL3pKeQQd X-Received: by 2002:a05:600c:871b:b0:499:8b00:5261 with SMTP id 5b1f17b1804b1-4998b0052cdmr11663425e9.8.1786708526352; Fri, 14 Aug 2026 04:55:26 -0700 (PDT) X-Received: by 2002:a05:600c:871b:b0:499:8b00:5261 with SMTP id 5b1f17b1804b1-4998b0052cdmr11662095e9.8.1786708525618; Fri, 14 Aug 2026 04:55:25 -0700 (PDT) Received: from maya.myfinge.rs (ifcgrfdd.trafficplex.cloud. [176.103.220.4]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-4815f2b1fc2sm7851154f8f.20.2026.08.14.04.55.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 14 Aug 2026 04:55:24 -0700 (PDT) From: Stefano Brivio To: David Gibson Subject: Re: [PATCH 5/5] fuzz: Add test server for bidirectional protocol fuzzing Message-ID: <20260814135523.5aace311@elisabeth> In-Reply-To: References: <20260812072630.3235261-1-anskuma@redhat.com> <20260812072630.3235261-6-anskuma@redhat.com> <20260813095323.4c2719bd@elisabeth> <20260814093558.41626ef3@elisabeth> Organization: Red Hat X-Mailer: Claws Mail 4.2.0 (GTK 3.24.49; x86_64-pc-linux-gnu) MIME-Version: 1.0 Date: Fri, 14 Aug 2026 13:55:24 +0200 (CEST) X-Mimecast-Spam-Score: 0 X-Mimecast-MFC-PROC-ID: DkgEAX7MEf4ll4DhdcaWu-uI8ZAZhjajhijzqHjpgQE_1786708527 X-Mimecast-Originator: redhat.com Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit Message-ID-Hash: OQXMYKUSQM2KQFTXAIE7RZPTZ2O5ZZIW X-Message-ID-Hash: OQXMYKUSQM2KQFTXAIE7RZPTZ2O5ZZIW X-MailFrom: sbrivio@redhat.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; emergency; loop; banned-address; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header CC: Anshu Kumari , passt-dev@passt.top, aerosound161@gmail.com, abdobngad@gmail.com, lvivier@redhat.com X-Mailman-Version: 3.3.8 Precedence: list List-Id: Development discussion and patches for passt Archived-At: Archived-At: List-Archive: List-Archive: List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On Fri, 14 Aug 2026 20:02:46 +1000 David Gibson wrote: > On Fri, Aug 14, 2026 at 09:35:59AM +0200, Stefano Brivio wrote: > > On Fri, 14 Aug 2026 15:40:35 +1000 > > David Gibson wrote: > > > > > On Thu, Aug 13, 2026 at 09:53:24AM +0200, Stefano Brivio wrote: > > > > On Wed, 12 Aug 2026 12:56:28 +0530 > > > > Anshu Kumari wrote: > > > > > > > > > Add fuzz-server that acts as passt's network peer > > > > > during fuzzing. > > > > > > > > To me, this part makes sense. But this one: > > > > > > > > > It connects to passt's UNIX socket > > > > > > > > much less, while: > > > > > > > > > and listens on 127.0.0.1:9999 for TCP connections. > > > > > > > > this is the part that I expected instead. Otherwise it's not just > > > > passt's network peer, it's the guest as well. > > > > > > > > > UNIX socket path: responds to ARP requests and TCP SYNs with > > > > > stateless replies (swapped addresses, fixed ISN). Responses > > > > > are XOR'd with AFL++ shared memory data so the fuzzer can > > > > > mutate server behavior. > > > > > > > > This looks rather complicated to me. > > > > > > It does. > > > > > > > The approach I was suggesting with a test server is the following: > > > > > > > > > > > > ,- exchanges guest-side data with ------------. > > > > | ,---------|---------. > > > > | ,--| passt | > > > > | / '-.---------------^-' > > > > ,---|---. / | connect(), | accept(), > > > > | AFL++ |-- shares memory with ---| | send data, | reply with > > > > '---|---' \ | etc. | data, etc. > > > > | \ ,-v---------------'-. > > > > | '--| test server | > > > > | '---------|---------' > > > > '- exchanges host-side data with -------------' > > > > > > > > ...at least in its basic form. Eventually, the test server should be > > > > able to connect to passt itself (and we could call it "test peer" at > > > > that point). > > > > > > So, I agree that to meaningfully fuzz things, we want the fuzzer to be > > > able to control data on both the guest and host side. I can see two > > > basic approaches: > > > > > > A) Alter passt/pasta so that instead of directly communicating with > > > external entities (either guest or host side) we use mocked > > > versions which retrieve data from AFL. We can do that either at > > > the system call level, or at a higher helper function level, the > > > lower level we go, the more of the "normal" passt code we're > > > exercising, but doing at a slightly higher level might be easier > > > to implement > > > > > > B) Run passt/pasta in an environment where we can intercept the > > > external transfers to a test server / test peer / test guest which > > > in turn responds based on data from AFL. > > > > > > Both the current draft and the sketch diagram Stefano has provided are > > > a hybrid of both approaches, so far I'm not seeing a clear advantage > > > to that over going all one way or the other. > > > > Of course, B) would be cleaner (and not that complicated, see below), > > but we don't want to do that guest-side (sending data from a test > > guest) because we would lose the speed advantage of having shared > > memory on the path that _really_ matters for fuzzing (the guest is > > untrusted, the kernel isn't). > > I agree we lose the speed advantage, but I don't really see why that > matters more on the guest side than the host. Because AFL++ tries to vary the input to discover new code paths, but if it's not necessary (and in general it's not for host-side payload), the input might remain relatively constant, and memory content that isn't changed it's cache hot. I have only profiled the "guest" side of things so far, though. In any case, this is minor. If we can have shared memory and no further transport on one path, it's better than having it on zero paths. > Yes, fuzzing the guest > side matters more, but most guest side operations will induce passt to > perform a host side operation, so the speed of the host side handling > matters even if the guest side is what we care about fuzzing. > > > That's something we already established a while ago when AbdAlRahman > > was working on it. We hadn't really looked into the host side yet, back > > then. > > > > So, host side: we can't do it (and it's much less important) because we > > need to use those sockets in the same way passt uses them. > > I'm not sure what you mean. I outlined a way to do this below. Sure, strictly speaking, we can, but we can't if we want to obtain a realistic approximation of what we would be normally doing on sockets. We would wrap all the host-side socket operations, in that case, which is really not ideal, to the point of questioning the whole effectiveness. > > The guest side interface is a trivial recv(), the host side is something > > complicated with iovecs and everything. > > I'm now not sure if you're saying this in relation to approach B, or > approach A. That's independent from the chosen approach: it's like that without fuzzing. If we wrap / skip recv() on the guest side, that's not a big loss. If we wrap host-side socket operations, that's a significant deviation and we'll miss bugs. > > > (B) is quite easy to do guest side - we just connect a "test guest" to > > > passt's socket. Approach B is much harder for host side. The current > > > draft has a test server listening on a single address. > > > > This is just to get something up and running though, it obviously needs > > to be changed later. > > I don't see how we are "up and running" if fuzzed packets from the > guest induce passt to forward them to random host side addresses that > we're not controlling - we won't generate reproducible results. If you look at patch 4/5, it hardcodes the destination port and uses a known destination address (again, for the moment). It's not random. It will need to be random of course. > > > But that means > > > the fuzzing is fundamentally incapable of finding bugs involving > > > talking to multiple peers at once. Worse, we have to constrain the > > > construction of guest side data so that we talk to the test server not > > > something else, and that's one of the things we most want to fuzz. > > > > > > To really take approach B for the host side we'd need to intercept > > > *all* host side network traffic regardless of address. Probably > > > easiest way to do that would be put the whole thing inside another > > > netns. That outside netns would have a default route to a tap device, > > > and on the other side of the tap device would be a test server serving > > > up frames built from the fuzzer output. > > > > As I was mentioning, this could be done in a network namespace without > > any interface, by making the test server listen to all ports and all > > addresses, with a non-local bind and a so-called AnyIP route. Tested: > > > > $ pasta -- sh -c 'ip route add local default dev lo; nc -l 1 & { sleep 1; echo x | nc -N 1.2.3.4 1; }' > > x > > True, but having the fuzzer synthesize L2 frames Slightly easier, perhaps, but that's beyond the scope of fuzzing we need (AFL++ will start trying to build malformed frames which we won't see anyway and effectively fuzz the kernel instead). > seems easier to me > than having it directly synthesize the various socket operations the > host side peer might perform. I don't see this as complicated. There just needs to be a way for AFL++ to say how much payload the test server might need to send in a given round, and that can be used as argument to send() or sendmsg(). That's what I meant by "play script." > > > It's probably easier to co-ordinate if the host side and guest side > > > test server is the same, so we'd have: > > > > > > ,-------. ,-------------. > > > | AFL++ |-- shares memory with ---| test peer | > > > '-------' '--v------v---' > > > | | > > > /-tap device-/ | > > > | | > > > /- test netns ----------^-------------------|-------------\ > > > | | | > > > | ,--------. | | > > > | | passt >----------------/ | > > > | '--------' | > > > \---------------------------------------------------------/ > > > > > > The test peer generates host side frames via the tap device, and guest > > > side frames via the Unix socket. > > > > The UNIX socket is something we want to avoid, it's really much slower > > compared to shared memory (we tried something like that) on the path > > where AFL++ is trying to mutate data fast (because it can hit a lot of > > different code paths with small changes, compared to changing > > socket-side payload). > > Right. That's why I conclude approch A is probably better further > down. But even if we only care about fuzzing guest side, we'll > usually incur the cost of operations on both sides, so I don't see > that using shared memory is more important for guest side than host > side interposition. > > > > It could also be done with pasta, like this: > > > > > > ,-------. ,-------------. > > > | AFL++ |-- shares memory with ---| test peer | > > > '-------' '--v------v---' > > > | | > > > /-tap device-/ packet > > > | socket > > > /- test netns ----------^-------------------|-------------\ > > > | | | > > > | ,--------. ,----------------|---. | > > > | | pasta >-tap device-< guest netns * | | > > > | '--------' '--------------------' | > > > \---------------------------------------------------------/ > > > > > > > > > The order it generates host vs. guest frames should also come from the > > > fuzzer, not be fixed. At least theoretically, this is non-invasive: > > > it could run with an unmodified passt/pasta. Except that > > > - AFL would still need coverage feedback from passt/pasta > > > - It wouldn't allow us to simulate odd timings (except by actually > > > expending real time) > > > - The various interposing layers will probably slow down fuzzing. > > > > > > So, I rather suspect it will work better to go fully to approach A: no > > > test peer at all, instead passt itself is modified to use mocked > > > versions of all the external syscalls to slurp data from AFL. The > > > draft series already does this for epoll_wait() and recv*(), but we'd > > > need to also do that for recv*() on the tap socket, connect(), > > > accept(), TCP_INFO and probably others. We'd also need to mock > > > "sending" calls, send(), write() and shutdown() at least - but those > > > could probably be no-ops. > > > > ...except that by mocking all those we lose a lot of complexity where > > historically we had a ton of bugs. If we just mock recv() it's much > > less (well yes we had bugs there as well but it was like 3 or 4 over the > > entire project history). > > If we mock as close as possible to the syscall level, I don't see that > bypass much of our complexity. To be clear, I'm suggesting mocks > where both returned data and error codes are derived from the fuzzer, > not just no-op stubs. Doing that is not a realistic test though, because the kernel won't return random error codes. Indeed it would be nice to be robust to kernel issues, but I don't see it as a priority (and I guess we would waste the whole time on those "issues" if we do that). > > > This is more invasive, of course. It also means we need to deal with > > > the case where AFL generates a syscall results that should be > > > impossible - that should move onto the next case ASAP, but not be > > > flagged as a passt bug. On the other hand, this approach should be > > > fast, and since we can also mock clock_gettime(), the fuzzer can > > > potentially find timer logic bugs that would only occur after hours or > > > days in real time. > > > > > > > As far as I understood, it's not trivial to make the same instance > > > > of AFL++ share memory with two processes at the same time, so the > > > > memory-sharing path might need to take a more complicated turn, for > > > > example there could be a wrapper starting both passt and the test > > > > server and sharing memory with them, or passt could _additionally_ > > > > (using a special out-of-band fuzzing channel) share data from AFL++ > > > > with the test server. > > > > > > > > An example of communication below (but events don't necessarily need > > > > to be in this order, this is just an example). For simplicity, let's > > > > ignore the fact that AFL++ might not directly share memory with passt > > > > and test server, and assume there are three areas of memory that > > > > AFL++ directly controls: > > > > > > > > a. shared with passt: an array of struct epoll_event, 'ev' > > > > > > > > b. shared with passt: the kind of tap-side buffer you implemented in > > > > 4/5, 'buf' > > > > > > > > c. shared with the test server: a separate buffer, 'test_buf' > > > > > > > > Example: > > > > > > > > 1. AFL++ writes an EPOLLIN event in 'ev' with type > > > > EPOLL_TYPE_TAP_PASST, of some data in 'buf', and some data in > > > > 'test_buf' > > > > > > > > 2. AFL++ starts passt and the test server > > > > > > > > 3. passt reads the EPOLL_TYPE_TAP_PASST event from 'ev', reads data > > > > from 'buf' and hands it to passt_tap_handler() > > > > > > > > 4. this happens to be have Ethernet, IP, and TCP headers, with the > > > > SYN flag set, and destination address set to the address of the > > > > test server (we might want to force all this, at least initially, > > > > or give it as a hint to AFL++ somehow), so passt connects to > > > > the test server > > > > > > > > 5. the test server accepts the connection, and sends the contents > > > > of 'test_buf' on it (for the test server, this is directly > > > > payload, without headers, as they don't make sense there). I'm > > > > not sure if we should have a different set of events (maybe we > > > > need a "play script" for the server, in case?) > > > > > > > > 6. this generates an EPOLLOUT event for passt. It's not in 'ev', > > > > it's a regular epoll_wait() (I think we could have an > > > > epoll_wait() loop where we additionally read one event from > > > > 'ev' for every iteration, or something like that) > > > > > > > > 7. passt marks the connection as established and inserts it in the > > > > flow table > > > > > > > > 8. passt reads the data sent from the test server and generates > > > > whatever TCP data packet to the "guest" (it might simply be > > > > a sink) > > > > > > > > ...and this attempt ends here because AFL++ generated a single > > > > event for passt, but there could be more (this should also be > > > > decided by AFL++). > > > > > > > > Would something like this make sense? -- Stefano