Skip to content

Bpftool sync 2026-10-02 - #292

Merged
qmonnet merged 42 commits into
libbpf:mainfrom
qmonnet:bpftool-sync-2026-10-02T14-10-23.033Z
Oct 2, 2026
Merged

qmonnet merged 42 commits into
libbpf:mainfrom
qmonnet:bpftool-sync-2026-10-02T14-10-23.033Z

Conversation

@qmonnet

@qmonnet qmonnet commented Oct 2, 2026

Copy link
Copy Markdown
Member

Pull latest libbpf from mirror and sync bpftool repo with kernel, up to the commits used for libbpf sync. This is an automatic update performed by calling the sync script from this repo:

$ ./scripts/sync-kernel.sh . <path/to/>linux

qmonnet and others added 30 commits October 2, 2026 15:25
We'll need it in a future commit.

Signed-off-by: Quentin Monnet <qmo@kernel.org>
Pull latest libbpf from mirror.
Libbpf version: 1.8.0
Libbpf commit:  6d0328e76bddfbfeb67a1c7217fe9dfaec96fa10

Signed-off-by: Quentin Monnet <qmo@kernel.org>
bpf_fib_lookup() returns the FIB-resolved egress ifindex straight
from the fib result. When the egress is a VLAN device, the returned
ifindex is the VLAN netdev's, which has no XDP xmit handler; XDP
programs that want to forward the frame (e.g. xdp-forward) must
instead target the underlying physical device and push the VLAN tag
themselves. Today the program has no way to learn either the
underlying ifindex or the VLAN tag without maintaining its own
VLAN-to-ifindex map in userspace and refreshing it on netlink
events.

Add BPF_FIB_LOOKUP_VLAN. When the caller sets this flag and the fib
result is a VLAN device whose immediate parent is a real (non-VLAN)
device in the same network namespace, populate the existing output
fields params->h_vlan_proto and params->h_vlan_TCI from the VLAN
device and replace params->ifindex with the parent's ifindex.
params->h_vlan_TCI carries the VID only, with PCP and DEI bits zero; a
consumer wanting to set egress priority writes PCP itself.
params->smac is the VLAN device's own address, which can differ from
the parent's.

Only the immediate parent is resolved, via vlan_dev_priv(dev)->real_dev
and not vlan_dev_real_dev(), which walks to the bottom of a stack. When
the immediate parent is not a real device in the same namespace, the
lookup returns BPF_FIB_LKUP_RET_VLAN_FAILURE and leaves params->ifindex
at the input. This covers a stacked VLAN (QinQ), where the immediate
parent is itself a VLAN device and one h_vlan_proto/h_vlan_TCI pair
cannot describe two tags, and a parent in another network namespace (a
VLAN device can be moved while its parent stays), whose ifindex would
be meaningless in the caller's namespace. A program that wants the
VLAN device's own ifindex re-issues the lookup, with a re-initialized
params, without BPF_FIB_LOOKUP_VLAN, so the unreducible case stays
distinct from a physical egress. That distinction matters for XDP: a
program cannot xmit on a VLAN device, so a success carrying the VLAN
ifindex would make it redirect to a device with no ndo_xdp_xmit and
drop the frame at xdp_do_flush(). The swap and the vlan fields are
written only on the reduce path; other output fields keep their
existing behaviour, so a frag-needed result still reports the route
mtu in params->mtu_result.

BPF_FIB_LOOKUP_VLAN is only useful to XDP, which cannot redirect to a
VLAN device. A tc program can redirect to the VLAN device directly, so
bpf_skb_fib_lookup() rejects the flag with -EINVAL; bpf_xdp_fib_lookup()
accepts it. When the flag is not set, behaviour is unchanged:
h_vlan_proto and h_vlan_TCI are zeroed and ifindex is left at the FIB
result.

The new block is compiled only under CONFIG_VLAN_8021Q since
vlan_dev_priv() is not defined otherwise; without that config
is_vlan_dev() is constant false and the flag is accepted but never
acts. That is safe because no VLAN device can exist there, so every
egress is already physical.

This lets an XDP redirect target the physical device and learn the
tag to push in a single lookup, which xdp-forward's optional VLAN
mode (xdp-project/xdp-tools#504) wants from the kernel side.

The helper's input semantics are unchanged; the reverse direction
(supplying a tag as lookup input) is added in the following patch.

Suggested-by: Toke Høiland-Jørgensen <toke@redhat.com>
Signed-off-by: Avinash Duduskar <avinash.duduskar@gmail.com>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Acked-by: David Ahern <dsahern@kernel.org>
Link: https://lore.kernel.org/bpf/20260713162305.1237211-2-avinash.duduskar@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
BPF_FIB_LOOKUP_VLAN resolves a VLAN egress. The reverse is also
useful: an XDP program receiving a VLAN-tagged frame on a physical
device wants the lookup to behave as if the packet had arrived on the
corresponding VLAN subinterface, so iif-based policy routing and VRF
table selection use the right ingress.

Add BPF_FIB_LOOKUP_VLAN_INPUT. When set, params->h_vlan_proto and
params->h_vlan_TCI are read as an input VLAN tag and the matching VLAN
device of params->ifindex is resolved with __vlan_find_dev_deep_rcu().
The device must be up and in the same network namespace as
params->ifindex (a VLAN device can be moved to another netns while
registered on its parent; receive would deliver into that other
namespace, which a lookup here cannot represent). If params->ifindex
is itself a VLAN device, its inner (QinQ) subinterface is matched.
For a bond or team, a tag on a port matches no device and returns
NOT_FWDED; pass the master's ifindex.
The lookup then runs with the resolved device as the ingress;
params->ifindex itself is not modified on the input side. When the
resolved device is enslaved to a VRF, both the full lookup (via the
l3mdev rule) and BPF_FIB_LOOKUP_DIRECT (via l3mdev_fib_table_rcu())
select the VRF's table from the resolved ingress. That follows from
feeding the resolved device to the flow as the ingress
(fl4.flowi4_iif = dev->ifindex), which is what makes l3mdev resolve
the VRF master from the subinterface rather than from
params->ifindex.

The two failure classes get different treatment on purpose. A
h_vlan_proto other than 802.1Q/802.1ad is API misuse and returns
-EINVAL, since it would otherwise reach the WARN in vlan_proto_idx()
with a program-controlled value. An unmatched VID, a device that is
down, or one in another namespace is a data outcome and returns
BPF_FIB_LKUP_RET_NOT_FWDED, matching the DIRECT path when
fib_get_table() finds no table and mirroring real ingress, where the
receive path drops such frames. A VID of 0 (a priority tag) is looked
up literally and normally fails the same way; receive instead
processes such frames untagged, so callers should not set the flag for
priority tags. Proceeding on the physical device for any of these
would be fail-open for the policy-routing cases above.

The h_vlan fields share a union with tbid, so the flag cannot be
combined with BPF_FIB_LOOKUP_TBID. It describes ingress, so it also
cannot be combined with BPF_FIB_LOOKUP_OUTPUT. Both combinations
return -EINVAL; restricting now keeps a later relaxation backward
compatible. Combining with BPF_FIB_LOOKUP_VLAN is allowed: the tag is
consumed on the ingress side and the egress tag is written on
success.

Under !CONFIG_VLAN_8021Q the __vlan_find_dev_deep_rcu() stub returns
NULL, so every lookup with a valid proto returns NOT_FWDED, which is
correct since no VLAN device can exist.

Suggested-by: Toke Høiland-Jørgensen <toke@redhat.com>
Signed-off-by: Avinash Duduskar <avinash.duduskar@gmail.com>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Link: https://lore.kernel.org/bpf/20260713162305.1237211-3-avinash.duduskar@gmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
BPF_RB_OVERWRITE_POS is supported by bpf_ringbuf_query() but was missing
from the helper documentation. Add it to the flags list in both the
kernel UAPI header and its tools/ mirror.

Signed-off-by: Jianlin Shi <shijianlin11@foxmail.com>
Acked-by: Xu Kuohai <xukuohai@huawei.com>
Link: https://lore.kernel.org/bpf/tencent_22134645443B75ED907D2A85A47AD554A709@qq.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Looking up a prog or map by name walks the whole id space. There is a
window between bpf_prog_get_next_id()/bpf_map_get_next_id() and getting
an fd for that id in which an unrelated object can be freed, and the
lookup then fails with ENOENT and aborts the whole command.

Skip such ids and keep walking, the same way do_show() already does.

Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://lore.kernel.org/bpf/20260720071520.396363-1-jiayuan.chen@linux.dev
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
To pick up the changes in:

  de9e2b3d88af3641 ("uapi: Provide DIV_ROUND_CLOSEST()")

That just rebuilds perf, silencing this build warning.

This addresses this perf build warning:

  Warning: Kernel ABI header differences:
    diff -u tools/include/uapi/linux/const.h include/uapi/linux/const.h

Please see tools/include/uapi/README for further details.

Cc: Cristian Ciocaltea <cristian.ciocaltea@collabora.com>
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
At the moment there are more callsites that want bpf_verbose_insn() to
not print a newline after the instruction, than callsites that want a
newline. Drop '\n' from disasm.c. Non-functional change.

The changes in bpftool are verified by writing a bpf program using a
variety of instructions and comparing `prog dump xlated` output in the
following modes: plain, opcodes, visual, visual opcodes. The output
before and after the changes is identical.

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Quentin Monnet <qmo@kernel.org>
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20260807-static-zext-v4-1-b6c270013c77@gmail.com
Enhance bpftool to generate skeletons that properly handle global percpu
variables. The generated skeleton now includes a dedicated structure for
percpu data, allowing users to initialize and access percpu variables more
efficiently.

For global percpu variables, the skeleton now includes a nested
structure, e.g.:

struct test_global_percpu_data {
	struct bpf_object_skeleton *skeleton;
	struct bpf_object *obj;
	struct {
		struct bpf_map *percpu;
	} maps;
	// ...
	struct test_global_percpu_data__percpu {
		int data;
		char run;
		struct {
			char set;
			int i;
			int nums[7];
		} struct_data;
		int nums[7];
	} *percpu;

	// ...
};

  * The "struct test_global_percpu_data__percpu *percpu" points to
    initialized data, which is actually "maps.percpu->mmaped".
  * Before loading the skeleton, updating the
    "struct test_global_percpu_data__percpu *percpu" modifies the initial
    value of the corresponding global percpu variables.
  * After loading the skeleton, "maps.percpu->mmaped" has been marked as
    read-only in libbpf. If users want to update the global percpu
    variables, they have to update the "maps.percpu" map instead.
  * For lightweight skeleton, "lskel->percpu" will be protected by
    "mprotect(p, sz, PROT_READ)".
  * For subskeleton, those variables of global percpu data will be
    skipped.

Assisted-by: Codex:gpt-5.5-xhigh
Signed-off-by: Leon Hwang <leon.hwang@linux.dev>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Acked-by: Quentin Monnet <qmo@kernel.org>
Link: https://lore.kernel.org/bpf/20260813152324.97937-7-leon.hwang@linux.dev
map_dump() closes the map fd in its error path, and do_dump() then
closes the same fd again after a successful dump. Closing an already
closed fd leaves errno set to EBADF, which poisons later errno checks
such as the batch file read check in do_batch(). Let do_dump() own the
fd and remove the close from map_dump().

The same double-close pattern exists in do_show_subset(): both
show_map_close_json() and show_map_close_plain() already close the fd,
so drop the extra close() there as well.

Also propagate the error when bpf_map_get_info_by_fd() fails on a
subsequent map in do_dump(): set err = -1 before breaking out of the
loop, so a later failure is not silently hidden after an earlier
iteration succeeded.

Fixes: 99f9863a0c45f ("bpftool: Match maps by name")
Signed-off-by: Yuan Chen <chenyuan@kylinos.cn>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20260810142224.2907373-2-chenyuan_fl@163.com
The existing anonymous enum for BPF_FUNC_skb_adjust_room flags is
named to enum bpf_adj_room_flags to enable CO-RE (Compile Once -
Run Everywhere) lookups in BPF programs.

Co-developed-by: Max Tottenham <mtottenh@akamai.com>
Co-developed-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Max Tottenham <mtottenh@akamai.com>
Signed-off-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Nick Hudson <nhudson@akamai.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://lore.kernel.org/bpf/20260812083115.73100-2-nhudson@akamai.com
Add new bpf_skb_adjust_room() decapsulation flags:

- BPF_F_ADJ_ROOM_DECAP_L4_GRE
- BPF_F_ADJ_ROOM_DECAP_L4_UDP
- BPF_F_ADJ_ROOM_DECAP_IPXIP4
- BPF_F_ADJ_ROOM_DECAP_IPXIP6

These flags let BPF programs describe which tunnel layer is being
removed, so later changes can update tunnel-related GSO state
accordingly during decapsulation.

This patch only introduces the UAPI flag definitions and helper
documentation.

Co-developed-by: Max Tottenham <mtottenh@akamai.com>
Co-developed-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Max Tottenham <mtottenh@akamai.com>
Signed-off-by: Anna Glasgall <aglasgal@akamai.com>
Signed-off-by: Nick Hudson <nhudson@akamai.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://lore.kernel.org/bpf/20260812083115.73100-4-nhudson@akamai.com
do_batch() checks errno after the read loop to detect read failures,
but fgets() does not clear errno on success, so a stale errno left by
a previously executed command (e.g. map dump's EBADF from a double
close) makes bpftool report a batch file read failure and exit with an
error even though every command succeeded.

Clear errno before each fgets() call, so the post-loop check only
sees the outcome of the last read: zero on success or EOF, E2BIG for
an overlong line, and a genuine errno when fgets() fails.

Since errno is now reset before every read in batch mode, drop the
USE_LIBCAP errno reset in main() that existed only to keep errno clean
for the batch mode.

Fixes: 71bb428fe2c1 ("tools: bpf: add bpftool")
Signed-off-by: Yuan Chen <chenyuan@kylinos.cn>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20260824092657.1789956-2-chenyuan_fl@163.com
do_batch() strips trailing comments by truncating the line at '#'
before checking whether fgets() filled the buffer. If a batch line
longer than the buffer contains a '#' within the first
sizeof(buf) - 1 bytes, the truncation makes strlen(buf) smaller and the
line-length check is bypassed. The unread remainder of the line then
stays in the file stream and is parsed and executed as a separate
command on the next loop iteration.

Continuation lines handled below are affected the same way: an overlong
continuation line containing '#' bypasses the "command is too long"
check, and its unread remainder is executed as a separate command.

Move the line-length checks before the comment is stripped, so they see
the full line as read from the file and overlong lines are rejected
regardless of comments. A line that fills the buffer exactly is now
rejected as well, which is fine: batch command lines are not expected
to come anywhere near the buffer limit.

Fixes: 71bb428fe2c1 ("tools: bpf: add bpftool")
Signed-off-by: Yuan Chen <chenyuan@kylinos.cn>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20260824092657.1789956-3-chenyuan_fl@163.com
Add bpftool support for ML-DSA program signing and drop the flag for
ML-DSA keys on affected OpenSSL versions, the same way as commit
0ad9a71933e7 ("modsign: Enable ML-DSA module signing").

Also, an ML-DSA-87 signature is 4627 bytes on its own, so the blob does
not fit into the 4 KiB of MAX_SIG_SIZE anymore and signing would fail
otherwise. Bump to 16 KiB. MAX_SIG_SIZE only sizes the buffer for
what bpftool itself emits (unrelated to BPF_PROG_MAX_SIGNATURE_SIZE).

Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/r/20260828175227.1537793-5-daniel@iogearbox.net
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
The C dump sorts types by default, so that generated headers are
diffable. The sorted dump emits one type fewer than the unsorted dump
of the same BTF.

dump_btf_c() starts its loop at index 1 to skip the void type at BTF
type ID 0. That holds for the unsorted dump, where the array index is
the type ID, but not after qsort(): position 0 is then the lowest
ranked type, and btf_type_rank() ranks an anonymous enum 0 while void
takes the default rank of 10. So the enum is skipped, and void is
emitted instead as a no-op.

Fixes: 94133cf24bb3 ("bpftool: Introduce btf c dump sorting")
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Ihor Solodrai <ihor.solodrai@linux.dev>
Link: https://lore.kernel.org/r/20260828215207.3105313-7-ihor.solodrai@linux.dev
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
When setting the XDP hints ifname, bpftool would set the
BPF_F_XDP_DEV_BOUND_ONLY without looking at the existing program flags,
overriding any other flag values. This was always a destructive action,
but after we change libbpf to carry the frags section flag in
prog_flags, this can impact bpftool loading of XDP frags programs.

Change the flag setting to be non-destructive by OR'ing it with the
existing flags.

Fixes: f46392ee3dec ("bpftool: Specify XDP Hints ifname when loading program")
Signed-off-by: Toke Høiland-Jørgensen <toke@redhat.com>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Reviewed-by: Larysa Zaremba <larysa.zaremba@intel.com>
Link: https://lore.kernel.org/bpf/20260901-libbpf-frags-flags-v3-1-4eb6f14968b0@redhat.com
The signed-load mnemonic table has entries for byte, half-word, and word
loads because BPF_MEMSX does not support double-word loads. A BPF_MEMSX
| BPF_DW instruction nevertheless selects index 3, past the end of this
table.

Program Structure diagnostics can disassemble a malformed instruction
before check_and_resolve_insns() rejects its opcode. Placing the invalid
signed double-word load at the end of a program therefore triggers an
out-of-bounds access while reporting subprogram fallthrough.

Treat signed double-word loads as invalid in the disassembler and use
the existing BUG_ldx fallback instead.

Fixes: a8f427835394 ("bpf: Report Program Structure CFG errors")
Reported-by: syzbot+3544d9b2a9206be8ba37@syzkaller.appspotmail.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://lore.kernel.org/bpf/20260820022020.3450479-2-memxor@gmail.com
bpf_convert_ctx_accesses() rewrites an atomic on an arena pointer from
BPF_STX | BPF_ATOMIC to BPF_STX | BPF_PROBE_ATOMIC, this patch adjusts
print_bpf_insn() to print such instructions as regular atomics with a
'probe_' prefix (instead of printing them as BUG_XX).

Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://lore.kernel.org/r/20260903171542.1438050-2-eddyz87@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Adopt the newly added bpf_program__add_flags() helper to
non-destructively add the BPF_F_XDP_DEV_BOUND_ONLY flag to XDP programs.

Signed-off-by: Toke Høiland-Jørgensen <toke@redhat.com>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Acked-by: Quentin Monnet <qmo@kernel.org>
Acked-by: Ihor Solodrai <ihor.solodrai@linux.dev>
Link: https://lore.kernel.org/bpf/20260912084109.432834-4-toke@redhat.com
A perf counter can be zero at fentry when PMU multiplexing has
not scheduled the event. Do not reject that sample.

Use enabled time to distinguish a successful snapshot. Read
directly into the per-CPU map value.

Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20260916-bpftool_cyles_per_run-v4-1-f2de02c480d5@meta.com
Independently scheduled perf events can cover different intervals,
making derived ratios inconsistent.

Open one event group per CPU and enable it only after all selected
metrics have joined.

Fail the profile setup when a selected event cannot join its group.

Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20260916-bpftool_cyles_per_run-v4-2-f2de02c480d5@meta.com
Perf counters undercount when the PMU multiplexes them. Scale each
per-CPU value before aggregation. Per-CPU event groups make PMU ratios
use matching scheduling intervals.

Keep run_cnt as the total number of program executions, independent of
counter scheduling.

Keep JSON value raw, report the scaled value in value_scaled, and
document the output.

Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20260916-bpftool_cyles_per_run-v4-3-f2de02c480d5@meta.com
This patch adds necessary infrastructure to attach a struct_ops
map to a cgroup. The initial need was to support migrating
the legacy BPF_PROG_TYPE_SOCK_OPS to a struct_ops.
Recently, there are other struct_ops use cases that
need to attach struct_ops to a cgroup. For example,
the recent BPF OOM and memcg discussion in LSFMMBPF 2026.

The motivation is to create a consistent expectation
for attaching struct_ops to cgroup instead of each subsystem
creating its own infrastructure. This logic includes
hierarchy expectation, ordering expectation,
attachment API, and rcu gp.

There is already an existing implementation for attaching
multiple bpf progs to a cgroup. There are also tools
built around it for querying. Attaching a struct_ops map
(which is a group of bpf programs) could also adhere to
a similar API and potentially reuse most of the existing
implementation.

A couple of ideas have been tried. One of them
is to use mprog.c. In terms of the amount of changes,
I eventually came to the same conclusion as in
commit 120933984460 ("bpf: Implement mprog API on top of existing cgroup progs").
I then shifted the focus to reusing the current
{update,compute,activate,purge}_effective_progs() which has
the main logic that implements the mprog API.

Since then, I tried to add a 'struct cgroup *cgroup' member
to the existing 'struct bpf_struct_ops_link' and link_create
will create a 'struct bpf_struct_ops_link' object to be stored
in the pl->link. This turns out to have more changes on
both cgroup.c and bpf_struct_ops.c than I like.

This patch directly reuses the 'struct bpf_cgroup_link' which
cgroup.c already understands. Add 'struct bpf_map *map'
to 'struct bpf_cgroup_link'. In the future, as more subsystems
are extended by struct_ops, we may consider to make
'struct bpf_map *map' as a primary citizen of a link
like 'struct bpf_prog *prog' and directly add
'struct bpf_map *map' to the generic 'struct bpf_link'.

The pl->link could be the traditional 'prog' link or the
new 'map' link. The places that need to handle them differently
have already been refactored into the new prog_list_*() added in
the earlier patch. In those new prog_list_*(), this patch will
check "pl->link && pl->link->map", learn that it is a 'map' link
and handle it correctly.

The bpf_prog_array also needs to handle that its item can store
the traditional 'prog' or it can store a struct_ops map.
The places that need to handle them differently have also
been refactored into the new bpf_cgroup_array_*() added
in the earlier patch. The two differences are:
  - different sentinel (dummy_bpf_prog in prog vs cfi_stub in struct_ops)
  - the array for struct_ops may need to go through different
    rcu gp.
The bpf_cgroup_array_*() functions use the cgroup_bpf_attach_type (ie atype)
to distinguish the array is storing prog or storing struct_ops map.

This patch also implements a separate struct bpf_link_ops
"cgroup_struct_ops_link_ops" to have a separate link_ops implementation
that only handles the cgroup's struct_ops link.

Questions:
- Although this patch did not change it, it is not obvious to me how
  the replace_effective_progs() and purge_effective_progs() handle
  cases when there are existing BPF_F_PREORDER progs attached
  in the hlist.

Misc notes:
- CGROUP_TCP_SOCK_OPS is added to the 'enum cgroup_bpf_attach_type'.
  The actual implementation of the tcp_bpf_ops (a struct_ops)
  will be added in the next patch.

- free_after_mult_rcu_gp is added to 'struct bpf_struct_ops' such that
  the bpf_prog_array can have a mix of sleepable and
  non-sleepable prog in a struct_ops. This can tell
  how the bpf_prog_array should be freed.

- For a struct_ops that supports cgroup attachment, it does not need to
  implement its own reg/unreg function. reg/unreg to a cgroup is
  done by the common infrastructure added in this patch.

- The cgroup's struct_ops link only supports BPF_F_ALLOW_MULTI.
  This is enforced internally in cgroup_bpf_struct_ops_attach.
  This should be consistent with the current prog's link
  behavior in cgroup_bpf_link_attach.

  In the future, we may allow each subsystem to choose differently.

- A cgroup_atype member is added to 'struct bpf_struct_ops'.
  When a subsystem struct_ops needs to support cgroup attachment,
  it needs to add a value to 'enum cgroup_bpf_attach_type'
  and then assign it to the newly added cgroup_atype member
  in the bpf_struct_ops.

- During LINK_CREATE in syscall, the patch uses the same
  BPF_STRUCT_OPS (in attr->link_create.attach_type).
  The bpf_struct_ops_link_create learns the map and
  from the map it learns the st_ops. If the st_ops->cgroup_atype
  is not 0, it will create a cgroup's link.

- When a subsystem registers a struct_ops that supports cgroup
  attachment, the struct_ops infrastructure will also ask the
  cgroup infrastructure to remember a few things. This is done
  by calling cgroup_bpf_struct_ops_register().

Signed-off-by: Martin KaFai Lau <martin.lau@kernel.org>
Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260917200542.3689605-10-ameryhung@gmail.com
Add the TCP header option callbacks to the bpf_tcp_ops struct_ops type:

  parse_hdr     - parse the options of an incoming skb on an established
                  connection
  hdr_opt_len   - reserve space in the TCP header for bpf options
  write_hdr_opt - write the reserved bpf options

These mirror the BPF_SOCK_OPS_PARSE_HDR_OPT_CB, _HDR_OPT_LEN_CB and
_WRITE_HDR_OPT_CB legacy sockops callbacks, but are exposed as struct_ops
members so a program can implement them with normal function signatures
and per-member helper sets.

The reserved header window is shared between the legacy sockops and
bpf_tcp_ops paths. tcp_{syn,synack,established}_options() first run the
legacy BPF_SOCK_OPS_HDR_OPT_LEN_CB and then call hdr_opt_len, so both
sources accumulate into opts->bpf_opt_len; at write time the legacy
options are emitted first and bpf_tcp_ops writes after them.

API design

bpf_tcp_ops overloads the sock_ops header-option helpers rather than
introducing a new API: bpf_reserve_hdr_opt(), bpf_store_hdr_opt() and
bpf_load_hdr_opt() are exposed per-member (reserve for hdr_opt_len,
store/load for write_hdr_opt, load for parse_hdr) and share the existing
kernel option-walking core via _bpf_sock_ops{store,load}hdr_opt(), with
the bpf_tcp_ops wrappers synthesizing a temporary bpf_sock_ops_kern from
the program ctx. This keeps a port from the legacy
BPF_SOCK_OPS*_HDR_OPT_CB callbacks mechanical (same helper calls) and
adds no new UAPI helper/kfunc surface.

An alternative considered was to drop the option helpers entirely: have
hdr_opt_len reserve space purely through its return value, and introduce
a dedicated TCP-header-option dynptr used for both reading and writing.
That is a cleaner, more self-contained interface, but it is a larger
change and does not reuse the legacy helpers, making a port from sockops
less mechanical. It can be pursued as a follow-up; the helper-based
interface here keeps this series focused on moving the hooks to
struct_ops.

The hdr_opt_len fast path in tcp_established_options() is gated by
cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS). Note this is a global,
per-attach-type static branch: it is enabled whenever any bpf_tcp_ops is
attached, even one that does not implement hdr_opt_len or that is attached
to a different cgroup. In those cases the block still runs but
bpf_tcp_ops_hdr_opt_len() no-ops via the per-member check in the dispatch
macro. A per-member/per-cgroup gate could be added later if the extra
fast-path work proves measurable.

Signed-off-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Link: https://patch.msgid.link/20260917200542.3689605-13-ameryhung@gmail.com
bpftool rounds memory-mapped data map sizes to the host page size when
generating a light skeleton. The generated code therefore uses a 64K
mapping size when bpftool runs on a 64K-page host, even if the skeleton
runs on a 4K-page target. The target rejects the oversized map mmap(),
causing failure of loading the light skeleton.

When try to run 64K-page selftests on 4K-page VM, the error message does
not provide the reason about page size.

 test_atomics:PASS:atomics skeleton open 0 nsec
 test_atomics:FAIL:atomics skeleton load unexpected error: -12 (errno 22)
 libbpf#15      atomics:FAIL

Pass the original map value size and max entries to the generated code and
round mmap size to the runtime page size in the user-space light
skeleton helpers. This keeps generated light skeletons independent of the
build host page size.

Fixes: d510296d331a ("bpftool: Use syscall/loader program in "prog load" and "gen skeleton" command.")
Signed-off-by: Leon Hwang <leon.hwang@linux.dev>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Acked-by: Quentin Monnet <qmo@kernel.org>
Link: https://lore.kernel.org/bpf/20260915161450.96249-1-leon.hwang@linux.dev
map event_pipe only accepts perf event arrays, leaving no built-in way
to inspect records produced through the BPF ring buffer API. Extend it
to consume queued and live BPF_MAP_TYPE_RINGBUF records in plain or
JSON output, retaining the existing perf event array path.

Ring buffers have a shared consumer position and no implicit CPU or
timestamp. Document that this command consumes records rather than
observing them passively, and reject perf-only CPU/index selectors.

A producer can keep ring_buffer__poll() busy after a stop signal, so
return -EINTR from the record callback when stopping. Keep stdio out of
the shared signal handler and preserve callback output errors.

Signed-off-by: Tianyi Chen <hi@tychen.cc>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Tested-by: Quentin Monnet <qmo@kernel.org>
Link: libbpf#54
Link: https://lore.kernel.org/bpf/20260921073556.99421-2-hi@tychen.cc
BPF programs that manage their own objects have no way to run their own
logic once an RCU grace period has elapsed.  bpf_obj_drop() defers a
free, but returning an index to an allocator or unpinning a resource
once readers are done has no equivalent.  sched_ext's BPF library works
around this today by pushing freed nodes onto a list and having a
userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
BPF program to reclaim them.

Add:

	int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
			 int (*callback)(struct bpf_map *map, void *key,
					 void *value));

@rh is a struct bpf_rcu_head embedded in a value of @Map, so the
callback runs as callback(map, key, value) for the element it lives in
and needs no cookie.  A head can only be armed once, which bounds
outstanding work by the number of elements.

struct bpf_rcu_head holds the callback state inline rather than a
pointer to it, as bpf_timer, bpf_wq and bpf_task_work do, because there
is nothing to cancel and so nothing that has to outlive the map value.
That avoids an allocation and a state machine on the arming path at the
cost of 48 bytes per element.

An RCU callback cannot be cancelled, so everything it touches has to
stay alive until it runs:

  - The callback is the program's text, so arming takes a program
    reference as bpf_timer, bpf_wq and bpf_task_work do, dropped once
    the callback returns.  bpf_prog_inc_not_zero() also fails the arm
    with -EBADF once the program is dying.

  - The map is held by that reference through used_maps.  An inner map
    is not, so bpf_rcu_head is rejected in one.

  - The field is only accepted in BPF_MAP_TYPE_ARRAY, whose elements
    are never freed individually.  A hash element can be deleted and
    recycled while a callback is queued on it.

  - The head is disarmed before the callback runs so it can be armed
    again from there, which takes a new program reference before the
    running callback drops its own.  Arming therefore fails with -EPERM
    once the map is held by neither a process nor bpffs.

bpf_iter hands a program a writable pointer to the live element, which
would let it overwrite a queued head, so bpf_iter_attach_map() rejects
maps carrying one.

The callback is verified non-sleepable even when the caller is
sleepable, and RCU invokes it with BH disabled.

Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260922200208.3203834-2-puranjay@kernel.org
Introduce BPF_JMP | BPF_CALL | BPF_X (opcode 0x8d) 'callx dst_reg'
instruction: indirect call of bpf subprog with address in dst_reg.
That's the encoding LLVM emits for calls via function pointer.
src_reg, off, imm are reserved and must be zero.

dst_reg must be PTR_TO_FUNC produced by ld_imm64 BPF_PSEUDO_FUNC.
check_ld_imm() allows it for static subprogs only, so callx cannot call
global subprogs or the main prog. Since every callee has its address
taken by ld_imm64, add_subprogs() and check_cfg() see all of them before
the main pass, and might_sleep, changes_pkt_data, might_throw of
the callee are already merged into the subprog that takes the address.

reg->subprogno is the callee. Verify callx as a direct call of that
static subprog: split check_func_call() into check_static_func_call()
that is shared with new check_func_callx(). Different paths through
the same callx may call different subprogs.

Arithmetic on PTR_TO_FUNC is allowed, so check that the pointer wasn't
modified. Allow callx while holding a lock like direct calls of static
subprogs.

The interpreter doesn't support callx. Set jit_required and add
bpf_jit_supports_callx() for JITs to opt in. No JIT does yet, so callx
is still rejected.

Print it as "callx rN" in the verifier log and xlated dump.

Adjust "invalid call insn1" test_verifier test that used opcode 0x8d as
unknown opcode. It fails with "R0 !read_ok" now.

Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260924031042.1690890-6-alexei.starovoitov@gmail.com
bpftool uses dense per-CPU value-buffer slots when printing per-CPU map
values. It also uses the slot index as the CPU ID, which produces
incorrect labels when the possible CPU mask is sparse, such as 0,2-3.

Parse the possible CPU mask and use the corresponding logical CPU ID in
plain, JSON, and BTF-formatted output. Keep the dense slot index for
accessing the per-CPU value buffer.

Pass the parsed CPU count together with the CPU ID array so that the
printed CPU IDs and loop bounds come from the same possible-CPU mask.
Propagate CPU-ID lookup and map output errors through the shared output
path to its callers.

Fixes: 71bb428fe2c1 ("tools: bpf: add bpftool")
Signed-off-by: Hui Su <sh_def@163.com>
Link: https://lore.kernel.org/bpf/20260813155131.1022745-3-sh_def@163.com/
Link: https://lore.kernel.org/bpf/20260923004243.969919-2-sh_def@163.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
liulangrenaaa and others added 12 commits October 2, 2026 15:25
bpftool prog profile currently treats the number of possible CPUs as
both the logical CPU ID range and the stride of the perf event array.
That misses valid logical CPUs when the possible CPU mask is sparse,
such as 0,2-3, and can use incorrect PERF_EVENT_ARRAY keys.

Keep the compact possible CPU count for per-CPU result buffers, while
enumerating the actual logical CPU IDs when creating per-CPU perf event
groups. Use the maximum logical CPU ID plus one as the metric stride in
the PERF_EVENT_ARRAY.

This keeps the current per-CPU event grouping intact while separating
the compact per-CPU buffer index from the logical CPU ID and event-array
key. Report the logical CPU ID when a per-CPU event was not counted.

Fixes: 47c09d6a9f67 ("bpftool: Introduce "prog profile" command")
Signed-off-by: Hui Su <sh_def@163.com>
Link: https://lore.kernel.org/bpf/20260813160858.1042834-3-sh_def@163.com/
Link: https://lore.kernel.org/bpf/20260923004243.969919-3-sh_def@163.com
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Add BTF_KIND_LOC_PARAM, BTF_KIND_LOC_PROTO and BTF_KIND_LOCSEC
to help represent location information for functions.

BTF_KIND_LOC_PARAM is used to represent how we retrieve data at a
location; either via register(s), or register+offset, a dereference
of a register+offset or a constant value.

BTF_KIND_LOC_PROTO represents location information about a location
with multiple BTF_KIND_LOC_PARAMs.

And finally BTF_KIND_LOCSEC is a set of location sites, each
of which has

- a BTF_KIND_FUNC function associated with the inline site
- a location prototype specifying where to find the function
  parameters
- an address offset relative to the kernel base address

This can be used to support representing

- a fully-inlined function at potentially multiple inline sites
  with potentially different parameter availability
- a partially-inlined function where some _LOC_PROTOs represent
  inlined sites as above and others have normal _FUNC representations

Also BTF_KIND_LOCSEC struct btf_loc will have two type id
references; one for the associated func, the other for the loc_proto.
Accordingly increase the number of m_offs references in btf_field_desc
to 2.

Signed-off-by: Alan Maguire <alan.maguire@oracle.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://patch.msgid.link/20260924111428.75957-2-alan.maguire@oracle.com
For bpftool to be able to dump .BTF.inline data in
/sys/kernel/btf/foo.inline for module foo, it needs to support
multi-split BTF because the parent-child relationship of BTF
inline data for modules is

vmlinux BTF data
	module BTF data
		module BTF inline data

So for example to dump BTF inline info for xfs we would run

$ bpftool btf dump -B /sys/kernel/btf/vmlinux -B /sys/kernel/btf/xfs file /sys/kernel/btf/xfs.inline

Multiple bases are specified with the vmlinux base BTF first (parent)
followed by the xfs BTF (child), and finally the XFS BTF inline info.

Update help text accordingly to reflect the ability to specify multiple
in-order root-to-branch bases.

Signed-off-by: Alan Maguire <alan.maguire@oracle.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924111428.75957-8-alan.maguire@oracle.com
Document the ability to pass multiple levels of split BTF, using
"-B base-btf" options.

Signed-off-by: Alan Maguire <alan.maguire@oracle.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924111428.75957-9-alan.maguire@oracle.com
In raw mode ensure we can dump new BTF kinds in normal/json format.
BTF_KIND_LOC_PARAMs are rendered as strings, for example a
const value of 0x2a and a dereference of %r10 + 0x20:

  [12] LOC_PARAM '(anon)' size=4 flags=0x2 vlen=1 values='0x2a'
  [13] LOC_PARAM '(anon)' size=8 flags=0x38 vlen=2 values='*(reg10 + 0x20)'

LOC_PROTOs render the associated values of each of their
LOC_PARAMs for easier readability:

  [14] LOC_PROTO '(anon)' vlen=2
  	type_id=12 value='0x2a'
  	type_id=13 value='*(reg10 + 0x20)'

and LOCSEC shows function name associated with site:

  [15] LOCSEC 'inline.text' vlen=1
  	name='foo' func_type_id=5 loc_proto_type_id=14 offset=64

Registers are displayed in an architecture-neutral form;
'regN' where N is the DWARF register number, or 'fbreg'
for the stack frame base.

Signed-off-by: Alan Maguire <alan.maguire@oracle.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260924111428.75957-10-alan.maguire@oracle.com
Commit 4b82b181a26c ("bpf: Allow pre-ordering for bpf cgroup progs")
introduced BPF_F_PREORDER to request pre-order execution across the
cgroup hierarchy. Furthermore, attachments legitimately use combinations
such as BPF_F_ALLOW_MULTI | BPF_F_PREORDER or
BPF_F_ALLOW_OVERRIDE | BPF_F_PREORDER.

With BPF_PROG_QUERY reporting the per-program BPF_F_PREORDER attribute,
bpftool's exact-match formatter falls back to "unknown(40)" when
BPF_F_PREORDER is present alone, or "unknown(41)" / "unknown(42)" when
combined with BPF_F_ALLOW_OVERRIDE or BPF_F_ALLOW_MULTI. Additionally,
do_attach() only accepts "multi" and "override", rejecting "preorder"
with "unknown option".

Before:
  $ bpftool cgroup show <cg>
  1234  cgroup_inet_ingress  unknown(42)  test_prog
  $ bpftool cgroup attach <cg> cgroup_inet_ingress id 5678 multi preorder
  Error: unknown option: preorder

After:
  $ bpftool cgroup show <cg>
  1234  cgroup_inet_ingress  multi,preorder  test_prog
  $ bpftool cgroup attach <cg> cgroup_inet_ingress id 5678 multi preorder
  (attaches successfully)

Refactor the attach flags formatter into a bitmask formatter that outputs
comma-separated flag names for plain text while preserving unrecognized
bits as "unknown(...)". Render attach flags as an array of flag names in
JSON output. Accept "preorder" in do_attach(), update the cgroup
documentation and synopsis to express valid flag combinations, and teach
bash completion about them ("multi" or "override" optionally combined
with "preorder").

Signed-off-by: Hui Su <sh_def@163.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Quentin Monnet <qmo@kernel.org>
Link: https://patch.msgid.link/20260922025442.3176057-4-sh_def@163.com
The existing BPF_PROG_STREAM_READ_BY_FD command only supports polling a
program stream through repeated bpf() calls. It cannot block for new data
or integrate with poll-based event loops.

Add BPF_PROG_STREAM_OPEN to return a read-only, close-on-exec file
descriptor for a selected program stream. Reads block by default and
BPF_F_STREAM_NONBLOCK, the only accepted flag, provides non-blocking
behavior. poll reports readable data and reports hangup once the program
has been freed. Like pipes and sockets, the descriptor is not seekable and
lseek fails with ESPIPE.

A stream descriptor deliberately does not retain the program. Move each
stream into a separately refcounted allocation so program teardown can mark
it dead and wake descriptor users while outstanding descriptors drain
buffered data safely. Readers sample the dead flag before looking for data,
so EOF is reported only when the stream was already dead before it was
found empty; data published right before teardown is never skipped.

Only programs loaded through BPF_PROG_LOAD get streams. Classic BPF
filters, JIT subprograms and shim programs never write to one, and
kernel-side writers already resolve a subprogram to its main program, so
those programs no longer carry stream state.

Readiness needs its own counter. Stream capacity is charged before
allocation and before an element is published to the stream log, so using
that reservation as the read and poll condition can report readable data
while no element exists: a blocking reader retries instead of sleeping and
a lone non-blocking reader can see POLLIN followed by EAGAIN. Publish bytes
with release ordering after adding elements to the lockless log, use
acquire loads before consuming them or reporting readiness, limit each read
to its readable snapshot and subtract only bytes actually copied. This
keeps the aggregate count correct even when concurrent publishers update it
out of publication order. With several readers on one stream, readiness
remains advisory, as it is for pipes. The capacity counter is kept solely
for enforcing the stream size limit.

Wakeups are always deferred through irq_work. Stream writers run in
whatever context the program runs in: NMI context for perf_event programs,
sections with interrupts disabled inside bpf_spin_lock or rqspinlock
critical sections since bpf_stream_vprintk() is KF_SPINLOCK_SAFE, and
tracing programs attached anywhere in the kernel, including inside the wait
queue and epoll code itself. Waking waiters directly from there can
deadlock, and no cheap context check covers every case: on PREEMPT_RT,
spinlock_t sections do not disable interrupts, so in_nmi() or
irqs_disabled() cannot tell such a program apart from a benign one. Queue
an irq_work item instead, as bpf_ringbuf does.

Queue it only when a publication turns an empty stream readable. Readers
block and pollers wait only after finding the stream empty, and the
readable count never drops below zero because each read is bounded by its
snapshot, so the first publication after such an observation is the one
that makes the count positive, and it is the one that queues the wakeup.
Publications into a stream that already holds data raise no interrupt, so
a program that prints while nobody drains its stream pays for a single
irq_work until the stream is emptied again. This matches bpf_ringbuf, which
notifies only once the consumer has caught up. Blocking readers and
level-triggered pollers re-check the readable count before waiting, so they
cannot miss data, and edge-triggered epoll consumers drain until EAGAIN
before waiting again, as epoll(7) requires.

Synchronize pending work before releasing the final stream reference so
the callback cannot outlive the stream, but only when the work was ever
queued: irq_work_sync() waits for an RCU grace period on PREEMPT_RT and on
architectures without an irq_work interrupt.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://patch.msgid.link/20260925045536.1480933-3-memxor@gmail.com
bpftool prog tracelog { stdout | stderr } PROG dumps the output buffered
in a program stream and exits, which is all BPF_PROG_STREAM_READ_BY_FD
allows. Waiting for further output means rerunning the command.

Add a -w/--wait option that opens the stream with bpf_prog_stream_open()
in its default blocking mode and keeps printing output as the program
produces it. Flush each chunk as it arrives so redirected output is not
held in stdio buffers while the next read blocks. Waiting is opt-in: the
default dump keeps using BPF_PROG_STREAM_READ_BY_FD and exits once the
buffered output is drained, so existing scripts behave the same on old and
new kernels.

A stream descriptor does not keep its program alive, so bpftool drops the
program descriptor once the stream is open. Waiting then ends with EOF
when the program is unloaded and otherwise runs until interrupted, as
bpftool prog tracelog already does for the trace pipe. Exit from the
SIGINT, SIGHUP and SIGTERM handlers like that command does. A flag set by
the handler and checked before each read would miss a signal that lands
between the check and the blocking read(), leaving bpftool asleep until the
next print.

Waiting requires BPF_PROG_STREAM_OPEN. The bpf() syscall fails with EINVAL
for an unknown command, which bpftool reports as missing kernel support
rather than degrading into a dump. Document the option and offer it in the
bash completion.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Quentin Monnet <qmo@kernel.org>
Link: https://patch.msgid.link/20260925045536.1480933-5-memxor@gmail.com
Augment func= output for LOCSEC entries to include a mapping from
function signature to where parameters are stored; for example:

[290179] LOCSEC 'inline.text' vlen=524941
        func='task_pid_nr(tsk [reg0])' func_type_id=136691 loc_proto_type_id=136693 offset=2097226
        func='get_current()' func_type_id=136694 loc_proto_type_id=136695 offset=2097247
        func='arch_static_branch(key [address 0x2e275e8], branch [const 0x0])' func_type_id=136697 loc_proto_type_id=136700 offset=2097296

Fixes: 321562c34d5b ("bpftool: Add ability to dump LOC_PARAM, LOC_PROTO and LOCSEC")
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Alan Maguire <alan.maguire@oracle.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Quentin Monnet <qmo@kernel.org>
Link: https://patch.msgid.link/20260926175113.2368566-3-alan.maguire@oracle.com
bpftool prog tracelog -w ends from its SIGINT, SIGHUP and SIGTERM handler,
which calls exit(). exit() is not async-signal-safe: it runs atexit
handlers and flushes stdio streams, and the signal may land while the read
loop is inside fwrite() or fflush() on the stream being flushed. Whether
that deadlocks or flushes a half-updated buffer depends on the C library.

Call _exit() instead. The loop flushes every chunk as soon as it is
printed, so the stdio buffer is empty except while a chunk is being
written, and skipping the exit-time flush loses nothing.

While here, note in the manual page that opening a stream as a file
descriptor arrives in Linux 7.4, and restrict the -w/--wait description to
the stdout/stderr form of the tracelog command, since the trace pipe form
does not take the option.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Quentin Monnet <qmo@kernel.org>
Link: https://lore.kernel.org/bpf/d64c6402-175f-4ea6-be6d-1a544a95c97c@qmon.net
Link: https://lore.kernel.org/bpf/20260926154515.191689-1-memxor@gmail.com
Update .mailmap based on bpftool's list of contributors and on the
latest .mailmap version in the upstream repository.

Signed-off-by: Quentin Monnet <qmo@kernel.org>
Syncing latest bpftool commits from kernel repository.
Baseline bpf-next commit:   9a3a07d06e7d74f4aecc51396c771149336ac55d
Checkpoint bpf-next commit: b5a4aa31abd6fe90009b63e35dc18c67d041ec0c
Baseline bpf commit:        7cbd0c4cebe4c9f678d15e6b9ba975e1155a107f
Checkpoint bpf commit:      de020dc8049bfb2b22e3b6d99c031feb2e22d112

Alan Maguire (5):
  btf: Extend UAPI to support BTF location (inline site) info
  bpftool: Handle multi-split BTF by supporting multiple base BTFs
  bpftool: Document support for multi-split BTF
  bpftool: Add ability to dump LOC_PARAM, LOC_PROTO and LOCSEC
  bpftool: Update func representation to include function signature

Alexei Starovoitov (1):
  bpf: Add callx instruction to call bpf subprogs indirectly

Amery Hung (1):
  bpf: tcp: Support parse/len/write header option hooks in bpf_tcp_ops

Arnaldo Carvalho de Melo (1):
  tools headers UAPI: Sync linux/const.h with the kernel sources

Avinash Duduskar (2):
  bpf: Add BPF_FIB_LOOKUP_VLAN flag to bpf_fib_lookup() helper
  bpf: Add BPF_FIB_LOOKUP_VLAN_INPUT flag to bpf_fib_lookup() helper

Daniel Borkmann (1):
  bpftool: Support ML-DSA program signing

Eduard Zingerman (2):
  bpf: Do not print a newline after disassembly in bpf_verbose_insn()
  bpf: update disasm.c to print BPF_PROBE_ATOMIC as atomics

Hui Su (3):
  bpftool: Fix CPU IDs in per-CPU map output
  bpftool: Fix sparse CPU IDs in prog profile
  bpftool: Add support for BPF_F_PREORDER cgroup attach flag

Ihor Solodrai (1):
  bpftool: Don't drop a type in the sorted C dump

Jianlin Shi (1):
  docs: bpf: Document BPF_RB_OVERWRITE_POS in bpf_ringbuf_query

Jiayuan Chen (1):
  bpftool: Skip prog/map that disappears while looking it up by name

Kumar Kartikeya Dwivedi (4):
  bpf: Reject invalid LDSX instruction in disassembly
  bpf: Add file descriptor interface for program streams
  bpftool: Add option to wait for program stream output
  bpftool: Exit from the stream wait signal handler with _exit()

Leon Hwang (2):
  bpftool: Generate skeleton for global percpu data
  bpftool: Compute map size of light skeletons at runtime

Martin KaFai Lau (1):
  bpf: Add infrastructure to support attaching struct_ops to cgroups

Mykyta Yatsenko (3):
  bpftool: Accept zero perf counter snapshots
  bpftool: Group profile events by CPU
  bpftool: Scale counters and report cycles per run

Nick Hudson (2):
  bpf: Name the enum for BPF_FUNC_skb_adjust_room flags
  bpf: Add BPF_F_ADJ_ROOM_DECAP_* flags for tunnel decapsulation

Puranjay Mohan (1):
  bpf: Add bpf_call_rcu() kfunc

Tianyi Chen (1):
  bpftool: Read ring buffer maps with event_pipe

Toke Høiland-Jørgensen (2):
  bpftool: Set BPF_F_XDP_DEV_BOUND_ONLY flag non-destructively
  bpftool: Adopt bpf_program__add_flags() helper

Yuan Chen (3):
  bpftool: Fix double close in map dump
  bpftool: Fix spurious batch file read error
  bpftool: Fix bypass of the batch line length check by comments

 bash-completion/bpftool     |  65 +++----
 docs/bpftool-btf.rst        |   7 +-
 docs/bpftool-cgroup.rst     |  25 ++-
 docs/bpftool-map.rst        |  14 +-
 docs/bpftool-prog.rst       |  25 ++-
 include/uapi/linux/bpf.h    | 171 ++++++++++++++++--
 include/uapi/linux/btf.h    |  83 ++++++++-
 include/uapi/linux/const.h  |  18 ++
 src/btf.c                   | 337 +++++++++++++++++++++++++++++++++++-
 src/cgroup.c                |  99 ++++++++---
 src/common.c                |  44 +++++
 src/gen.c                   |  75 +++++---
 src/kernel/bpf/disasm.c     | 111 ++++++------
 src/main.c                  |  69 +++++---
 src/main.h                  |   4 +-
 src/map.c                   |  99 +++++++----
 src/map_perf_ring.c         |  87 +++++++---
 src/prog.c                  | 300 ++++++++++++++++++++++----------
 src/sign.c                  |  24 ++-
 src/skeleton/profiler.bpf.c |  29 ++--
 src/xlated_dumper.c         |  19 +-
 21 files changed, 1318 insertions(+), 387 deletions(-)

Signed-off-by: Quentin Monnet <qmo@kernel.org>
@qmonnet
qmonnet merged commit 83f764e into libbpf:main Oct 2, 2026
7 checks passed
@qmonnet
qmonnet deleted the bpftool-sync-2026-10-02T14-10-23.033Z branch October 2, 2026 14:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.