Skip to content

[improve][test] Record Netty's allocator events and other JFR settings per profiled component, and profile on Alpine - #26787

Merged
lhotari merged 4 commits into
apache:masterfrom
lhotari:lh-improve-perf-netty-alloc-events
Sep 30, 2026
Merged

lhotari merged 4 commits into
apache:masterfrom
lhotari:lh-improve-perf-netty-alloc-events

Conversation

@lhotari

@lhotari lhotari commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Motivation

The performance tests' profiled runs (tests/performance) record JDK Flight Recorder's events alongside async-profiler, with the JDK's profile configuration, which each scenario set with async-profiler's jfrsync option. They had no way to record other JFR events, such as those of Netty's buffer allocators: Netty 4.2's AdaptivePoolingAllocator and PooledByteBufAllocator emit JFR events for their buffer and chunk allocations (io.netty.*, see Netty's microbench/src/main/resources/Netty Allocator Events.jfc). With them, a profile shows how much memory the allocators allocate and free, which buffers don't fit the pooled memory, and how often buffers grow. For example, in the IoT telemetry max-rate scenario the broker's allocator allocated and freed about 660 chunks per second, about 94 MB/s, at a steady load, which led to #26786.

Adding them brought up three more problems with the profiling setup:

  • Profiled runs used a different image. They used the glibc-based Wolfi test image, while the unprofiled runs and Pulsar's default Docker image are Alpine-based. A profile of another image doesn't carry over: the libc's memory allocator, for one, behaves differently ([improve][broker] Netty's direct memory chunk churn causes 14x more broker page faults with musl's malloc in the Alpine-based image #26786).
  • On Alpine, stacks stopped at musl. Alpine strips musl, so every frame in it read as /lib/ld-musl-x86_64.so.1, and the JNI frames above it were missing. That was 44 % of a broker's CPU samples in the max-rate scenario.
  • Failed deliveries left no logs. A run whose applications received duplicate messages counted as successful, so the launcher deleted launcher.log, which has the containers' logs. A broker that ran out of direct memory and restarted mid-run left no trace of why.

Modifications

  • JFR configurations per component. A profiled component's settings take jfrConfigurations, a list of JFR configurations that the launcher merges with the JDK's jfr configure. The default is [profile].

    • A name, such as profile or default, is one of the JDK's configurations.
    • A name ending with .jfc is a file of the new tests/performance/jfr directory.
    • none is an empty configuration.
  • JFR event settings per component. jfrEventConfig lists JFR events to enable, or event settings, each with event and optionally setting and value, such as {event: jdk.CPULoad, setting: period, value: 100 ms}. They apply after the configurations. With jfrConfigurations: [] they apply to the JDK's default configuration, as jfr configure does without --input, and with jfrConfigurations: [none] to an empty one.

  • How the launcher sets jfrsync.

    • A single configuration of the JDK without jfrEventConfig, such as the default profile, goes to async-profiler's jfrsync option as it is.
    • Otherwise, before the cluster starts, the launcher runs jfr configure in a one-off container of the component's image, so that the JDK's configurations are those of the JVM that records with them. It writes the merge into jfr-configuration.jfc beside the component's recordings and passes that file to jfrsync.
    • jfrConfigurations: [] or [none] without jfrEventConfig leaves jfrsync out, so the component records only async-profiler's samples.
    • The profiler options no longer set jfrsync, and the launcher rejects options that do.
    • When the image's JDK can't merge them, such as a released Pulsar's image whose JDK has no jfr tool, the component records with the JDK's profile configuration, and the launcher says so.
  • Netty's allocator events. netty-allocations.jfc enables all of them: io.netty.AllocateBuffer, ReallocateBuffer, FreeBuffer, AllocateChunk, FreeChunk and ReturnChunk.

    • Overhead. They're costly: there is an event for every buffer, and JFR can't sample them, since its throttle setting doesn't apply to events without @Throttle. In the max-rate scenario there were about 2.5 buffer events of each kind per message in each profiled component, and a broker recording of 1 GB instead of 10 MB.
    • Which Netty versions have them. Netty 4.2.4 and later emit them, and ReturnChunk isn't in a Netty release yet. Netty 4.1, which Pulsar 4.x uses, has none, so profiling a Pulsar 4.x image records none.
  • The Netty allocations report. A component's nettyAllocationsReport: true makes the launcher summarize its measurement recording's Netty allocator events after the run into <recording>.measurement.netty-allocator.json, which the profile report shows in a "Netty allocator events" section:

    • the events per second and per message;
    • buffer and chunk allocations by allocator, memory, and pooled or one-off chunk;
    • buffer sizes, reallocations, and the thread pools that allocate.

    The report is separate from recording the events, so other runs don't scan their recordings. A recording with only some of the events, or none, gives a summary of those it has. The summarizeNettyAllocatorEvents task summarizes other recordings.

  • Netty allocation profiles. configs/profile-{broker,gateways,applications}-netty-allocations.yaml set jfrConfigurations: [profile, netty-allocations.jfc] and nettyAllocationsReport: true for a component, in place of configs/profile-<component>.

  • Profiling on Alpine.

    • Profiled runs use the Alpine test image by default, like the unprofiled runs; -Pinttest.testImageVariant=wolfi still profiles on the Wolfi image.
    • The Alpine test image installs musl's debug symbols, musl-dbg. apk pins it to the installed musl version, and it isn't installed on the Wolfi image. With them, a stack reads, for example, Socket.recvAddress → netty_unix_socket_recvAddress → recvfrom → __syscall_cp_c. Pulsar's own Docker image is unchanged.
  • Keeping the logs of incorrect deliveries. The launcher keeps launcher.log when the applications received duplicates, ordering violations or invalid messages.

  • Documentation.

    • docs/profiling.md ("The JFR configuration", "Netty allocator events") and docs/analyzing-profiles.md ("Netty allocator events").
    • AGENTS.md, including the overhead of profiling Netty's allocations.
    • The scenario docs and the configurations' comments.
    • The docs say that asyncProfilerOptions is a comma-separated list of async-profiler's options, as the "Launch as agent" column of its ProfilerOptions.md names them.

Verifying this change

  • Make sure that the change passes the CI checks.

This change added tests and can be verified as follows:

  • Unit tests:
    • ProfilingSettingsTest covers jfrConfigurations, jfrEventConfig and nettyAllocationsReport, including the jfr configure command line, default and none, [] and [none] turning JFR off, and rejecting invalid settings and a jfrsync in the options.
    • NettyAllocatorEventsTest records stand-in events with Netty's event names and fields, and checks the summary and the report section, also for a recording without them.
    • DeliveryCheckTest covers keeping launcher.log.
  • jfr configure checks in the test image. Merging profile with netty-allocations.jfc, and applying event settings such as +jdk.CPULoad#period=100 ms after it, gives the expected effective settings, which override those of the configurations.
  • Profiled runs of the IoT telemetry max-rate scenario on one host, on the Alpine image:
    • The broker with the default configuration and the applications with netty-allocations.jfc and nettyAllocationsReport: true. The broker recorded with jfrsync=profile, with no merge and no Netty summary. The applications recorded with the merged jfr-configuration.jfc, and their profile report's Netty section shows their buffer events, about 0.05 buffer allocations per message. Every message was delivered.
    • The broker with musl-dbg: CPU samples with an unnamed ld-musl frame went from 44.3 % (2,600 of 5,869) to none (0 of 5,798).
    • jfrEventConfig was verified with the unit tests and the jfr configure checks, not in a profiled run.

Does this pull request potentially affect one of the following parts:

If the box was checked, please highlight the changes

  • Dependencies (add or upgrade a dependency)
  • The public API
  • The schema
  • The default values of configurations
  • The threading model
  • The binary protocol
  • The REST endpoints
  • The admin CLI options
  • The metrics
  • Anything that affects deployment

This change was prepared with the assistance of Claude Code (claude-opus-5-5); I have reviewed and verified it.

…ce runs, merge JFR configurations per component and profile on Alpine

Profiled JVMs record JDK Flight Recorder's events with the merge of their component's jfrConfigurations, by default
[profile, pulsar.jfc]: a name is one of the JDK's configurations, and a name ending with .jfc a file of
tests/performance/jfr. Before the cluster starts, the launcher merges them with "jfr configure" in a one-off container
of the component's image, so that the JDK's configurations are those of the JVM that records with them, into
jfr-configuration.jfc beside the component's recordings, and passes it to async-profiler's jfrsync option; scenarios
no longer set jfrsync, and the launcher rejects options that do. When the image's JDK can't merge them, the component
records with the JDK's profile configuration.

pulsar.jfc adds Netty's allocator chunk and reallocation events (io.netty.*). netty-allocations.jfc adds the events of
every buffer allocation and free, which are costly, since there is an event for every buffer and JFR can't sample
them (a broker recording grew from 10 MB to 1 GB in the max-rate scenario); configs/profile-<component>-netty-allocations
profile a component with them.

After a profiled run, NettyAllocatorEvents summarizes each measurement recording's Netty allocator events into
<recording>.measurement.netty-allocator.json: events per second and per message, buffer and chunk allocations by
allocator, memory and pooled or one-off chunk, buffer sizes, reallocations and allocating thread pools. The profile
report shows it, and the summarizeNettyAllocatorEvents task summarizes other recordings.

Profiled runs use the Alpine test image by default, as the unprofiled runs and Pulsar's default image do, since a
profile of the Wolfi image, such as of its libc's memory allocation, doesn't carry over to them;
-Pinttest.testImageVariant=wolfi profiles on the glibc-based image. The Alpine test image installs musl's debug
symbols (musl-dbg), which Alpine strips from musl: without them, 44 % of a broker's CPU samples ended in an unnamed
/lib/ld-musl-x86_64.so.1 frame, and the JNI frames above them were missing; with them, stacks read, for example,
Socket.recvAddress, netty_unix_socket_recvAddress, recvfrom.

The launcher also keeps launcher.log, with the containers' logs, when the applications received duplicates, ordering
violations or invalid messages: such a run counted as successful and deleted it, which hid a broker that ran out of
direct memory and restarted mid-run.

Assisted-by: Claude Code (claude-opus-5-5)
@lhotari lhotari added this to the 5.0.0 milestone Sep 30, 2026
…ents, allow profiling without JFR's events

- jfrConfigurations defaults to [profile], and an empty list records without jfrsync, only async-profiler's events. A
  single configuration of the JDK is passed to jfrsync as it is, without merging.
- netty-allocations.jfc has all of Netty's allocator events, enabled; pulsar.jfc is removed.
- nettyAllocationsReport: true summarizes a component's Netty allocator events after the run, so that other runs don't
  scan their recordings for them. A recording with only some of the events, or none, gives a summary of those it has.
- configs/profile-<component>-netty-allocations set both settings.
- The docs, AGENTS.md and the configurations' comments describe the settings consistently: when jfrsync is added, the
  heavy overhead of the buffer events, that Netty 4.1, which Pulsar 4.x uses, has no allocator events, and that
  asyncProfilerOptions is a comma-separated list of async-profiler's options, as the "Launch as agent" column of its
  ProfilerOptions.md names them.

Assisted-by: Claude Code (claude-opus-5-5)
Assisted-by: Claude Code (claude-opus-5-5)
…nfigurations

A profiled component's jfrEventConfig lists JFR events to enable, or event settings, each with event and optionally
setting and value, such as {event: jdk.CPULoad, setting: period, value: 100 ms}. The launcher applies them with
"jfr configure" after merging the component's jfrConfigurations into the jfr-configuration.jfc that it passes to
async-profiler's jfrsync option. With jfrConfigurations: [], they apply to the JDK's default configuration, as
"jfr configure" starts from without --input, and jfrConfigurations: [none] starts from an empty one. Without event
settings, a single configuration of the JDK still goes to jfrsync as it is, and [] or [none] leave jfrsync out.

Assisted-by: Claude Code (claude-opus-5-5)
@lhotari lhotari changed the title [improve][test] Record Netty's allocator events in profiled performance runs, merge JFR configurations per component and profile on Alpine [improve][test] Record Netty's allocator events and other JFR settings per profiled component, and profile on Alpine Sep 30, 2026
@lhotari
lhotari merged commit 89f11e6 into apache:master Sep 30, 2026
44 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants