An off-by-a-message offset in bpf_msg_push_data()

Some of the most satisfying kernel fixes are one-liners. This is the story of one I landed in the sockmap BPF helper bpf_msg_push_data(): a single wrong operand in net/core/filter.c that corrupted the scatterlist layout of a socket message every time data was inserted into the middle of a non-first scatterlist entry. The bug was born with the helper itself in October 2018 and survived for about seven and a half years, hidden by the fact that the most common use case — prepending a header at offset zero — never triggers it.

This writeup walks through the whole thing slowly: what the helper does, how the entry-splitting path works, exactly which coordinate system each variable lives in, a concrete numeric walkthrough of the corruption, the one-line fix, and how I confirmed through git archaeology that the bug was there from day one.

What bpf_msg_push_data() does

Sockmap (and its companion sk_msg program type) is a BPF-based framework for steering and transforming socket payloads in the kernel. A sk_msg BPF program is attached to a socket map and runs when a message is sent; through a set of helpers it can inspect and reshape the in-flight message. One of those helpers is bpf_msg_push_data(), introduced by John Fastabend in commit 6fff607e2f14 ("bpf: sk_msg program helper bpf_msg_push_data", 2018):

bpf_msg_push_data(msg, offset, len, flags)

It inserts len bytes into the message at byte position offset. The canonical example from the original commit message is prepending 10 bytes at the front:

bpf_msg_push_data(msg, 0, 10, 0);

After the call, the message size has grown by len, the newly inserted region is writable scratch space the BPF program can fill, and — importantly — the packet's data pointers (data / data_end) have been invalidated, so the program must redo all bounds checks. This is the mechanism sk_msg programs use to inject metadata headers, much like XDP metadata, to pass information between the sk_msg layer and lower layers.

Internally, a sk_msg message is not a flat buffer. It is a ring of scatterlist entries (msg->sg.data[]), where each entry describes one page fragment: a struct page pointer, a page-local offset saying where inside that page the fragment begins, and a length. The logical message byte stream is the concatenation of all these fragments, walked via sk_msg_iter_var_next() around the ring between msg->sg.start and msg->sg.end.

So "insert len bytes at offset start" decomposes into a few cases:

The third case — the split path — is where the bug lived.

The split path

Here is the relevant part of the function as it existed just before the fix (parent of f72eed9b84fb). First, the helper walks the ring to locate the entry containing the insertion point:

/* First find the starting scatterlist element */
i = msg->sg.start;
do {
	offset += l;
	l = sk_msg_elem(msg, i)->length;

	if (start < offset + l)
		break;
	sk_msg_iter_var_next(i);
} while (i != msg->sg.end);

if (start > offset + l)
	return -EINVAL;

This loop establishes the two coordinate systems that matter for everything that follows:

So after the loop, the insertion point sits inside entry i, at entry-local position start - offset. Note the invariant the loop guarantees: offset <= start < offset + l.

Later, once the helper has decided there is ring space and no copy fallback is needed, the split branch runs (guarded by start - offset != 0, i.e. the insertion point is strictly inside the entry):

if (start - offset) {
	if (i == msg->sg.end)
		sk_msg_iter_var_prev(i);
	psge = sk_msg_elem(msg, i);
	rsge = sk_msg_elem_cpy(msg, i);

	psge->length = start - offset;
	rsge.length -= psge->length;
	rsge.offset += start;

	sk_msg_iter_var_next(i);
	sg_unmark_end(psge);
	sg_unmark_end(&rsge);
}

Line by line, this is what the code is supposed to do:

The crucial detail: an SG entry's offset field is page-local. It says where inside the backing page the fragment starts. When the right fragment is carved out of the same page fragment, its new page-local offset must be the original page-local offset plus the number of bytes consumed from the front — that is, offset advanced by the left fragment's length, start - offset. But the code advances it by start, the message-global insertion point.

What happens to rsge afterwards

It is worth following rsge a little further, because its later treatment is why a wrong rsge.offset is a silent corruption rather than an immediate, loud failure. After the split, the helper shifts the ring entries one or two slots forward to make room, and then, in the place_new epilogue:

place_new:
	/* Place newly allocated data buffer */
	sk_mem_charge(msg->sk, len);
	msg->sg.size += len;
	__clear_bit(new, msg->sg.copy);
	sg_set_page(&msg->sg.data[new], page, len + copy, 0);
	if (rsge.length) {
		get_page(sg_page(&rsge));
		sk_msg_iter_var_next(new);
		msg->sg.data[new] = rsge;
	}

	sk_msg_reset_curr(msg);
	sk_msg_compute_data_pointers(msg);
	return 0;

Three things to note:

Everything downstream — sk_msg_compute_data_pointers(), the verdict and redirect logic, the eventual transmit or receive path — trusts the scatterlist. A right fragment pointing 4096 bytes into the weeds is simply served up as message data.

Where it goes wrong

Let me make this concrete with a numeric walkthrough. (This is an illustrative example constructed to show the arithmetic, not a captured trace.)

Say the message ring holds two entries:

The whole message is 7096 bytes. A sk_msg program calls bpf_msg_push_data(msg, 5000, 64, 0) — insert 64 bytes at message offset 5000, which falls inside entry 1.

Trace the find-entry loop:

So offset = 4096 (message-global start of entry 1), l = 3000, and the entry-local insertion position is start - offset = 5000 - 4096 = 904. There is ring space, so we take the split branch:

psge->length = start - offset;   /* 5000 - 4096 = 904  */
rsge.length -= psge->length;    /* 3000 - 904 = 2096  */
rsge.offset += start;            /* 512 + 5000 = 5512  ← BUG */

Now look at the resulting layout:

The over-advance is exactly offset — the message-global position of the entry's beginning: start - (start - offset) = offset = 4096. The lengths still add up (904 + 2096 = 3000), so nothing that checks only lengths catches it. But the right fragment now points 4096 bytes past where it should: the 2096 bytes it claims to describe are not the tail of the original entry at all. The bytes in page range [1416, 3512) — the real tail of the message — are silently dropped from the logical stream, and the fragment may run off the end of the page entirely, so whatever consumer later walks this scatterlist reads the wrong data or wanders out of bounds. Left and right fragments no longer reconstruct the original entry.

And here is why the bug survived seven and a half years: for an insertion into the first SG entry, the find-entry loop breaks with offset = 0, so start == start - offset and the buggy line computes the correct value by accident. Prepending a header at offset 0 — the use case from the original commit message, and by far the most common one — never triggers it. You need a message that already spans multiple SG entries and an insertion point strictly inside an entry other than the first. That is a much rarer shape, and when it happens the failure is data corruption downstream rather than an obvious crash at the call site.

The fix

Advance the page-local offset by the fragment-local delta — the number of bytes the left fragment absorbed from the front of the original entry:

-rsge.offset += start;
+rsge.offset += start - offset;

The invariant check is immediate, because the line above already defined the left fragment's length as exactly that delta:

With the fix, our walkthrough's right fragment starts at page-local offset 512 + 904 = 1416 with length 2096, covering [1416, 3512) — precisely the tail of the original entry. Left fragment, new data slot, right fragment: the message is reassembled correctly.

How it was found / verified

To be honest about the process: this came out of reading the split arithmetic, not from a crash report of my own. The pattern that jumps out once you stare at the three lines is that two of them use the entry-local quantity start - offset while the third uses the message-global start — in a block where every other operand is entry-local or page-local. Mixed coordinate systems in adjacent lines of the same computation is a classic smell, and the numeric walkthrough above is the cheapest way to confirm the suspicion: pick any non-zero offset and the right fragment visibly detaches from the original entry.

The second half of verification was git archaeology, to answer "when did this break, and was it ever right?" Two commands do most of the work:

# Follow the function's whole history, diff by diff:
git log -L '/BPF_CALL_4(bpf_msg_push_data/,/^}/:net/core/filter.c'

# Inspect the commit that introduced the helper:
git show 6fff607e2f14 -- net/core/filter.c

The git log -L history shows the helper being touched many times over the years — fixes for the end mark, for copy-state bookkeeping, for bounds overestimation, for len == 0 — but the three split lines themselves are untouched since birth. And git show 6fff607e2f14 confirms it: the 2018 commit that introduced bpf_msg_push_data() already contained rsge.offset += start; verbatim. The bug was not introduced by a later refactor; it was born with the helper and had simply never been exercised hard enough in the non-first-entry case to be pinned down. That also settled the Fixes: tag: it points at 6fff607e2f14, which means the fix is a stable-backport candidate for every tree that has the helper.

I deliberately did not dress the patch up with fabricated test output — no BPF program, no KASAN splat was attached. The commit message makes the argument the same way this post does: state the two coordinate systems, state which one rsge.offset lives in, and show that the corrected operand matches the length removed from the front of the original entry.

Upstream

I posted the patch on 2026-05-24. Jakub Kicinski applied it, and it landed in mainline on 2026-05-29 as commit f72eed9b84fb ("bpf: sockmap: fix tail fragment offset in bpf_msg_push_data"). The commit carries:

Takeaway

The fix was one operand. The actual work was the audit: writing down which coordinate system each variable lives in, and noticing that one line out of three disagreed with its neighbors. Message-global versus page-local, absolute versus relative — wherever two coordinate systems meet in the same expression is where this class of bug lives. When you find such a junction in code you're reading, it is worth the ten minutes to run one concrete numeric example through it.