Open vSwitch (OVS) lets you cap the number of conntrack entries per zone — the "CT limit" feature, useful when several VMs or containers share one network namespace and you don't want a single noisy tenant to exhaust the conntrack table. The per-zone accounting state lives in kernel memory attached to the network namespace: a struct ovs_ct_limit_info hanging off the per-netns struct ovs_net.
This post walks through a slab-use-after-free I fixed in the way that state was torn down. The fix itself is only about a hundred lines, but almost none of them are where a reader would first look: the correct repair turned out to be a restructuring of the namespace teardown sequence, not another NULL check in the reader. The post goes through the bug, the vulnerable code line by line, the fix hunk by hunk, and the git archaeology that got me there.
The bug
The commit message states it compactly:
Packet processing uses CT limit state under RCU, while netns teardown
frees that state under ovs_mutex. The CT limit pointer was neither removed
from readers nor protected by a grace period, allowing packet processing to
dereference the freed state.
An unprivileged user can trigger this bug from a user and network
namespace, causing a slab-use-after-free in ovs_ct_execute() when the
netns is torn down.
Unpacking that, there are two sides touching the CT limit state with completely different locking:
- The datapath side. Every packet that hits a conntrack (
ct) action in an OVS flow goes throughovs_ct_execute()innet/openvswitch/conntrack.c. For unconfirmed connections it callsovs_ct_check_limit(), which readsovs_net->ct_limit_infoand walks the per-zone limit state. This is the packet fast path: it runs underrcu_read_lock()(taken inovs_dp_process_packet()indatapath.c), and it may not sleep or take a mutex. - The teardown side. When a network namespace dies, the pernet core runs the OVS exit callback
ovs_exit_net(), which underovs_mutex(viaovs_lock()) calledovs_ct_exit()→ovs_ct_limit_exit(), which freed the very sameovs_ct_limit_infostructure.
The critical detail: the teardown freed the structure without ever clearing ovs_net->ct_limit_info and without waiting for RCU readers. A mutex and a read-side RCU critical section provide zero mutual exclusion against each other. So if a packet was being processed on another CPU while the namespace was being torn down, the datapath could load a pointer to memory that the teardown had already handed back to the slab allocator — a slab-use-after-free in ovs_ct_execute().
What makes this more than a theoretical race is who can reach it. Creating a user namespace requires no privileges on the host, and the creator of a user namespace gets full capabilities inside it — including CAP_NET_ADMIN over any network namespaces created within. So an unprivileged user can: create a user+network namespace pair, bring up OVS inside it, install flows with conntrack actions, set zone limits, push traffic, and tear the namespace down. The whole trigger is reachable from an unprivileged shell on a default kernel with user namespaces enabled.
Reading the vulnerable code
All snippets below are the pre-fix code, taken from the parent of the fix commit (git show 403f96c32c9e~1:net/openvswitch/conntrack.c). First, the reader on the fast path:
static int ovs_ct_check_limit(struct net *net,
const struct sk_buff *skb,
const struct ovs_conntrack_info *info)
{
struct ovs_net *ovs_net = net_generic(net, ovs_net_id);
const struct ovs_ct_limit_info *ct_limit_info = ovs_net->ct_limit_info;
u32 per_zone_limit, connections;
u32 conncount_key;
conncount_key = info->zone.id;
per_zone_limit = ct_limit_get(ct_limit_info, info->zone.id);
if (per_zone_limit == OVS_CT_LIMIT_UNLIMITED)
return 0;
connections = nf_conncount_count_skb(net, skb, info->family,
ct_limit_info->data,
&conncount_key);
if (connections > per_zone_limit)
return -ENOMEM;
return 0;
}
Note the plain, un-annotated load of ovs_net->ct_limit_info — no rcu_dereference(), because the field wasn't an RCU pointer at all. The function then dereferences it twice: once to look up the per-zone limit in the hash table, once to pass ct_limit_info->data (the nf_conncount accounting tree) to nf_conncount_count_skb(). Both dereferences assume the structure is alive.
Now the teardown, as it existed before the fix:
static void ovs_ct_limit_exit(struct net *net, struct ovs_net *ovs_net)
{
const struct ovs_ct_limit_info *info = ovs_net->ct_limit_info;
int i;
nf_conncount_destroy(net, info->data);
for (i = 0; i < CT_LIMIT_HASH_BUCKETS; ++i) {
struct hlist_head *head = &info->limits[i];
struct ovs_ct_limit *ct_limit;
struct hlist_node *next;
hlist_for_each_entry_safe(ct_limit, next, head, hlist_node)
kfree_rcu(ct_limit, rcu);
}
kfree(info->limits);
kfree(info);
}
This was called from ovs_ct_exit(), which the pernet operations table wired directly into .exit:
static void __net_exit ovs_exit_net(struct net *dnet)
{
...
ovs_lock();
ovs_ct_exit(dnet);
...
}
Read it slowly and the half-finished RCU discipline jumps out:
- The individual
struct ovs_ct_limitzone entries were freed withkfree_rcu()— someone knew these hash nodes are walked by RCU readers, and did the right thing for them. - But the container structure —
info->data(destroyed vianf_conncount_destroy()), theinfo->limitsbucket array, andinfoitself — were all released synchronously, in the middle of the exit callback. - And the pointer
ovs_net->ct_limit_infowas never cleared. After this function returned, the per-netns struct still pointed at freed memory.
So the reader above could find itself in three bad states: following a live pointer to a structure whose data member was already destroyed, walking a limits array that was already freed, or reading a container struct that was already reallocated as somebody else's slab object. No synchronization primitive anywhere closes this window, because ovs_mutex and rcu_read_lock() don't exclude each other, and no grace period was inserted between "stop using" and "free". The missing invariant: a pointer read under RCU must be unpublished, and an RCU grace period must elapse, before its memory is released.
Why the obvious fixes are wrong
Two tempting patches don't work.
"Just NULL-check the reader." The reader never observes a NULL pointer in the buggy scenario — it observes a stale pointer, loaded before the free, pointing at memory the allocator may already have recycled. A NULL check only helps if the teardown first stores NULL and then waits for all in-flight readers to finish their loads before freeing; and once you've built that machinery you've reinvented RCU publication plus a grace period. A bare if (!p) return; on its own fixes nothing.
"Just take ovs_mutex in the reader." The reader is the datapath. Packets are processed under rcu_read_lock() in softirq context, potentially at line rate, on every CPU. Taking a sleeping mutex there is a non-starter — it would be illegal (you cannot sleep in an RCU read-side critical section with preemptible RCU semantics the code relies on) and it would serialize all packet forwarding on one global lock. The asymmetry is fundamental: the fast path must stay RCU, so the fix has to live entirely on the teardown side.
The right shape of the fix is therefore: publish the pointer through RCU so loads are well-defined; on teardown, detach first (store NULL, under the mutex, so no new reader can pick the state up); then wait for a grace period; then free. The pleasant surprise — and the reason the final patch adds no synchronize_rcu() call of its own — is that the pernet core already provides exactly the detach-then-free sequencing with a grace period in between. More on that below.
The fix, hunk by hunk
The merged commit is 403f96c32c9e, touching four files for +98/−46. Walking through every hunk.
datapath.h — make the pointer honestly RCU
- * @ct_limit_info: A hash table of conntrack zone connection limits.
+ * @ct_limit_info: Hash table of conntrack zone connection limits. Protected
+ * by RCU; updates and teardown are serialized by ovs_mutex. May be NULL during
+ * netns teardown.
+ * @ct_limit_exit_data: CT limit state detached at .pre_exit, freed at .exit.
...
- struct ovs_ct_limit_info *ct_limit_info;
+ struct ovs_ct_limit_info __rcu *ct_limit_info;
+ struct ovs_ct_limit_info *ct_limit_exit_data;
Two things change. The field gains the __rcu annotation so sparse and human readers both know the protocol. And a new ct_limit_exit_data field acts as a holding pen: it carries the detached state from the .pre_exit callback (where detaching happens) to the .exit callback (where freeing happens), because those are two separate functions with no shared local variable.
The datapath reader — proper rcu_dereference, and tolerate a dying netns
struct ovs_net *ovs_net = net_generic(net, ovs_net_id);
- const struct ovs_ct_limit_info *ct_limit_info = ovs_net->ct_limit_info;
+ const struct ovs_ct_limit_info *ct_limit_info;
u32 per_zone_limit, connections;
u32 conncount_key;
+ ct_limit_info = rcu_dereference(ovs_net->ct_limit_info);
+ if (!ct_limit_info)
+ return 0;
+
conncount_key = info->zone.id;
After the fix, the load is a real rcu_dereference(). The NULL check is meaningful now (unlike the naive "fix" dismissed above) because teardown stores NULL before freeing: a reader that loads NULL simply skips limit enforcement and lets the packet through. That's the correct policy — the namespace is already dying; enforcing per-zone limits on its last packets buys nothing. A reader that loaded a non-NULL pointer just before the store is protected by the grace period that follows.
ovs_ct_limit_init() — publish only fully-built state
- ovs_net->ct_limit_info = kmalloc_obj(*ovs_net->ct_limit_info);
- if (!ovs_net->ct_limit_info)
+ info = kmalloc_obj(*info);
+ if (!info)
return -ENOMEM;
...
+ rcu_assign_pointer(ovs_net->ct_limit_info, info);
return 0;
Before, the code allocated directly into the global pointer and then filled in the fields, so a concurrent reader could (in principle) observe a half-initialized structure. The rewrite builds everything in a local info — limits array, hash buckets, conncount tree — and only then publishes it with a single rcu_assign_pointer(). Readers now see either NULL or a complete object, never anything in between.
The teardown split — exit_start / exit_finish
This is the heart of the fix. The old monolithic ovs_ct_limit_exit() becomes two functions:
+static void *ovs_ct_limit_exit_start(struct ovs_net *ovs_net)
+{
+ return rcu_replace_pointer(ovs_net->ct_limit_info, NULL,
+ lockdep_ovsl_is_held());
+}
+
+/* The CT limit state must be detached by ovs_ct_limit_exit_start() and an
+ * RCU grace period must elapse before this function runs. The pernet core
+ * guarantees the grace period between the .pre_exit and .exit callbacks.
+ */
+static void ovs_ct_limit_exit_finish(struct net *net, void *data)
+{
+ const struct ovs_ct_limit_info *info = data;
int i;
+ if (!info)
+ return;
+
nf_conncount_destroy(net, info->data);
for (i = 0; i < CT_LIMIT_HASH_BUCKETS; ++i) {
...
hlist_for_each_entry_safe(ct_limit, next, head, hlist_node)
- kfree_rcu(ct_limit, rcu);
+ kfree(ct_limit);
}
kfree(info->limits);
kfree(info);
}
ovs_ct_limit_exit_start() does exactly one thing: atomically swap the published pointer to NULL (with the lockdep assertion that ovs_mutex is held, since updates are serialized by it) and hand the old pointer back. From this instant, no new RCU reader can find the state.
ovs_ct_limit_exit_finish() runs later and does the actual freeing. Note that the per-entry kfree_rcu() calls were downgraded to plain kfree(): by the time this function runs, a grace period has already elapsed, so deferring again would be pure overhead. Who guarantees that grace period? The pernet core — the next hunk wires that up.
datapath.c — hook into .pre_exit
+static void __net_exit ovs_pre_exit_net(struct net *dnet)
+{
+ ovs_lock();
+ ovs_ct_exit_start(dnet);
+ ovs_unlock();
+}
+
static void __net_exit ovs_exit_net(struct net *dnet)
{
...
ovs_lock();
- ovs_ct_exit(dnet);
+ ovs_ct_exit_finish(dnet);
...
}
static struct pernet_operations ovs_net_ops = {
.init = ovs_init_net,
+ .pre_exit = ovs_pre_exit_net,
.exit = ovs_exit_net,
...
};
The old ovs_ct_exit() is split into ovs_ct_exit_start() / ovs_ct_exit_finish() (the conntrack.h hunk is just the declaration change plus stubs for the !CONFIG_NF_CONNTRACK case). Detach happens in a brand-new .pre_exit pernet callback under ovs_mutex; freeing stays in .exit. Between the two, the pernet core in net/core/net_namespace.c runs a grace period — that's the load-bearing guarantee, and I verify it from source in the investigation section below. ovs_ct_exit_start() stashes the detached pointer in ovs_net->ct_limit_exit_data, which ovs_ct_exit_finish() picks up and frees.
SET/DEL handlers — dereference under ovs_mutex
-static int ovs_ct_limit_set_zone_limit(struct nlattr *nla_zone_limit,
- struct ovs_ct_limit_info *info)
+static int ovs_ct_limit_set_zone_limit(struct ovs_net *ovs_net,
+ struct nlattr *nla_zone_limit)
{
...
ovs_lock();
+ info = ovsl_dereference(ovs_net->ct_limit_info);
info->default_limit = zone_limit->limit;
ovs_unlock();
The netlink SET and DEL paths used to receive ct_limit_info as an argument, read once by the caller from the (previously unprotected) field. Now that the field is __rcu, the handlers take ovs_net instead and fetch the pointer with ovsl_dereference() — the "under ovs_lock" dereference helper — at each of the three mutation sites (set default limit, set zone limit, delete zone limit), inside the existing ovs_lock()/ovs_unlock() critical sections. The mutex serializes these updates against .pre_exit's rcu_replace_pointer(), so a mutation can never act on a half-detached state.
GET handler — rcu_read_lock around the traversal
+ rcu_read_lock();
+ ct_limit_info = rcu_dereference(ovs_net->ct_limit_info);
if (a[OVS_CT_LIMIT_ATTR_ZONE_LIMIT]) {
err = ovs_ct_limit_get_zone_limit(...);
- if (err)
- goto exit_err;
} else {
err = ovs_ct_limit_get_all_zone_limit(...);
- if (err)
- goto exit_err;
}
+ rcu_read_unlock();
+ if (err)
+ goto exit_err;
Previously the two GET helpers each took and dropped their own inner rcu_read_lock() around list walks, but the top-level load of ct_limit_info happened before any of that, unprotected. The fix hoists one read-side critical section to cover both the rcu_dereference() and the entire traversal, drops the now-redundant inner locking (the helpers get "Called with RCU read lock held" comments), and simplifies the error paths to a single return err / shared goto exit_err after the unlock.
Why the netlink handlers need no NULL checks
A natural review question: after the pointer can become NULL, don't SET/DEL/GET need to handle that? The commit message answers it directly, and it's worth understanding rather than pattern-matching:
The netlink command handlers do not need NULL checks because the userspace
netlink socket holds an active reference to its network namespace while a
request is processed. The per-netns exit path therefore cannot run
concurrently with SET, DEL, or GET for that socket's namespace.
While a genetlink request for namespace N is being processed, the socket pins N. cleanup_net() cannot get past the namespace's reference count, so ovs_pre_exit_net() for N simply cannot run concurrently with a CT limit netlink command for N — the pointer seen by those handlers is always live. The datapath is different: packets can still be in flight on a vport while teardown begins, which is exactly why ovs_ct_check_limit() alone got the NULL check.
How it was investigated
I want to be honest about the shape of this investigation: it was code analysis and review, not a heroic debugging session. There was no KASAN splat on my machine and no runtime reproducer logs to paste here — the bug was identified by auditing the lifetime rules of the CT limit state, and confirmed by reading the locking on both sides. The git archaeology looked like this.
First, confirm the reader and its context — that the CT limit check really is reachable from packet processing under RCU:
$ git grep -n "ovs_ct_execute" -- net/openvswitch/
net/openvswitch/actions.c:1386: err = ovs_ct_execute(ovs_dp_get_net(dp), skb, key, ...)
$ grep -n "rcu_read_lock" net/openvswitch/datapath.c
685: rcu_read_lock(); /* ovs_dp_process_packet() */
706: err = ovs_execute_actions(dp, packet, sf_acts, &flow->key);
Then trace the lifetime of the teardown function back to its origin, to find which commit introduced the pattern (needed for the Fixes: tag and the stable backport decision):
$ git log -L :ovs_ct_limit_exit:net/openvswitch/conntrack.c
...
commit 11efd5cb04a184eea4f57b68ea63dddd463158d1
Author: Yi-Hung Wei <yihung.wei@gmail.com>
Date: Thu May 24 17:56:43 2018 -0700
openvswitch: Support conntrack zone limit
$ git show 11efd5cb04a1 --stat
net/openvswitch/conntrack.c | 551 +++++++++++++++++++++++++++++++++-
...
So the pattern — synchronous free of RCU-read state, pointer left dangling — was born with the feature itself in May 2018 and survived roughly eight years and many later cleanups (git log --oneline -- net/openvswitch/conntrack.c shows the file churning around it). That earned the patch its Fixes: 11efd5cb04a1 ("openvswitch: Support conntrack zone limit") tag and a Cc: stable@vger.kernel.org.
The last piece of diligence was verifying the load-bearing assumption of the fix design: that the pernet core really does insert an RCU grace period between .pre_exit and .exit. Reading net/core/net_namespace.c:
$ grep -n "pre_exit\|synchronize_rcu" net/core/net_namespace.c
...
list_for_each_entry_continue_reverse(ops, ops_list, list) {
hold_rtnl |= !!ops->exit_rtnl;
ops_pre_exit_list(ops, net_exit_list);
}
/* Another CPU might be rcu-iterating the list, wait for it.
* This needs to be before calling the exit() notifiers, so the
* rcu_barrier() after ops_undo_list() isn't sufficient alone.
* Also the pre_exit() and exit() methods need this barrier.
*/
if (expedite_rcu)
synchronize_rcu_expedited();
else
synchronize_rcu();
...
list_for_each_entry_continue_reverse(ops, ops_list, list)
ops_exit_list(ops, net_exit_list);
cleanup_net() calls ops_undo_list(..., true), so the expedited grace period runs after all .pre_exit callbacks and before any .exit callback — the source comment even spells out "the pre_exit() and exit() methods need this barrier." That's exactly the sequencing the fix relies on, which is why ovs_ct_limit_exit_finish() can use plain kfree() with no synchronization of its own. Design verified against the implementation, not against documentation.
Upstream
The timeline, from the public record:
- 2026-07-22 — first posting to netdev and dev@openvswitch.org (cover letter on lore).
- July–August 2026 — a long review cycle; the series went through nine revisions. The final posting is [PATCH net v9] on lore, sent 2026-08-21.
- 2026-08-21 —
Reviewed-by: Ilya Maximets(the OVS userspace/upstream maintainer most active in this area), in addition toReviewed-by: Ren Weicarried from earlier rounds. Applied by Jakub Kicinski the same day. - 2026-08-24 — reached mainline as commit 403f96c32c9e24600093d7d0c61c17daeedca957 ("openvswitch: Fix CT limit teardown use-after-free").
Credits on the commit: Reported-by: Vega, Co-developed-by: Nan Li (who shares the Signed-off-by), and the Fixes: 11efd5cb04a1 plus Cc: stable@vger.kernel.org tags route it to the stable trees, since the bug has existed since the feature was merged in 2018.
Takeaway
When the fast path is RCU and the slow path frees, the bug is rarely in the reader — it's in the teardown that forgot RCU existed. And before reaching for synchronize_rcu(), check the framework you're already inside: here the pernet core had been providing exactly the right grace period all along, at exactly the right place, for anyone who split their teardown across .pre_exit and .exit.